MMarketing Against The Grain
← All frameworks
Strategy

AI Chat Outcome Scorecard

Judge automation by experience, conversion, and economic value together

Difficulty
Easy
Time to result
~weeks to results
Steps
5
Confidence
97%

Evaluate an AI chat system through three complementary lenses. Customer satisfaction tests whether automation preserves or improves the visitor experience relative to human chat. Conversion from handed-off conversations to qualified leads tests whether the model is identifying and routing stronger commercial intent. Value per chat tests whether those improvements translate into meaningful economic output rather than superficial activity. These measures should be reviewed together and segmented by page because visitor intent and expected behavior vary across environments. A system that reduces labor but harms satisfaction fails the scorecard; so does a system that delights visitors without improving the intended business outcome. HubSpot treated parity with human satisfaction as a major milestone, then validated the broader hypothesis through increased qualified-lead conversion and value per conversation.

Origin

Extracted from Marketing Against The Grain from the success criteria and results of HubSpot's AI chat experiment.

Core principles

  • 01Automation must preserve the quality of the customer experience
  • 02Better routing should improve the quality of human-assisted demand
  • 03Commercial value and customer satisfaction must be measured together
  • 04Human performance provides a meaningful benchmark for automation

How to run it

  1. 1

    Set the Experience Benchmark

    Record the customer-satisfaction level achieved by human-handled chats and define the minimum acceptable automated performance.

    Pro tip Use comparable pages and visitor intents when establishing the baseline.

    Watch out An aggregate benchmark can hide weak performance for important visitor groups.

  2. 2

    Measure Routing Quality

    Track how often visitors passed to humans become qualified leads or reach another defined high-value outcome.

    Pro tip Compare rates rather than raw lead totals when traffic allocation changes.

    Watch out Higher handoff volume does not necessarily mean better qualification.

  3. 3

    Calculate Economic Yield

    Determine the revenue or expected value associated with each chat and compare it with the baseline.

    Pro tip Segment value per chat by page because intent and product economics can differ.

    Watch out Do not attribute all downstream revenue to chat without a consistent attribution rule.

  4. 4

    Review the Metrics Jointly

    Accept the experiment only when experience quality and intended business outcomes move within agreed boundaries.

    Pro tip Create explicit guardrails so a commercial gain cannot silently override a customer-experience decline.

    Watch out Optimizing a single KPI invites harmful trade-offs.

  5. 5

    Track Performance Over Time

    Monitor whether satisfaction, conversion, and value remain stable as the system expands to new pages and question types.

    Pro tip Retain page-level reporting after rollout.

    Watch out Initial success in a bounded knowledge base does not guarantee success in open-ended sales conversations.

In the wild

From Early Satisfaction Decline to Commercial Lift

HubSpot initially saw customer satisfaction decline after releasing its AI bot. After annotation and tuning, satisfaction reached parity with human chats. The team then reported a 43% increase in conversion to qualified leads and increases of more than 50% in value per chat on some pages.

The combined scorecard showed that automation preserved experience while improving lead quality and economic performance.

Common mistakes

Treating Cost Savings as Success

Lower staffing requirements can conceal worse answers, weaker trust, or lost high-intent opportunities.

Ignoring the Human Benchmark

A bot may improve over its own earlier version while remaining materially worse than the existing customer experience.

Averaging Across Every Page

Aggregate performance can conceal poor outcomes in high-intent or strategically important environments.

Is it for you?

Best for

It is best for customer-facing automation programs that affect support quality, lead routing, and revenue simultaneously.

Not ideal for

It is not ideal for internal AI tools whose outcomes have no customer-satisfaction or revenue component.

From the transcript

We'll probably see an increased value per chat.

Emmy Jonathan · 10:00

And now what we see over time is that our AI chat bots are on par with our human chats.

Emmy Jonathan · 11:30

We saw meaningful increases in the conversion rate between people, chatters who are being passed to the ISC team and the conversion of those people…

Emmy Jonathan · 12:00

From the episode

How We Hacked Hubspot With Ai To Make Free Money