MMarketing Against The Grain
← All frameworks
Marketing

Hypothesis-Driven A/B Testing

Use experiments to explain customer behavior, not merely move a metric

Difficulty
Moderate
Time to result
~weeks to results
Steps
5
Confidence
99%

Frame each A/B test as customer research. Start with an observed customer problem and propose a causal explanation for the behavior, then design a variation that directly tests that explanation. Define the metric as evidence of whether the mechanism worked, not as the hypothesis itself. After the test, explain why the result supports or challenges the original insight and record what the organization can reuse elsewhere. This approach distinguishes meaningful experimentation from random color changes or repeated attempts to force a five-percent lift. It makes even unsuccessful tests valuable because they eliminate incorrect explanations and sharpen the team's understanding of customer needs.

Origin

Extracted from Marketing Against The Grain's discussion of Brian Chesky's objection to metric-only A/B testing at Airbnb.

Core principles

  • 01Begin with a customer insight rather than a random variation
  • 02State why the proposed change should affect behavior
  • 03Treat metric movement as an outcome, not the insight
  • 04Use experiments to prove or disprove a mechanism
  • 05Carry validated learning into future decisions

How to run it

  1. 1

    Observe a customer problem

    Use journey evidence, feedback, or behavior to identify a specific obstacle or unmet need.

    Pro tip Describe what the customer is trying to accomplish before discussing interface changes.

    Watch out Do not begin with a favorite variation.

  2. 2

    Form a causal hypothesis

    State why a proposed change should improve the customer's ability or motivation to complete the task.

    Pro tip Use an if-then-because structure to make the mechanism explicit.

    Watch out A target such as increasing conversion by five percent is an outcome, not a hypothesis.

  3. 3

    Design the discriminating test

    Create a control and variation whose key difference isolates the proposed mechanism.

    Pro tip Change only what is necessary to test the customer insight.

    Watch out Large bundles of unrelated changes obscure what caused the result.

  4. 4

    Measure the outcome

    Run the test against a predetermined success metric and sufficient sample.

    Pro tip Include guardrail metrics for retention, quality, or downstream behavior.

    Watch out Do not stop early when preliminary results match expectations.

  5. 5

    Explain and preserve the learning

    Decide whether the evidence proved, disproved, or left the hypothesis unresolved, then document the reusable insight.

    Pro tip Apply validated insights to adjacent pages or product experiences.

    Watch out Do not report only the percentage lift.

In the wild

Testing clarity rather than button color

A signup page has high abandonment after visitors encounter an unfamiliar pricing term. The team hypothesizes that uncertainty about billing creates hesitation and tests a plain-language explanation beside the call to action. It measures completed signups and early cancellations rather than testing an arbitrary button color.

The team learns whether billing clarity changes customer behavior and gains an insight usable throughout the purchase journey.

Common mistakes

Random-element testing

Changing a green button to red without a customer-based rationale may produce noise but little reusable learning.

Calling the target the hypothesis

A desired five-percent increase says what the team wants, not why customer behavior should change.

Reporting lift without causality

A winning variation is less valuable when the team cannot articulate the customer insight behind it.

Is it for you?

Best for

Growth, marketing, and product teams running recurring controlled experiments.

Not ideal for

Situations with insufficient traffic, unreliable measurement, or no plausible causal hypothesis.

From the transcript

if you're going to do an A B test, it should be a hypothesis driven test.

Kieran Flanagan · 11:00

That is actually the outcome, not the insight.

Kieran Flanagan · 11:00

if an A B test works, you should be able to say the why.

Kieran Flanagan · 11:30

From the episode

Airbnb Just Copied Apple’s Product Development Strategy... Here’s Why (#138)