MMarketing Against The Grain
← All frameworks
Marketing

AI Search Geographic Split Test

Compare US and international traffic to isolate AI search effects

Difficulty
Easy
Time to result
~weeks to results
Steps
5
Confidence
96%

This natural-experiment framework uses a geographically limited product rollout to estimate how AI-driven search changes user behavior. When the new experience is initially available only in the United States, marketers segment search performance into US and non-US cohorts. They establish baseline trends, then inspect traffic and behavioral metrics daily or several times per week. The non-US cohort serves as a provisional comparison for legacy search, while the US cohort reflects greater exposure to AI-generated results. Differences are investigated rather than automatically attributed to AI, because seasonality, market mix, rankings, and campaigns may also matter. The resulting evidence guides focused optimization plays while ongoing monitoring reveals whether those changes improve visibility, engagement, or traffic under the emerging search experience.

Origin

Extracted from Marketing Against The Grain as Kipp Bodner described Google's US-first AI search rollout as a split-test opportunity for businesses.

Core principles

  • 01Use phased rollouts as natural experiments
  • 02Separate exposed traffic from an appropriate comparison group
  • 03Measure trends repeatedly rather than relying on one snapshot
  • 04Translate observed differences into search experiments
  • 05Preserve conventional SEO while gathering evidence

How to run it

  1. 1

    Define rollout exposure

    Verify where and when the AI search experience is being introduced. Identify an exposed geographic cohort and a credible comparison cohort.

    Pro tip Record major rollout changes because cohort exposure may expand over time.

    Watch out Do not assume every user in an announced market receives the feature simultaneously.

  2. 2

    Create geographic segments

    Split organic search data into US and non-US traffic, then preserve consistent filters across reporting periods.

    Pro tip Also segment by device or query type if those dimensions differ substantially by geography.

    Watch out Aggregate international traffic may conceal large differences between individual markets.

  3. 3

    Establish the baseline

    Measure pre-rollout or early-rollout trends for traffic, click-through behavior, conversions, and query visibility in both cohorts.

    Pro tip Use several prior weeks when available to understand normal volatility.

    Watch out A single prior day is not a reliable baseline.

  4. 4

    Monitor frequently

    Review the two cohorts daily or multiple times each week and note when their trends diverge.

    Pro tip Annotate promotions, ranking changes, site releases, and other events that could affect only one cohort.

    Watch out Correlation with rollout timing alone does not prove causation.

  5. 5

    Build optimization plays

    Use repeated, query-level evidence to propose changes for AI-driven search, then test those changes against measurable outcomes.

    Pro tip Start with pages and queries showing the clearest exposure-related decline or visibility shift.

    Watch out Do not abandon proven search practices based on an early, noisy signal.

In the wild

Comparing search traffic after rollout

A software company receives substantial organic traffic from the US, UK, and Canada. After AI search begins rolling out in the US, the team creates separate dashboards for US and non-US traffic and compares query-level clicks, impressions, and conversions against the previous month. It finds that informational US queries lose clicks while high-intent product queries remain stable, so it tests more distinctive evidence and brand-led material on the affected pages.

The company identifies where AI answers are most likely changing click behavior and focuses optimization work on the exposed query class.

Common mistakes

Treating every difference as an AI effect

Geographic demand, campaigns, rankings, seasonality, and device mix can produce divergent trends even without an AI search rollout.

Comparing incompatible markets

A control cohort is weak when its audience, language, product availability, or historical behavior differs dramatically from the exposed cohort.

Checking only total traffic

Aggregate traffic can hide different outcomes across informational, branded, and transactional queries.

Is it for you?

Best for

Businesses with meaningful organic traffic from both the United States and countries where the AI search rollout has not yet occurred.

Not ideal for

Businesses with negligible international traffic or major geographic differences that make the comparison groups fundamentally incompatible.

From the transcript

Google's essentially giving most businesses a split test right because they're only rolling out AI search in the US initially

Kipp Bodner · 28:30

daily multiple times a week look at my search traffic Trends and especially cut it by US versus non- us

Kipp Bodner · 29:00

based on those learnings I would start putting plays together to hopefully better optimize for that a search engine

Kipp Bodner · 29:00

From the episode

Ai Agents: The Future Marketers You Can't Afford to Ignore