High-Volume, Low-Risk AI Test Selection
Choose bounded, well-supported environments that generate fast learning loops
- Difficulty
- Easy
- Time to result
- ~weeks to results
- Steps
- 5
- Confidence
- 98%
Evaluate candidate AI environments on three connected dimensions: traffic volume, failure risk, and knowledge quality. First, identify what visitors are trying to accomplish in each environment because different pages attract different personas and levels of intent. Next, favor an environment with enough usage to produce a fast feedback loop. Constrain that choice by asking how damaging an inaccurate response would be and whether the model can ground its answers in a strong, bounded repository. HubSpot selected its knowledge base because it combined substantial traffic with product-focused questions and extensive documentation. The result is a practical starting point where unstructured questions remain valuable to the model, but their subject matter and possible answers are structured enough to support controlled learning.
Origin
Extracted from Marketing Against The Grain during HubSpot's explanation of how it selected the first page for AI-powered chat.
Core principles
- 01Begin with a bounded problem rather than unrestricted user intent
- 02Use high traffic to accelerate evidence collection
- 03Reduce hallucination risk with authoritative source material
- 04Match each experiment to the intent of its environment
How to run it
- 1
Map Visitor Intent
List candidate pages or workflows and define the persona, intent level, and job associated with each one.
Pro tip Separate research, support, pricing, and purchase intent rather than treating all website traffic alike.
Watch out A page's total traffic says little about whether its users share a coherent intent.
- 2
Measure Learning Volume
Estimate how many real interactions each candidate can generate and how quickly those interactions will support iteration.
Pro tip Prefer recurring, high-frequency questions that make patterns visible quickly.
Watch out A low-volume test can appear safe while taking too long to reveal important failures.
- 3
Bound the Failure Risk
Assess the consequences of hallucinations, irrelevant answers, and incorrect recommendations in each candidate environment.
Pro tip Start where mistakes can be detected and corrected without causing irreversible harm.
Watch out Do not equate technical feasibility with acceptable customer risk.
- 4
Audit Grounding Material
Confirm that accurate, comprehensive, and maintained content exists to support answers in the chosen environment.
Pro tip Use a repository whose scope closely matches the questions visitors normally ask.
Watch out Large quantities of stale or loosely related content do not create reliable grounding.
- 5
Select and Instrument
Choose the strongest high-volume, low-risk candidate and define the interactions, outcomes, and corrections that will be captured.
Pro tip Design the data collection process before exposing the model to live users.
Watch out Launching without an explicit feedback loop wastes the main advantage of a high-volume environment.
In the wild
HubSpot compared pages with different user intents and selected its knowledge base. The page attracted substantial traffic, visitors mostly asked bounded product questions, and extensive knowledge-base articles gave the model reliable material for answering them.
→ The team obtained frequent real-world interactions while limiting the risks associated with an early customer-facing model.
Common mistakes
Choosing Traffic Without Considering Risk
The busiest environment may expose the company to costly errors if questions are open-ended or purchase-critical.
Testing Without Authoritative Content
A model cannot reliably resolve bounded questions when the underlying documentation is incomplete, inaccurate, or inaccessible.
Ignoring Page-Specific Intent
Treating every website page as equivalent hides major differences in personas, questions, and expected outcomes.
Is it for you?
Best for
It is best for teams choosing among several customer-facing AI use cases with different traffic, intent, and failure consequences.
Not ideal for
It is not ideal when every available use case is safety-critical or lacks reliable source material.
From the transcript
“But we also wanted to test on a page that was lower risk, because this was a very new motion for us.”
“And so the page that we decided to start on was the knowledge base, because it was a very high volume page as far as…”
“And then the third thing you said that I think is super important is enough people have to be going there that you get the…”
From the episode
How We Hacked Hubspot With Ai To Make Free Money