LLM Brand Visibility Monitor
Run fixed queries across models and track how each one represents your brand.
- Difficulty
- Moderate
- Time to result
- ~weeks to results
- Steps
- 5
- Confidence
- 99%
Create a fixed panel of queries that represent how customers research the category, compare alternatives, and ask for recommendations. Run those queries across multiple language models on a recurring schedule, preserving the complete responses and model details. Classify each result by whether the brand receives a positive recommendation, a secondary or afterthought mention, no mention, or a negative representation. Track competitors and cited sources alongside the classification. Over time, the dataset reveals volatility, provider differences, and shifts following model updates or content changes. The method adapts rank tracking to probabilistic systems without pretending that one response is a stable or universal position.
Origin
Extracted from Marketing Against The Grain through an experiment run by Zapier's SEO leader, Matt, to observe brand mentions across language models.
Core principles
- 01LLM answers vary across providers and updates.
- 02A fixed query set makes changes observable over time.
- 03Mention quality matters, not just mention presence.
- 04Optimization requires longitudinal evidence rather than occasional screenshots.
How to run it
- 1
Build the query panel
Select recurring questions that reflect customer discovery, comparison, and recommendation behavior for the brand's category.
Pro tip Include broad category prompts and specific use-case prompts.
Watch out Queries designed solely to force the brand name will produce misleading visibility scores.
- 2
Choose the model panel
Select the major language models relevant to the audience and record their model versions or available identifiers.
Pro tip Include both general chat products and search-connected experiences.
Watch out Results from different versions should not be treated as one continuous model.
- 3
Run controlled checks
Execute the same queries on a regular schedule and preserve complete outputs, timestamps, and citations where available.
Pro tip Use fresh sessions and consistent settings to reduce avoidable variance.
Watch out One run per query cannot capture probabilistic variation.
- 4
Score brand treatment
Classify each output as a positive mention, secondary mention, absence, or negative mention. Record competitor treatment using the same rules.
Pro tip Keep a written rubric and periodically audit ambiguous classifications.
Watch out Counting any mention as success ignores whether the model actually recommends the brand.
- 5
Analyze and respond
Review trends around model updates, source changes, campaigns, and competitor movement. Use the evidence to prioritize content and authority-building experiments.
Pro tip Separate durable multi-run changes from isolated response noise.
Watch out Do not promise deterministic rankings in systems that generate variable answers.
In the wild
Zapier's SEO lead runs the same queries across multiple large language models every day. For each model, he checks whether Zapier receives a positive mention, appears as an afterthought, or is omitted.
→ The team can observe how rapidly AI visibility changes and begin learning how brand influence differs from conventional search rankings.
Common mistakes
Testing only one model
A brand may be represented differently across providers, so one chatbot cannot stand in for the whole discovery environment.
Recording only presence
A passing or unfavorable mention is materially different from a strong recommendation and needs a separate score.
Reacting to isolated variance
A single changed answer may reflect probabilistic generation rather than a meaningful visibility trend.
Is it for you?
Best for
Brands whose customers increasingly use chatbots or generative search during discovery and evaluation.
Not ideal for
Teams without a stable query set or enough relevance to expect legitimate category-level mentions.
From the transcript
“he has the same queries being run across a multitude of different large language models every single day.”
“Are we a positive mention and afterthought or not mentioned at all?”
“We've seen the same thing. It's like you can't expect consistent answers.”
From the episode
Why Snap MyAI is the Sleeping Giant of AI Chat (#148)