MMarketing Against The Grain
← All frameworks
Marketing

LLM Brand Visibility Monitor

Run fixed queries across models and track how each one represents your brand.

Difficulty
Moderate
Time to result
~weeks to results
Steps
5
Confidence
99%

Create a fixed panel of queries that represent how customers research the category, compare alternatives, and ask for recommendations. Run those queries across multiple language models on a recurring schedule, preserving the complete responses and model details. Classify each result by whether the brand receives a positive recommendation, a secondary or afterthought mention, no mention, or a negative representation. Track competitors and cited sources alongside the classification. Over time, the dataset reveals volatility, provider differences, and shifts following model updates or content changes. The method adapts rank tracking to probabilistic systems without pretending that one response is a stable or universal position.

Origin

Extracted from Marketing Against The Grain through an experiment run by Zapier's SEO leader, Matt, to observe brand mentions across language models.

Core principles

  • 01LLM answers vary across providers and updates.
  • 02A fixed query set makes changes observable over time.
  • 03Mention quality matters, not just mention presence.
  • 04Optimization requires longitudinal evidence rather than occasional screenshots.

How to run it

  1. 1

    Build the query panel

    Select recurring questions that reflect customer discovery, comparison, and recommendation behavior for the brand's category.

    Pro tip Include broad category prompts and specific use-case prompts.

    Watch out Queries designed solely to force the brand name will produce misleading visibility scores.

  2. 2

    Choose the model panel

    Select the major language models relevant to the audience and record their model versions or available identifiers.

    Pro tip Include both general chat products and search-connected experiences.

    Watch out Results from different versions should not be treated as one continuous model.

  3. 3

    Run controlled checks

    Execute the same queries on a regular schedule and preserve complete outputs, timestamps, and citations where available.

    Pro tip Use fresh sessions and consistent settings to reduce avoidable variance.

    Watch out One run per query cannot capture probabilistic variation.

  4. 4

    Score brand treatment

    Classify each output as a positive mention, secondary mention, absence, or negative mention. Record competitor treatment using the same rules.

    Pro tip Keep a written rubric and periodically audit ambiguous classifications.

    Watch out Counting any mention as success ignores whether the model actually recommends the brand.

  5. 5

    Analyze and respond

    Review trends around model updates, source changes, campaigns, and competitor movement. Use the evidence to prioritize content and authority-building experiments.

    Pro tip Separate durable multi-run changes from isolated response noise.

    Watch out Do not promise deterministic rankings in systems that generate variable answers.

In the wild

Zapier daily query tracking

Zapier's SEO lead runs the same queries across multiple large language models every day. For each model, he checks whether Zapier receives a positive mention, appears as an afterthought, or is omitted.

The team can observe how rapidly AI visibility changes and begin learning how brand influence differs from conventional search rankings.

Common mistakes

Testing only one model

A brand may be represented differently across providers, so one chatbot cannot stand in for the whole discovery environment.

Recording only presence

A passing or unfavorable mention is materially different from a strong recommendation and needs a separate score.

Reacting to isolated variance

A single changed answer may reflect probabilistic generation rather than a meaningful visibility trend.

Is it for you?

Best for

Brands whose customers increasingly use chatbots or generative search during discovery and evaluation.

Not ideal for

Teams without a stable query set or enough relevance to expect legitimate category-level mentions.

From the transcript

he has the same queries being run across a multitude of different large language models every single day.

Kieran Flanagan · 15:00

Are we a positive mention and afterthought or not mentioned at all?

Kieran Flanagan · 15:00

We've seen the same thing. It's like you can't expect consistent answers.

Kip Bodner · 15:30

From the episode

Why Snap MyAI is the Sleeping Giant of AI Chat (#148)