MMarketing Against The Grain
← All frameworks
Marketing

Monthly LLM Brand Poll

Track how AI models discover, rank, and describe your brand each month.

Difficulty
Easy
Time to result
~months to results
Steps
4
Confidence
90%

Create a recurring benchmark of how major language models and AI agents represent your company. Start with a stable set of prompts covering product recommendations, category leaders, buyer-specific needs, brand associations, strengths, weaknesses, and comparisons. Ask the same questions of each selected model every month, preserving the complete responses and recording structured indicators such as inclusion, rank, sentiment, cited attributes, and competitor mentions. Compare each new run with the historical baseline to identify persistent patterns or meaningful movement. The mechanism turns opaque AI recommendations into a longitudinal monitoring practice, although it cannot explain every ranking change or establish how real customers feel. Treat the results as an emerging discovery signal that complements search analytics, customer interviews, focus groups, and conventional brand tracking.

Origin

Extracted from Marketing Against The Grain during a discussion about how agents and language models may reshape product discovery and brand measurement.

Core principles

  • 01AI assistants increasingly mediate product discovery.
  • 02Consistent prompts make month-to-month comparisons more useful.
  • 03Recommendation visibility and brand perception are separate signals.
  • 04Model outputs reveal patterns, not definitive consumer truth.
  • 05Archived results matter more than isolated responses.

How to run it

  1. 1

    Define the benchmark prompts

    Write a fixed series of questions that test category recommendations, brand perception, buyer needs, and comparisons with competitors. Keep the wording stable so changes are more likely to reflect model behavior than prompt variation.

    Pro tip Include both unaided questions and prompts that mention the brand explicitly.

    Watch out Do not use leading prompts that predetermine a favorable answer.

  2. 2

    Poll relevant models monthly

    Submit the benchmark to each model or agent on a consistent schedule. Save complete answers alongside the model name, date, and any available version information.

    Pro tip Use fresh conversations to reduce contamination from prior context.

    Watch out Model updates and nondeterministic outputs can create noise.

  3. 3

    Score discovery and perception

    Record whether the brand appears, its position in recommendation lists, the attributes attached to it, and the competitors mentioned. Separate recommendation visibility from broader perception signals.

    Pro tip Use a simple recurring scorecard rather than changing metrics after every run.

    Watch out An LLM response is not equivalent to measured customer sentiment.

  4. 4

    Compare trends and investigate changes

    Review results against previous months and flag sustained changes across multiple prompts or models. Use customer research, content analysis, and market evidence to test possible explanations.

    Pro tip Prioritize repeated cross-model movement over a single surprising answer.

    Watch out Correlation between marketing activity and an output change does not prove causation.

In the wild

Monitoring an emerging CRM brand

A CRM startup asks four major models the same 12 questions each month, including which tools suit small agencies and how its brand compares with established platforms. It logs recommendation rank, recurring attributes, objections, and cited competitors. After three months, the company appears more often for automation-focused prompts but remains absent from enterprise-security queries, prompting a targeted evidence and content initiative.

The team gains a repeatable view of AI discovery gaps and a focused hypothesis to validate through customer and market research.

Common mistakes

Treating one answer as a trend

A single response may reflect randomness, prompt wording, or a temporary model version. Require repetition across time, questions, or models before acting.

Equating model output with customer opinion

Language models synthesize online material under their own constraints; they do not directly measure buyer perception. Pair the poll with interviews, surveys, or established brand studies.

Changing the benchmark every month

Frequent prompt changes destroy comparability. Preserve a stable core and label any experimental prompts separately.

Is it for you?

Best for

It is best for marketers and founders preparing for product discovery through AI assistants and agents.

Not ideal for

It is not ideal as a standalone substitute for customer research, brand-lift studies, or causal attribution.

From the transcript

I do think it's good to start to PLL them every month and keep a record

Kieran Flanagan · 20:30

just pull them every month look to see like have a series of questions

Kieran Flanagan · 20:30

you can pull it to see not just are you in the top x% of products that it recommends for certain queries but also PLL…

Kieran Flanagan · 21:00

From the episode

OpenAI Agents 2.0… Is This The End Of Google Search?