Monthly LLM Brand Poll
Track how AI models discover, rank, and describe your brand each month.
- Difficulty
- Easy
- Time to result
- ~months to results
- Steps
- 4
- Confidence
- 90%
Create a recurring benchmark of how major language models and AI agents represent your company. Start with a stable set of prompts covering product recommendations, category leaders, buyer-specific needs, brand associations, strengths, weaknesses, and comparisons. Ask the same questions of each selected model every month, preserving the complete responses and recording structured indicators such as inclusion, rank, sentiment, cited attributes, and competitor mentions. Compare each new run with the historical baseline to identify persistent patterns or meaningful movement. The mechanism turns opaque AI recommendations into a longitudinal monitoring practice, although it cannot explain every ranking change or establish how real customers feel. Treat the results as an emerging discovery signal that complements search analytics, customer interviews, focus groups, and conventional brand tracking.
Origin
Extracted from Marketing Against The Grain during a discussion about how agents and language models may reshape product discovery and brand measurement.
Core principles
- 01AI assistants increasingly mediate product discovery.
- 02Consistent prompts make month-to-month comparisons more useful.
- 03Recommendation visibility and brand perception are separate signals.
- 04Model outputs reveal patterns, not definitive consumer truth.
- 05Archived results matter more than isolated responses.
How to run it
- 1
Define the benchmark prompts
Write a fixed series of questions that test category recommendations, brand perception, buyer needs, and comparisons with competitors. Keep the wording stable so changes are more likely to reflect model behavior than prompt variation.
Pro tip Include both unaided questions and prompts that mention the brand explicitly.
Watch out Do not use leading prompts that predetermine a favorable answer.
- 2
Poll relevant models monthly
Submit the benchmark to each model or agent on a consistent schedule. Save complete answers alongside the model name, date, and any available version information.
Pro tip Use fresh conversations to reduce contamination from prior context.
Watch out Model updates and nondeterministic outputs can create noise.
- 3
Score discovery and perception
Record whether the brand appears, its position in recommendation lists, the attributes attached to it, and the competitors mentioned. Separate recommendation visibility from broader perception signals.
Pro tip Use a simple recurring scorecard rather than changing metrics after every run.
Watch out An LLM response is not equivalent to measured customer sentiment.
- 4
Compare trends and investigate changes
Review results against previous months and flag sustained changes across multiple prompts or models. Use customer research, content analysis, and market evidence to test possible explanations.
Pro tip Prioritize repeated cross-model movement over a single surprising answer.
Watch out Correlation between marketing activity and an output change does not prove causation.
In the wild
A CRM startup asks four major models the same 12 questions each month, including which tools suit small agencies and how its brand compares with established platforms. It logs recommendation rank, recurring attributes, objections, and cited competitors. After three months, the company appears more often for automation-focused prompts but remains absent from enterprise-security queries, prompting a targeted evidence and content initiative.
→ The team gains a repeatable view of AI discovery gaps and a focused hypothesis to validate through customer and market research.
Common mistakes
Treating one answer as a trend
A single response may reflect randomness, prompt wording, or a temporary model version. Require repetition across time, questions, or models before acting.
Equating model output with customer opinion
Language models synthesize online material under their own constraints; they do not directly measure buyer perception. Pair the poll with interviews, surveys, or established brand studies.
Changing the benchmark every month
Frequent prompt changes destroy comparability. Preserve a stable core and label any experimental prompts separately.
Is it for you?
Best for
It is best for marketers and founders preparing for product discovery through AI assistants and agents.
Not ideal for
It is not ideal as a standalone substitute for customer research, brand-lift studies, or causal attribution.
From the transcript
“I do think it's good to start to PLL them every month and keep a record”
“just pull them every month look to see like have a series of questions”
“you can pull it to see not just are you in the top x% of products that it recommends for certain queries but also PLL…”
From the episode
OpenAI Agents 2.0… Is This The End Of Google Search?