Two AI Visibility Tools, Same Brand, Different Scores. Here's What's Actually Being Measured
Buy two AI visibility tools, point them at the same brand in the same week, and you will get two different scores. Teams treat this as a reason to distrust the category. It is better understood as a reason to read the methodology before you sign anything, because the number a platform reports is a function of four choices, and almost none of them appear on the pricing page.
1. The prompt set
Every platform runs a set of buyer questions on a schedule and counts how often you appear. The set is the measurement. A tool tracking 15 prompts and a tool tracking 56 prompts are not measuring the same thing at different resolutions. They are measuring different things, because the marginal prompts added between 15 and 56 are typically longer-tail and more specific, which is exactly where a niche brand's mention rate diverges most from a generic one.
Worse, who writes the prompts matters enormously. A prompt set generated from your own marketing copy will flatter you: it uses your framing, your category name, your differentiators. A prompt set built from how buyers actually describe the problem before they know your category exists will score you far lower, and is far more useful. When a vendor auto-generates prompts, ask which of those two they are doing.
2. The engine mix
Coverage varies from one engine to ten across the market, and the engines do not agree with each other. A brand can be well represented in Gemini and absent in Perplexity. So a composite "visibility score" is a weighted average over whichever engines that vendor happens to query, with weights you usually cannot see.
This creates a specific trap when comparing vendors: the cheaper tool that tracks four engines may report a higher score than the expensive one tracking ten, simply because the six extra engines are ones where you do badly. That is not the cheap tool being wrong. It is a smaller, more favorable sample.
3. Non-determinism and sampling
Ask the same model the same question twice and you can get two different answers, with different brands named. Any single run is a sample from a distribution, not a reading of a fixed value.
Platforms handle this differently and mostly quietly. Some run each prompt once per refresh cycle. Some run it several times and average. Some run once but refresh daily, so the averaging happens across time rather than within a cycle. A weekly-refresh tool that samples once per prompt is producing a noisier number than a daily tool, and that noise will read as movement: you will see your score "improve" and "decline" on changes you did not make. Before you attribute any week-over-week swing to your own work, find out how many samples are behind it.
4. What counts as a mention
The definitions diverge more than you would expect. Does an unlinked brand name in prose count the same as a cited source link? Does appearing sixth in a list of ten count the same as being the single recommendation? Does a mention of your parent company count? Does a negative mention count as visibility?
The better platforms separate these: mention rate, citation rate, sentiment and position tracked as distinct dimensions rather than blended into one figure. A single 0-100 number that blends them is more presentable to a board and less useful to the person who has to act on it.
How to use scores without being misled
Treat the absolute number as close to meaningless across vendors, and the trend within one vendor as the actual signal. Never benchmark your score against a competitor's score from a different tool, or against a vendor's published industry average.
Fix the prompt set early and change it rarely. Every edit resets your trend line, and the temptation to add prompts you do well on is strong. If you must change it, version it and annotate the chart.
Insist on seeing raw answers, not just aggregates. The single highest-value output of any of these tools is the actual text of the answer where a competitor beat you, because it tells you which source the engine used and therefore what to go and fix. A platform that shows you a score but not the underlying responses is selling you a dashboard, not a diagnosis.
And decide up front what the score is for. If it exists to prove ROI to a board, you will be tempted to game the prompt set, and you will succeed. If it exists to find gaps, then a falling score after you broaden the prompt set is good news: you just learned where you actually stand.