Three numbers get misread constantly: a mention counted as a recommendation, a single-run result treated as a rate, and a branded prompt result reported as discovery. Each one inflates the picture in a different direction.
The fix is not a better tool. It is knowing which question each number can answer and refusing to let it answer the others.
A citation report is easy to read badly because every number in it looks like the same kind of number. They are percentages, they sit in the same column, and nothing on the page tells you that one of them is a rate and one of them is a coin flip.
What are the three numbers people misread?
They fail independently, which is why fixing one does not protect you from the other two. A report can be honest about sampling and still tell you nothing, because the prompts were branded.
- What the number says
- Your name appeared in the answer.
- What it does not say
- Whether it appeared as the recommendation or as the thing being compared against.
- What the number says
- On this occasion, the model produced this answer.
- What it does not say
- What it does most of the time, which is the only thing a percentage can mean.
- What the number says
- Asked about you by name, the engine knows who you are.
- What it does not say
- Whether anyone who did not already know your name would ever reach you.
Is a mention the same as a recommendation?
No, and the gap between them is where most of the useful diagnosis lives. Three distinct things can happen to your brand inside one answer, and they carry different weight.
| Outcome | What happened | What it tells you |
|---|---|---|
| Mentioned | Your name appears in the prose, with no link. | The model holds your brand in memory for this category. It says nothing about your site. |
| Cited | A specific page of yours is attached as a source. | Your content was retrieved and judged usable for this question. |
| Recommended | You are named as the answer, not as context. | The one that moves pipeline. Also the rarest, and the one most reports do not separate. |
The distinction is not academic. Anthropic’s documentation on Claude’s citation feature describes citations as returning the exact passages that support each claim, which is a retrieval event with a specific page behind it. A mention has no page behind it at all. Microsoft draws the same line from the other side: the AI Performance report in Bing Webmaster Tools counts “the total number of citations that are displayed as sources,” so a Copilot answer that names you without linking you never enters that count.
If your report collapses all three into one “visibility” figure, you cannot tell an entity problem from a content problem, which is the first fork in entity problem or content problem. A brand mentioned everywhere and cited nowhere has a content problem. A brand cited but never named has the opposite.
Why does one run tell you almost nothing?
Because the same question asked twice does not reliably produce the same answer. This is documented behaviour, not a quirk of any one product. OpenAI’s own advanced usage guide states plainly that “Chat Completions are non-deterministic by default (which means model outputs may differ from request to request),” and even with a fixed seed the documentation offers only mostly deterministic output rather than a guarantee.
Google says something similar about its own surfaces. Its documentation on AI features notes that “AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary.”
“This engine can produce an answer that names us.” One observation is enough to prove something is possible.
“We appear 40% of the time.” A percentage derived from one observation is either 0% or 100%, and neither is a measurement.
This is why run count belongs on the methodology page of any tool you are evaluating, and why so few publish it. A tool sampling once a day and a tool sampling five times will report different numbers from identical reality.
Does a branded prompt measure discovery?
It measures recall, which is a different thing and usually the thing you already know. Asking an engine “what does Rivarise do” tests whether your brand is resolved in the model. Asking “what tools track brand visibility in AI answers” tests whether you get discovered by somebody who has never heard of you.
The second question is the one with revenue attached, and it is the one a branded prompt cannot answer. If your prompt set is mostly branded, your report will look strong and your pipeline will not move. The unbranded set in prompts to test AI brand visibility is built for exactly this, and it deliberately contains no company names.
How do you read the report properly?
Six checks, in order. The first three establish whether the number means anything. The last three establish what it means.
Run the same six checks on any tool you are evaluating, including this one. Several vendors in the Peec AI alternatives roundup do not publish a run count anywhere, which is a finding about the tool rather than an omission on your part.
What should you do with a number you distrust?
Reproduce it by hand before you act on it. Take the prompt, run it five times in the engine itself, and read all five answers. If the tool says 40% and you see two out of five, the tool is measuring. If you see five out of five or zero out of five, ask how many runs went into that 40%.
A dashboard you cannot reproduce by hand is a dashboard you cannot defend in a meeting.
When the answer text disagrees with the score, trust the text. That is the whole method in AI giving wrong information about my product, and it applies here: the raw answer is evidence, the score is an interpretation of evidence, and only one of them is a primary source. If the two conflict on your own product, they conflict on your competitors too, and the same reading applies to AI recommends competitor instead of us.
If none of the six checks resolve it, the problem is upstream of the report. Start at why is my brand not showing up in AI search and work forward, because a measurement problem and a visibility problem produce the same flat line.
