Blog
6 min read Diagnose Your AI Visibility

How to Read a Citation Report Without Fooling Yourself

On this page
  1. What are the three numbers people misread?
  2. Is a mention the same as a recommendation?
  3. Why does one run tell you almost nothing?
  4. Does a branded prompt measure discovery?
  5. How do you read the report properly?
  6. What should you do with a number you distrust?
The short version

Three numbers get misread constantly: a mention counted as a recommendation, a single-run result treated as a rate, and a branded prompt result reported as discovery. Each one inflates the picture in a different direction.

The fix is not a better tool. It is knowing which question each number can answer and refusing to let it answer the others.

A citation report is easy to read badly because every number in it looks like the same kind of number. They are percentages, they sit in the same column, and nothing on the page tells you that one of them is a rate and one of them is a coin flip.

What are the three numbers people misread?

They fail independently, which is why fixing one does not protect you from the other two. A report can be honest about sampling and still tell you nothing, because the prompts were branded.

Three ways a visibility number lies to you
1 Mention read as endorsement
What the number says
Your name appeared in the answer.
What it does not say
Whether it appeared as the recommendation or as the thing being compared against.
2 One run read as a rate
What the number says
On this occasion, the model produced this answer.
What it does not say
What it does most of the time, which is the only thing a percentage can mean.
3 Branded prompt read as discovery
What the number says
Asked about you by name, the engine knows who you are.
What it does not say
Whether anyone who did not already know your name would ever reach you.
Each failure inflates the number, never deflates it. That is why a report almost never reads worse than reality, and why a flattering dashboard deserves more suspicion than a bleak one.

Is a mention the same as a recommendation?

No, and the gap between them is where most of the useful diagnosis lives. Three distinct things can happen to your brand inside one answer, and they carry different weight.

OutcomeWhat happenedWhat it tells you
MentionedYour name appears in the prose, with no link.The model holds your brand in memory for this category. It says nothing about your site.
CitedA specific page of yours is attached as a source.Your content was retrieved and judged usable for this question.
RecommendedYou are named as the answer, not as context.The one that moves pipeline. Also the rarest, and the one most reports do not separate.

The distinction is not academic. Anthropic’s documentation on Claude’s citation feature describes citations as returning the exact passages that support each claim, which is a retrieval event with a specific page behind it. A mention has no page behind it at all. Microsoft draws the same line from the other side: the AI Performance report in Bing Webmaster Tools counts “the total number of citations that are displayed as sources,” so a Copilot answer that names you without linking you never enters that count.

If your report collapses all three into one “visibility” figure, you cannot tell an entity problem from a content problem, which is the first fork in entity problem or content problem. A brand mentioned everywhere and cited nowhere has a content problem. A brand cited but never named has the opposite.

Why does one run tell you almost nothing?

Because the same question asked twice does not reliably produce the same answer. This is documented behaviour, not a quirk of any one product. OpenAI’s own advanced usage guide states plainly that “Chat Completions are non-deterministic by default (which means model outputs may differ from request to request),” and even with a fixed seed the documentation offers only mostly deterministic output rather than a guarantee.

Google says something similar about its own surfaces. Its documentation on AI features notes that “AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary.”

What a single run can and cannot support
A single run supports Existence claims

“This engine can produce an answer that names us.” One observation is enough to prove something is possible.

A single run cannot support Frequency claims

“We appear 40% of the time.” A percentage derived from one observation is either 0% or 100%, and neither is a measurement.

The word “rate” is the tell. If a report shows a percentage, ask how many runs produced it. If the answer is one, the percentage is a label on a single event.

This is why run count belongs on the methodology page of any tool you are evaluating, and why so few publish it. A tool sampling once a day and a tool sampling five times will report different numbers from identical reality.

Does a branded prompt measure discovery?

It measures recall, which is a different thing and usually the thing you already know. Asking an engine “what does Rivarise do” tests whether your brand is resolved in the model. Asking “what tools track brand visibility in AI answers” tests whether you get discovered by somebody who has never heard of you.

The second question is the one with revenue attached, and it is the one a branded prompt cannot answer. If your prompt set is mostly branded, your report will look strong and your pipeline will not move. The unbranded set in prompts to test AI brand visibility is built for exactly this, and it deliberately contains no company names.

Branded prompt“Is Acme any good?”Tests recall and sentiment
Unbranded prompt“Best tool for X under $100”Tests discovery and shortlisting
Competitor prompt“Alternatives to Acme”Tests displacement, both directions
Only one of the three is a growth metric. Branded results are worth tracking as a regression check, not as evidence that AI search is working for you.

How do you read the report properly?

Six checks, in order. The first three establish whether the number means anything. The last three establish what it means.

Reading a citation report, in order
1
Find the run countPer prompt, per engine, per day. If it is not published, treat every percentage as a single observation until proven otherwise.
2
Count the branded promptsSplit the set. Report branded and unbranded separately, always. A blended figure hides which one moved.
3
Check the engine splitOne number averaged across engines that retrieve differently is not a measurement of anything. Keep them in separate rows.
4
Separate mention from citationIf the tool reports one figure, you cannot diagnose. Two brands with identical scores can need opposite fixes.
5
Read the answer text, not the scoreBeing named as the thing a competitor beats counts as a mention. The score cannot see tone; you can.
6
Date everythingA number without a date is not comparable to next month’s number, because the model underneath may not be the same model.
Steps one to three decide whether to keep reading. If a report fails any of them, the remaining numbers are decoration and no amount of interpretation rescues them.

Run the same six checks on any tool you are evaluating, including this one. Several vendors in the Peec AI alternatives roundup do not publish a run count anywhere, which is a finding about the tool rather than an omission on your part.

What should you do with a number you distrust?

Reproduce it by hand before you act on it. Take the prompt, run it five times in the engine itself, and read all five answers. If the tool says 40% and you see two out of five, the tool is measuring. If you see five out of five or zero out of five, ask how many runs went into that 40%.

A dashboard you cannot reproduce by hand is a dashboard you cannot defend in a meeting.

When the answer text disagrees with the score, trust the text. That is the whole method in AI giving wrong information about my product, and it applies here: the raw answer is evidence, the score is an interpretation of evidence, and only one of them is a primary source. If the two conflict on your own product, they conflict on your competitors too, and the same reading applies to AI recommends competitor instead of us.

If none of the six checks resolve it, the problem is upstream of the report. Start at why is my brand not showing up in AI search and work forward, because a measurement problem and a visibility problem produce the same flat line.