AI share of voice is the percentage of answers in a defined prompt set where your brand is named, divided across every named competitor. It is only meaningful when the prompt set and engine mix are fixed.
Change either one and the number moves without anything about your visibility changing. That is the whole difficulty with the metric, and the reason most reported figures are not comparable to each other.
Share of voice was a broadcast idea before it was a search idea: your spend as a fraction of category spend. Ported to AI answers it keeps the shape and loses the guarantee, because nobody controls the inventory and the inventory is not the same twice.
What exactly is share of voice in AI search?
It is a ratio with a strict denominator. The numerator is the count of answers naming you. The denominator is the count of brand-naming slots across the whole set, yours and everyone else’s.
Which questions you asked. Fixed, written down, and unbranded if the number is meant to describe discovery.
Changing one prompt changes the metric.Which engines, in what proportion. An average across engines that retrieve differently is not one measurement.
Report per engine, then aggregate.How many times each prompt ran. One run gives 0% or 100%, and neither is a share.
Published, or the figure is unverifiable.How do you calculate it?
Five steps. The arithmetic is trivial; the definitions are where it goes wrong.
Worked through: 30 prompts, 5 runs, 3 engines is 450 answers. If those answers contain 1,300 brand-naming events and 260 of them are you, your share of voice is 20%. Report it as “20% across 450 answers, 30 unbranded prompts, 5 runs per prompt per engine, three engines, week of [date]” or do not report it.
Why does the denominator decide the number?
Because you choose it, and every choice moves the result. Add two small competitors nobody shortlists and your share falls without anything changing in the answers. Drop them and it rises.
| Denominator choice | What the number then means | Honest use |
|---|---|---|
| All brands named, no list | Your share of everything the engines mention, including irrelevant ones. | Category breadth. Poor for tracking, because the set changes week to week. |
| A fixed competitor list | Your share among the companies you actually compete with. | The default. Requires publishing the list alongside the number. |
| Top three named per answer | Your share of shortlist positions rather than of mentions. | Closest to pipeline. Hardest to compute, because it needs the answer text. |
None of the three is wrong. Mixing them between periods is, and it is the most common way a share of voice chart shows a trend that did not happen.
The sensitivity is easy to underestimate. Take the same 450 answers and the same 260 mentions of you. Against a denominator of 1,300 naming events you are at 20%. Add two adjacent vendors that the engines name often and the denominator grows to 1,600, so the identical performance now reports as 16%. Nothing about your visibility moved. A four-point drop appeared because the list changed.
This is why the competitor list belongs in version control alongside the prompt set. If you add a competitor mid-quarter, recompute the earlier periods against the new list before drawing the line, or the chart shows a decline that is an artefact of bookkeeping.
How many runs before the percentage means anything?
More than one, and the reason is documented rather than theoretical. OpenAI’s advanced usage guide states that “Chat Completions are non-deterministic by default (which means model outputs may differ from request to request),” and notes that even a fixed seed gives only mostly deterministic output.
So each run is a draw, and a share of voice figure is an estimate of a proportion from a sample. The uncertainty on that estimate is a solved problem: NIST’s engineering statistics handbook gives the interval for a proportion, and the practical consequence is that small samples produce intervals wide enough to swallow the movement you are trying to detect.
The practical rule: if the change between two periods is smaller than the change you see between two runs in the same period, you have not detected anything. Run the current period twice before believing a movement, which is the same discipline described in how to interpret an AI visibility report.
Should you blend engines into one figure?
No. Google says the surfaces differ from each other, let alone from other vendors: its documentation on AI features notes that “AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary.”
Engines also differ in what a naming event even is. A live-search engine and a training-recall engine can produce the same brand list for completely different reasons, which is the split covered in the AI engines your buyers actually use. Averaging across that split hides which mechanism you are winning or losing on, and the fixes are different for each.
Keep a row per engine. If leadership wants one number, give them the weighted figure with the weights stated, and keep the per-engine table underneath it.
What does share of voice not tell you?
Whether you were recommended. Naming events count appearances, and an appearance as the incumbent a challenger beats counts identically to an appearance as the answer. Anthropic’s citations documentation describes citations as returning the exact passages supporting a claim, which is a stricter and more useful event than a mention, and share of voice does not distinguish the two.
A rising share of voice with flat pipeline usually means you are being named more often as the comparison, not as the choice.
It also says nothing about why. A share that falls on one engine while holding on others is a retrieval story, and where you look next depends on the engine: how ChatGPT chooses sources, how does Perplexity rank sources, and how Gemini grounds answers each have a different failure mode behind the same drop in one number.
Track it as a trend line with its methodology attached, and read the answers when it moves. The metric is a pointer, and the answer text is the evidence.
