Copilot’s retrieval runs on Microsoft’s search infrastructure, which means a Bing indexing problem is a Copilot visibility problem, and Bing Webmaster Tools is the only free diagnostic that reports directly on an AI engine’s citations of you.
That report exists, it is free, and almost nobody in this category has noticed. It shows citation counts, the pages cited, and the queries used to retrieve them.
Every other engine in this field is a black box you can only probe from outside by asking it questions. Copilot is the exception, and the reason is that Microsoft already ran a webmaster programme before any of this and simply extended it.
What does Microsoft actually give you?
A report called AI Performance, in Bing Webmaster Tools, which shows “how publisher content appears across Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations”.
How often your content is referenced as a source in AI-generated answers. A count you did not have to sample for.
Unique pages shown as sources per day across supported AI surfaces.
The key phrases AI used when retrieving your content. This is the one worth the whole exercise.
Citation counts for specific URLs, so you can see which pages are carrying you.
Why is the grounding-queries report such a big deal?
Because the hardest problem in building a prompt set is not measuring it. It is knowing which questions to put in it.
Everywhere else, your prompt set is a hypothesis: somebody wrote thirty questions they believe buyers ask, and every number afterwards is conditional on that guess. Grounding queries invert that. They are a record of the phrases that actually retrieved your pages, which is the closest thing to ground truth this discipline offers, and it is free.
Use the grounding queries to build your prompt set for the other engines. The questions that retrieve you on one surface are a reasonable starting hypothesis for the rest.
How does this change what you do for Copilot?
It collapses AI visibility work on this engine into search visibility work you may already be doing.
If retrieval runs on Microsoft’s index, then being indexed there is the entry requirement, and Bing Webmaster Tools is where you verify it. That is a different situation from the engines that run their own crawlers: OpenAI and Perplexity both document dedicated crawlers with their own configuration, and neither gives you a report on how you performed.
What are the limits of using it as your only signal?
Two, and they are large enough that this cannot be your whole programme.
It covers Microsoft’s surfaces and nothing else. Your citation counts here say nothing about ChatGPT, Perplexity, Claude, or Google’s AI surfaces, which run separate retrieval over separate corpora. Treating a healthy AI Performance report as evidence of general AI visibility is exactly the generalisation error this field is full of.
And it reports citations rather than mentions. Being named in a Copilot answer without a link to you does not appear, which is the same blind spot every link-based measure has. The gap between being mentioned and being cited is diagnostic, and this report only sees one side of it.
First-party, free, no sampling required, and with the retrieving queries attached.
Separate corpora, separate retrieval, and a citation-shaped view of a problem that is partly about being named at all.
Why does one engine give data away when none of the others do?
Because Microsoft was already in this business, and the answer is more mundane than strategic.
Bing Webmaster Tools has existed for years as a search product, with an established relationship with site owners, verification flows and a reporting habit. Extending it to cover AI citations was a feature addition to something that already worked. The assistant-first companies have no equivalent relationship with publishers to extend, so building one would mean starting from nothing.
That is worth knowing because it predicts where the next free data will appear. Engines built on top of a search index have a webmaster surface available to them. Engines that started as assistants do not, and probing them from outside by asking questions repeatedly is likely to remain the only method for some time.
Grounded in a search index
Copilot, and Google’s AI surfaces. A webmaster relationship already exists, so first-party reporting is at least possible.
Data plausible
Assistant-first
ChatGPT, Claude, Perplexity. No publisher programme to extend, and sampling from outside is the only method available.
Sampling only
Does Google offer the same thing?
Not in this form, which is the comparison that makes the Microsoft report unusual.
Google publishes eligibility rules for its AI surfaces, stating there are “no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary” and that a page must be “indexed and eligible to be shown in Google Search with a snippet”. That tells you what qualifies. It does not tell you how often you were actually cited, on which pages, or for which queries.
So the position is asymmetric: Google tells you the rules, Microsoft shows you the results. Neither of the assistant-first engines does either.
What should you take from this?
One engine gives you first-party data for free, and the correct response is to use it without over-reading it.
Verify the site, read the grounding queries, feed those queries into the prompt set you use everywhere else, and keep measuring the other engines by sampling because there is no alternative. The Copilot data is a gift and a sample of one engine at the same time.
How should this change your reporting?
Give Copilot its own row, sourced from Bing Webmaster Tools rather than from your tracking tool, and label the number differently from the others, because it is a different kind of measurement.
Everywhere else you are reporting a sampled rate: how often you appeared across a fixed prompt set at a stated run count. Here you are reporting a census: how many times Microsoft observed your content cited. Those are not the same unit and putting them in one chart implies a comparison that does not hold.
The practical convention that works is to report Copilot’s figure as a count with its own axis, and the sampled engines as rates with theirs. It looks less tidy and it stops somebody asking why one engine’s number is an order of magnitude different from the rest.
How the engines divide into retrieval families is in the engine guide, the two Google surfaces that behave differently from each other are in AI Overviews vs AI Mode, and if the grounding queries suggest your prompt set is wrong, the no-budget audit covers rebuilding it properly.
