A minimum viable prompt set covers four intents: category discovery, shortlist building, head-to-head comparison, and objection checking. Most audits only test the first and overstate the result.
Twelve prompts, three per intent, is enough to produce a defensible baseline. The list is below and you can copy it.
The most common way an AI visibility audit goes wrong is not the tool, the run count or the session state. It is that every prompt tests the same moment in the buying process, so the result describes one narrow slice of behaviour and gets reported as though it described all of it.
Buyers ask different questions at different stages, and an engine answers them from different material. A set that covers only one stage is measuring one stage.
What are the four intents?
Category discovery
“What kind of tool solves this?”
- Tests
- Whether your category is understood and whether you are placed in it.
- Usually answered from
- Stable, general material. Often no lookup at all.
Shortlist building
“Which ones should I look at?”
- Tests
- Whether you are in the named set. The commercially decisive one.
- Usually answered from
- Comparative third-party material, not vendor sites.
Head-to-head
“How does A compare to B?”
- Tests
- Accuracy, and whether the comparison favours you or a competitor.
- Usually answered from
- Comparison pages and review platforms.
Objection checking
“Is it any good, and what goes wrong?”
- Tests
- Sentiment, and whether a known weakness has become the headline.
- Usually answered from
- Communities, reviews, and complaint threads.
What are the twelve prompts?
Swap the bracketed parts for your own category and buyer. Change nothing else, and do not put your company name in the first nine.
# 1. Category discovery, unbranded
Q1 What kind of software helps a [buyer] with [problem]?
Q2 What should a [size] [industry] team use to handle [problem]?
Q3 What is the difference between [your category] and [adjacent category]?
# 2. Shortlist building, unbranded. The commercially decisive group.
Q4 What are the best [category] tools for a [size] [industry] team?
Q5 Which [category] tools work well for [specific constraint]?
Q6 What are the leading [category] tools in 2026?
# 3. Head to head. Use competitors the engine itself named in Q4.
Q7 How does [you] compare to [competitor A]?
Q8 [Competitor A] vs [competitor B]: which is better for [buyer]?
Q9 What are the alternatives to [competitor A]?
# 4. Objection checking. Branded, and the only group that should be.
Q10 What are the downsides of [you]?
Q11 Is [you] worth the price?
Q12 What do users complain about with [you]?
Why does the branded and unbranded split matter?
Because they measure different things and mixing them inflates your result.
A prompt containing your name hands the engine the answer. It will describe you, because you asked it to. That tests accuracy, which is worth knowing, and it tests visibility not at all. Nine of the twelve above are unbranded for exactly that reason.
The engine has to choose you from a category. This is the number that describes your visibility.
Report these separately, always. Folding them into a visibility rate is the most common way an audit flatters itself.
Why does intent change which store answers?
Because engines do not look things up on every question, and the shape of the question decides whether they bother.
Anthropic documents this explicitly for its web search tool: “Claude determines when to search based on the prompt”, searching when a request depends on information that is “current, changing, or outside its training data” and answering directly from stable knowledge otherwise. The same pattern holds across assistants that can search.
That maps straight onto the four intents. Group one is general and stable, so it often gets answered from memory, and nothing you did this quarter can reach it. Groups two, three and four are specific, comparative and current, which is exactly the shape that triggers a lookup.
Useful for checking category placement, and mostly out of reach of anything you publish this quarter.
Specific, comparative and current. This is where your evidence work can actually change the answer.
On Google’s surfaces the mechanics differ but the conclusion does not. Google states there are “no additional requirements to appear in AI Overviews or AI Mode” beyond ordinary Search eligibility, so the qualifying work is your existing SEO, and inclusion in a named list is still decided from the retrieved set.
How do you run them properly?
Three constants, and breaking any one makes the result incomparable to itself later.
Twelve prompts, three engines, three runs is 108 answers, and it takes an afternoon. That is the whole cost of having a baseline rather than an anecdote.
What do you do with the results?
Compute a rate per intent group rather than one number for the set, because the four groups fail for different reasons and a blended figure hides which one is failing.
| If you are weak on | The likely cause | Where to look next |
|---|---|---|
| Category discovery | Entity or category placement | Consistent naming and description everywhere |
| Shortlist building | Thin third-party comparative evidence | Review platforms and independent comparisons |
| Head-to-head | Competitors have comparison coverage and you do not | The sources cited in those answers |
| Objection checking | An old complaint has become the canonical description | The specific source, traced and dated |
The most common shape is respectable on group one, poor on group two. That combination means the engine knows what you are and will not put you forward, which is an evidence problem rather than an identity one.
What should you not do with this list?
Three things, and all three destroy the comparison you are building.
Do not edit it once you start. Adding prompts because you are losing on the current ones changes the denominator, and the new set cannot be compared to the old one. Freeze it for a quarter minimum.
Do not use your own vocabulary. Ask somebody who bought recently how they searched, and use their words. A set written in your category page’s language tests how visible you are to people who already talk like you.
There is one free shortcut worth knowing here. Bing Webmaster Tools publishes an AI Performance report showing “grounding queries”, the phrases AI used when retrieving your content. Those are observed questions rather than invented ones, and they make an excellent starting point for groups one and two.
Do not run it once and conclude. A single pass gives you a baseline, not a trend. The value of these twelve prompts comes entirely from running the same twelve again next month.
Is twelve enough?
For a first baseline, yes. For a mature programme, no, and it is worth knowing which direction to grow in.
Twelve prompts across four intents is enough to tell you which stage is failing, which is the decision the first audit has to support. It is not enough to segment: one buyer persona, one geography, one product line. If you sell into three segments that ask genuinely different questions, twelve prompts will average across them and hide the segment that is failing.
Grow by adding whole intent groups for a new segment rather than by adding prompts to the existing groups. Four more prompts for a second persona keeps the structure intact and stays comparable. Sprinkling extras into group two because you are losing there changes the denominator and quietly resets your baseline.
One more caution about growth. Every prompt you add multiplies by engines and by runs, so twelve prompts on three engines at three runs is 108 answers, and twenty-four is 216. The recording is the work, not the asking.
The full method, including what to record and how to compute the rates, is in the no-budget audit. If you would rather spend five minutes than an afternoon first, the five-minute test uses three of these prompts to point you at the right problem before you commit to the full set.
