Blog
7 min read Diagnose Your AI Visibility

12 Prompts Every B2B SaaS Should Run Against Every Engine This Week

On this page
  1. What are the four intents?
  2. What are the twelve prompts?
  3. Why does the branded and unbranded split matter?
  4. Why does intent change which store answers?
  5. How do you run them properly?
  6. What do you do with the results?
  7. What should you not do with this list?
  8. Is twelve enough?
The short version

A minimum viable prompt set covers four intents: category discovery, shortlist building, head-to-head comparison, and objection checking. Most audits only test the first and overstate the result.

Twelve prompts, three per intent, is enough to produce a defensible baseline. The list is below and you can copy it.

The most common way an AI visibility audit goes wrong is not the tool, the run count or the session state. It is that every prompt tests the same moment in the buying process, so the result describes one narrow slice of behaviour and gets reported as though it described all of it.

Buyers ask different questions at different stages, and an engine answers them from different material. A set that covers only one stage is measuring one stage.

What are the four intents?

Four stages, four different retrievals
1

Category discovery

“What kind of tool solves this?”

Tests
Whether your category is understood and whether you are placed in it.
Usually answered from
Stable, general material. Often no lookup at all.
2

Shortlist building

“Which ones should I look at?”

Tests
Whether you are in the named set. The commercially decisive one.
Usually answered from
Comparative third-party material, not vendor sites.
3

Head-to-head

“How does A compare to B?”

Tests
Accuracy, and whether the comparison favours you or a competitor.
Usually answered from
Comparison pages and review platforms.
4

Objection checking

“Is it any good, and what goes wrong?”

Tests
Sentiment, and whether a known weakness has become the headline.
Usually answered from
Communities, reviews, and complaint threads.
Stage two is the one that decides revenue and stage four is the one that surprises people. Most audits run three variants of stage one and stop.

What are the twelve prompts?

Swap the bracketed parts for your own category and buyer. Change nothing else, and do not put your company name in the first nine.

The prompt setThree per intent
# 1. Category discovery, unbranded
Q1  What kind of software helps a [buyer] with [problem]?
Q2  What should a [size] [industry] team use to handle [problem]?
Q3  What is the difference between [your category] and [adjacent category]?

# 2. Shortlist building, unbranded. The commercially decisive group.
Q4  What are the best [category] tools for a [size] [industry] team?
Q5  Which [category] tools work well for [specific constraint]?
Q6  What are the leading [category] tools in 2026?

# 3. Head to head. Use competitors the engine itself named in Q4.
Q7  How does [you] compare to [competitor A]?
Q8  [Competitor A] vs [competitor B]: which is better for [buyer]?
Q9  What are the alternatives to [competitor A]?

# 4. Objection checking. Branded, and the only group that should be.
Q10 What are the downsides of [you]?
Q11 Is [you] worth the price?
Q12 What do users complain about with [you]?
Q9 is the sneakiest one on the list. Asking for alternatives to a competitor is how you find out whether you are in their consideration set, which is a different question from whether you are in your own.

Why does the branded and unbranded split matter?

Because they measure different things and mixing them inflates your result.

A prompt containing your name hands the engine the answer. It will describe you, because you asked it to. That tests accuracy, which is worth knowing, and it tests visibility not at all. Nine of the twelve above are unbranded for exactly that reason.

Unbranded, Q1 to Q9 Tests whether you get surfaced

The engine has to choose you from a category. This is the number that describes your visibility.

Branded, Q10 to Q12 Tests what is said about you

Report these separately, always. Folding them into a visibility rate is the most common way an audit flatters itself.

Two numbers, never one. A blended figure across both is not comparable to anything, including your own figure next quarter if the ratio shifts.

Why does intent change which store answers?

Because engines do not look things up on every question, and the shape of the question decides whether they bother.

Anthropic documents this explicitly for its web search tool: “Claude determines when to search based on the prompt”, searching when a request depends on information that is “current, changing, or outside its training data” and answering directly from stable knowledge otherwise. The same pattern holds across assistants that can search.

That maps straight onto the four intents. Group one is general and stable, so it often gets answered from memory, and nothing you did this quarter can reach it. Groups two, three and four are specific, comparative and current, which is exactly the shape that triggers a lookup.

Group 1 Often answered from stored knowledge

Useful for checking category placement, and mostly out of reach of anything you publish this quarter.

Groups 2, 3, 4 Usually trigger a lookup

Specific, comparative and current. This is where your evidence work can actually change the answer.

An audit made only of group-one prompts will conclude nothing you do matters. It will be measuring the store you cannot influence.

On Google’s surfaces the mechanics differ but the conclusion does not. Google states there are “no additional requirements to appear in AI Overviews or AI Mode” beyond ordinary Search eligibility, so the qualifying work is your existing SEO, and inclusion in a named list is still decided from the retrieved set.

How do you run them properly?

Three constants, and breaking any one makes the result incomparable to itself later.

Engines
Three minimumFrom different retrieval families, because they disagree for real reasons
Runs each
ThreeAnswers vary between identical runs. One run is a coin flip
Session
Fresh every timeTemporary or incognito, memory off, nothing pasted from before
Recorded
Full answer textPlus brands named, their order, and every cited URL

Twelve prompts, three engines, three runs is 108 answers, and it takes an afternoon. That is the whole cost of having a baseline rather than an anecdote.

What do you do with the results?

Compute a rate per intent group rather than one number for the set, because the four groups fail for different reasons and a blended figure hides which one is failing.

If you are weak onThe likely causeWhere to look next
Category discoveryEntity or category placementConsistent naming and description everywhere
Shortlist buildingThin third-party comparative evidenceReview platforms and independent comparisons
Head-to-headCompetitors have comparison coverage and you do notThe sources cited in those answers
Objection checkingAn old complaint has become the canonical descriptionThe specific source, traced and dated

The most common shape is respectable on group one, poor on group two. That combination means the engine knows what you are and will not put you forward, which is an evidence problem rather than an identity one.

What should you not do with this list?

Three things, and all three destroy the comparison you are building.

Do not edit it once you start. Adding prompts because you are losing on the current ones changes the denominator, and the new set cannot be compared to the old one. Freeze it for a quarter minimum.

Do not use your own vocabulary. Ask somebody who bought recently how they searched, and use their words. A set written in your category page’s language tests how visible you are to people who already talk like you.

There is one free shortcut worth knowing here. Bing Webmaster Tools publishes an AI Performance report showing “grounding queries”, the phrases AI used when retrieving your content. Those are observed questions rather than invented ones, and they make an excellent starting point for groups one and two.

Do not run it once and conclude. A single pass gives you a baseline, not a trend. The value of these twelve prompts comes entirely from running the same twelve again next month.

Is twelve enough?

For a first baseline, yes. For a mature programme, no, and it is worth knowing which direction to grow in.

Twelve prompts across four intents is enough to tell you which stage is failing, which is the decision the first audit has to support. It is not enough to segment: one buyer persona, one geography, one product line. If you sell into three segments that ask genuinely different questions, twelve prompts will average across them and hide the segment that is failing.

Grow by adding whole intent groups for a new segment rather than by adding prompts to the existing groups. Four more prompts for a second persona keeps the structure intact and stays comparable. Sprinkling extras into group two because you are losing there changes the denominator and quietly resets your baseline.

One more caution about growth. Every prompt you add multiplies by engines and by runs, so twelve prompts on three engines at three runs is 108 answers, and twenty-four is 216. The recording is the work, not the asking.

The full method, including what to record and how to compute the rates, is in the no-budget audit. If you would rather spend five minutes than an afternoon first, the five-minute test uses three of these prompts to point you at the right problem before you commit to the full set.