Run three prompts in a fresh session with memory off: an unbranded category question, a branded description question, and a direct comparison. The gap between answers two and three is your real problem.
It takes five minutes and it costs nothing. What it gives you is not a score, it is a direction: whether to work on being known, being trusted, or being reachable at all.
Most people check their AI visibility by typing their own company name into ChatGPT and reading the answer. That is the one test guaranteed to tell you nothing, because you handed the engine the answer inside the question. The three-prompt version below takes the same five minutes and is actually diagnostic.
Why does the session state matter so much?
Two things quietly corrupt this test before you have typed anything.
Memory, chat history and past conversations all shape what comes back. If you have discussed your company before, the engine may name you for reasons no buyer will ever benefit from.
Temporary chat, memory off, logged out, or a private window. You want the answer a stranger gets, because a stranger is who you are trying to reach.
Set that up first. Whichever engine you use, start a temporary or incognito session, turn memory off if the product has it, and do not paste anything from a previous chat.
The second contaminant is subtler and catches careful people. If you run the three prompts, get a disappointing answer, and then reword the first prompt to see whether that helps, everything after the reword is happening in a session that has already seen your original question. Start over in a new session instead. It costs thirty seconds and it is the difference between a test and a conversation.
What are the three prompts?
Ask them in this order, in the same session, without correcting the engine between them. The order matters because each answer sets up how you read the next one.
1. Unbranded category question. Your name appears nowhere.
What are the best [category] tools for a [size] [industry] team?
2. Branded description question.
What is [your company] and who is it for?
3. Direct comparison against a name it gave you in answer 1.
How does [your company] compare to [competitor it named]?
How do you read the three answers?
Do not score them. Compare them. Four patterns cover almost everything, and each points somewhere different.
Absent, vague, vague
Not in answer 1, thin in answer 2, and answer 3 hedges or invents. The engine does not have a confident idea of what you are.
Entity problem
Absent, good, good
Not named in answer 1, but described well in 2 and compared sensibly in 3. It knows you. It just does not bring you up unprompted.
Evidence problem
Absent, good, wrong
Described accurately alone, then compared using stale pricing, a removed feature, or the wrong segment. A third-party source is outranking your own.
Source problem
Present in answer 1
You were named unprompted in the category question. Rare on a first try, and worth re-running several times before believing it.
Measure, do not celebrate
Which engine should you run it on?
At least two, from different retrieval families, because being absent on one says nothing about the others.
Google’s surfaces run on the Google index with the same eligibility rules as ordinary Search, and the company states there are “no additional requirements to appear in AI Overviews or AI Mode”. ChatGPT and Perplexity run their own crawlers and their own indexes, documented respectively by OpenAI and Perplexity. Those are genuinely different systems and they disagree regularly.
If you only have five minutes, use ChatGPT and one Google surface. If you have ten, add Perplexity, which tends to cite more sources per answer and therefore shows you more about where the category’s evidence actually lives.
Common, and not conclusive. Different corpora, different retrieval, different day.
Three separate systems reaching the same conclusion about you is a signal worth acting on.
What does this test not tell you?
It is a direction-finder, not a measurement, and treating it as one is how a five-minute check turns into a bad quarter.
Ask again in a minute and the named brands can differ. One run is a sample of one, whichever way it lands.
ChatGPT, Perplexity and Google’s surfaces run different retrieval. Being absent on one says nothing about the others.
You cannot tell whether today’s result is better or worse than last month’s, because there is no last month.
If the result points at a real problem, the next step is the full version: the same logic applied across ten prompts, three engines and three runs each. That is the four-question diagnostic, and it is still an afternoon rather than a project.
Two definitions are worth having straight before you interpret any of this, because the difference between them decides which of the four readings you are looking at: being mentioned and being cited are not the same event, and a test that conflates them will point you at the wrong fix.
What should you do with the answer text?
Keep it. All three answers, in full, pasted somewhere dated. It takes ten seconds and it is the only part of this exercise that gains value over time.
In three months you will want to know whether the description in answer 2 has changed, and a screenshot of a verdict cannot tell you that. The raw text can, and it can also answer questions you have not thought of yet, which is the entire reason our own pipeline stores full answers rather than a tally of hits.
What if the engine gets you wrong?
If answer 2 describes you inaccurately, resist the instinct to correct it in the chat. The correction lasts one session and teaches you nothing about why it was wrong.
Instead ask it where it got that. Most engines will name a source, and the source is the finding. A stale pricing figure usually traces to a directory listing or a review profile nobody has updated in two years, and that page is the fix, not your own site. Your site is already right, which is precisely why the engine is not using it.
This is the most valuable thing the five-minute test surfaces, and the one people most often miss, because they are scanning for their own name rather than reading what was said around it.
How often is it worth repeating?
Monthly is plenty for a check this small, and there is a specific reason not to do it more often: you will start reacting to variance. The same three prompts run a week apart can differ for no reason connected to anything you did, and a team that runs the test on Monday mornings will eventually build a narrative out of noise.
Run it monthly, keep every answer, and treat a change as interesting only when it persists across two consecutive checks. If you need to know sooner than that, you have outgrown the five-minute test and want the sampled version instead.
That is the whole test. Three prompts, one clean session, five minutes, and a direction rather than a number.
