There is no stable list of prompts that mention your brand. If a tool vendor tells you it found yours, it is selling you a 2012 keyword report with a new label.
I spent most of a decade building match type structures by hand and mining search term reports at 1am. The lesson I learned there is playing out again in ChatGPT, Gemini and AI Overviews: the list you track is a wish; what buyers actually ask is the truth. Ask the same buying question 100 times and you get close to 100 different answer lists. Track one phrasing and you learn little. Track the intent behind it and the sources each engine pulls, and you learn where to work.
The exact-match lesson I learned the slow way
I used to tell clients exact match meant control. I was wrong. In 2018, Google expanded close variants so exact match started matching same-meaning paraphrases and implied words, with [yosemite camping] matching campground language based on intent rather than the words you typed. Say you were spending $20k a month. Your keyword list said 200 tight exacts. Your search term report said 2,000 messy variants you never added. The list was the wish. The report showed what buyers actually typed and what you actually paid for.

That split changed how I worked. I stopped polishing the keyword tab and started living in the search terms tab: adding negatives, splitting out the intent that converted, feeding the winners back into the account structure. CPA moved when I managed the auction I was actually in, not the keyword list I wanted to be in. AI answers work the same way, only looser. The tracked prompt is the old keyword. Repeated runs across phrasings and engines are the new search term report. If you are not looking there, you are managing a wish.
One prompt produces a distribution, not a rank
Run the same prompt 20 times and you do not get an answer. You get a distribution. SparkToro and Gumshoe ran 2,961 prompts through hundreds of volunteers and found less than 1 in 100 repeat runs returned the same brand list, while identical ordering showed up less than 1 in 1,000. Ask 100 times and nearly every response is unique in list, order or count. That is not a bug you can prompt around. Even at temperature 0, the same prompt can return different answers because of sampling, server-side arithmetic and model updates. A single-run check of “does this prompt mention us?” is noise dressed as data.
A vendor screenshot that says you rank No. 3 for one prompt is the old exact-match fantasy. It pretends one phrasing equals one stable result. I used to sell that kind of certainty with keyword lists. The search term report kept embarrassing me, and repeated AI runs will embarrass this report the same way. If your visibility number changes every time you hit refresh, you are not measuring visibility. You are measuring randomness.
Buyers do not stick to your tracked phrasing, either. In that same SparkToro test, 142 people asked to write headphone-buying prompts produced almost no two alike, with semantic similarity of just 0.081. Yet the engines kept pulling from a stable consideration set, with Bose, Sony, Sennheiser and Apple appearing in 55 to 77% of 994 varied responses. That is what broad match did to my old keyword builds. A hundred wordings, one buying job underneath. Track a single wording and you miss it. Track the intent cluster across those wordings and you see which brands keep making the cut.
Then there is engine drift, which is measured, not theoretical. Ahrefs compared 730,000 AI Mode versus AI Overview pairs for the same queries and found only 13.7% citation overlap, or 16.3% among the top three, alongside 16% word overlap but 86% semantic similarity. One engine’s citation report tells you little about the other. Add ChatGPT and Gemini, which retrieve differently again, and a tool that tracks 50 prompts in one engine is reporting on one slot machine in a casino. I learned the placement lesson in PPC when a keyword that printed money on exact search bled on Performance Max. Different retrieval, different winners. Break visibility out by engine or you will mistake one view for the market.
Why a bigger prompt list does not fix the problem
Sales decks love a big prompt count. Five hundred tracked prompts sounds like coverage. It is a vanity number in a new costume. Backlinko’s 2026 tool test reached the same conclusion from the testing side and calls per-prompt position statistically meaningless, prescribing visibility percentage across a large prompt sample, run from the UI and not just APIs, broken out by engine. In Brand24’s analysis of 46,350 AI-visibility mentions from March to May 2026, the top complaint was unreliable tools and scores. Users described scores as lottery tickets: a dashboard says you are missing from 60% of answers but offers no fix or attribution. I recognize that report format. I used to receive it as a Google Ads Grader PDF. Pretty chart, no CPA change.
Position in the answer is not rank. Google rank was stable enough to track because ten blue links sat still long enough to measure. An AI answer has no comparable fixed slots. The model writes a paragraph, drops three citations, then rewrites it on the next run with different sources. Chasing whether you were brand No. 2 or No. 4 in one generation is like arguing about ad position while the auction reshuffles. Count how often you show up across repeated runs. Ignore where you sat in any single one.
More prompts can give you more observations. They cannot make one observation representative, and they cannot tell you what to change if the report never shows the sources behind the answer. That is the hole in the prompt-list pitch. It sells the inventory of questions as the asset when the useful information sits underneath them.
Measure the buying job and the sources behind it
Stop treating prompts as individual targets. Group phrasings of the same decision into one intent cluster: best headphone for remote calls under $300, emergency plumber near me that answers at night, CRM for a 20-person agency that hates manual entry. Then run 10 to 20 variants of each cluster per engine and count how often you appear across repeated answers. Treat each answer as one observation instead of pretending any one is the truth. Say you are spending $20k a month on search. You would never judge a keyword on one click. Judge an intent the same way.
Mention rate beats prompt rank for the same reason conversion rate beat ad position in my old accounts. I used to tell clients position 1.4 was the goal. I was wrong. Position told me where I sat. Conversion rate told me if I got paid. Mention rate across repeated runs tells you whether buyers meet you at all. Run it separately for ChatGPT, Gemini and AI Overviews, from the interfaces buyers use where you can, and watch the gap between engines before you spend a dollar fixing pages. For a repeatable way to set that panel up, use this method for measuring share of answer across repeated runs instead of screenshotting single prompts.
Look at retrieval before you rewrite the page
Your page does not lose every mention at write time. It can lose at retrieval time, when an engine pulls a different pile of sources to work from. Only 44.3% of Google top-10 pages appear in any AI answer at all, and for ChatGPT that number drops to 2.1% for the same queries. Read that twice. Ranking still helps, but your own ranking page is not the whole field. Reviews, comparisons, Reddit threads, videos and encyclopedia entries can shape the answer before your page gets a look. Backlinko’s test of 3.7 million citations found 91% of cited URLs appear on only one LLM. Fix your own page without checking what else gets pulled and you may be working on the wrong part of the problem.
The piles differ by engine. In the Ahrefs comparison, AI Mode leans encyclopedic while Overviews prefers video and Reddit. That is why one dashboard that only checks Overviews tells you nothing about what ChatGPT buyers see. I ran into the same wall with Shopping feeds years ago: same product, different placement logic, different winners. If your report lists prompts but not cited domains and URLs, it is hiding the part you can act on. Ask which of your pages got pulled, which third-party sources got pulled when you were missing, and how often you appeared for each intent cluster in each engine. That is the new search term report.
Put the visibility tool on trial in the demo
Sit in the demo and ask for the distribution, not the screenshot. Any tool can show you one answer where you appear. A useful one shows repeated runs of the same buying job across ChatGPT, Gemini and AI Overviews, how often you appeared, and which sources got pulled when you did not. If the rep keeps steering you back to total prompt count, you have your answer about what they sell. What the deck calls coverage, I call a phone book. Thick, impressive, and nobody finds customers in it.
Ask these questions before you pay for the report:
- How many runs do I see per prompt and per engine? Can I see the spread, not just a selected answer? One run is noise.
- Do you group prompts into intent clusters? I want mention rate per cluster, not a position beside each phrasing.
- Which domains and URLs got cited when I was missing? Break them out by engine. That list is the work order.
- Do runs come from the interfaces buyers use, or from an API only? I want to know how closely the report reflects what people actually see.
- What changes after the report? Show me which of my pages needs work and which third-party sources I need to earn a place in.
Walk away from any pitch that sells you the prompt list as the asset. The asset is the retrieval layer underneath it. That is what the earned search work groas runs for businesses is built to move: track buyer questions and citations by engine, then work on pages and third-party presence until mention rate moves.

The narrow case for a single-prompt check
This will not fit everyone. If almost all your AI traffic comes from people typing your exact brand name plus one question, single-prompt checks are fine. Navigational demand is stable enough to screenshot. If you buy visibility reports for compliance or franchise reporting rather than to change pages, a prompt list gives you the paper trail you need. Keep buying it. Everyone else who wants pipeline from buyers who have not chosen yet should stop paying for the list and start paying for the work underneath it.
Keep polishing the list and you will miss the buyer
I kept my old keyword builds for too long because they looked organized. The search term report kept telling me the organization was mine, not the buyer’s. Prompt lists are the same comfort object. Keep optimizing toward them and you will polish pages the engines rarely retrieve while a competitor owns the Reddit thread, the comparison site and the review set that actually get pulled. You will find out the way I found out with exact match: the report looked clean and CPA kept climbing. Measure the intent cluster across repeated runs, track which sources get retrieved, and fix what you can act on. The auction is where you get paid.

