Only 18.8% of the AI bot hits on our homepage were verified crawlers. Over 30 days on groas.com, the page logged 11,793 requests claiming to be AI bots; 2,219 passed verification. A chart of the first number would have looked busy. It would not have told us much about AI visibility.
The more useful number sat elsewhere: three dated, specific posts received 524 live-answer fetches between them, while three category pages with more than 3,000 claimed hits each received none. That does not prove every fetch produced a citation. It does show why I no longer treat crawl volume as the prize. Answer retrieval concentrated on a few pages built to answer specific questions.
The homepage number that fooled me
Our edge logs counted requests whose user agents said they were AI bots. That is a claim, not an identity: anyone can send User-Agent: GPTBot with curl from a laptop. Scrapers can wear the same costume, particularly where a site whitelists known AI crawlers.
We checked the claimed traffic against publisher information: GPTBot against OpenAI’s gptbot.json, OAI-SearchBot against its searchbot.json, and Claude traffic against Anthropic’s combined bots.json feed. These are our site logs, not an industry sample; the percentages below describe this path in this 30-day window.
| Path | Claimed AI hits | Verified crawler hits | Verified share |
|---|---|---|---|
Homepage / | 11,793 | 2,219 | 18.8% |
I used to tell clients to watch crawl volume as a proxy for AI visibility. I was wrong. On our homepage, roughly four out of five hits in that chart failed the verification step. Before asking whether a number went up, ask what the number counts.
Separate the bots before you compare pages
I now keep three buckets rather than one:
- Claimed hits: requests carrying a bot user agent, whether or not the requester is that bot.
- Verified hits: requests whose network identity passes the relevant check. For OpenAI, we match published IP lists; for Anthropic, its published feed; for PerplexityBot, forward-confirmed reverse DNS and published bot files.
- Live-answer fetches: verified requests classified as retrieval for an answer, rather than a steady training crawl. This is a measure of fetching, not a count of citations displayed to users.
The jobs differ. OpenAI distinguishes GPTBot, OAI-SearchBot, and ChatGPT-User: training, search and citations, and on-demand page fetches are not interchangeable. Allowing GPTBot does nothing for OAI-SearchBot or PerplexityBot. Blocking GPTBot does not by itself remove you from ChatGPT answers; blocking OAI-SearchBot can. Nor does reverse DNS settle OpenAI identity when OpenAI publishes no stable rDNS suffix. Use the check that fits the bot.
Our robots.txt shows why even the third bucket needs interpretation. It drew 1,156 fetches classified as live Claude traffic in the same window. That is a retrieval check, not evidence that anyone cited a robots file. And an Anthropic IP verifies the organization, not which of ClaudeBot, Claude-User, or Claude-SearchBot made the request; the user agent and request pattern still matter when classifying its job.
There are limits outside the log, too. Cloudflare notes that edge blocking can keep requests out of origin logs. A separate 28-day comparison reported roughly 2,237 ClaudeBot crawls and 217 GPTBot crawls per referral. That study measures crawls against referrals, not our answer fetches against citations, so I would not paste its ratios onto our site. Its useful warning is narrower: a large crawl count need not produce a large visible return. Count the bot’s job before you assign value to its visit.
Three posts took the answer-fetch traffic
In our logs, 524 live-answer fetches went to three URLs. The figures below come from the same groas.com 30-day window as the homepage count. The final column describes what each page offers a reader or fetcher; it is my explanation of the pattern, not a reason recorded by a bot.
| Page | Live-answer fetches | Material available to retrieve |
|---|---|---|
| YouTube ads 2026 guide | 290 | Dated format prices, frequency rules, and policy cutoffs |
| AI Max setup and data guide | 144 | Step order, data requirements, and settings in sequence |
| Google Ads updates 2026 | 90 | Dated changes, effective dates, and whom they affect |
These are not three versions of an ultimate guide to everything. Each gives a fetcher something bounded: a date, a price or limit, a requirement, or a step in an order. The page does less interpretive work for the system retrieving it. That is my reading of the concentration, not a claim that a log can reveal why a model chose a URL.
Outside work points in a similar direction, with different measures and limits. One citation study reports AI-cited content as about 25.7% fresher than standard search results and reports higher AI visibility for content with statistics, sources, and quotes, alongside lower visibility for keyword stuffing. Its tracker of 1,429 AI answers also found ChatGPT returning to pages with specific, checkable facts. Those are outside observations, not a controlled explanation of our 524 fetches. Our logs support the practical hypothesis: make the answer easy to locate and check. They do not let me claim that each request became a displayed citation.
Three busy category pages got none
The contrast is blunt. Each of our three busiest category pages cleared 3,000 claimed AI hits in the same window. Their live-answer fetch counts were zero, zero, and zero. Older explainer posts also received verified bot visits without showing up in the answer-fetch pattern we could identify. Traffic reached those pages; the retrieval activity we were looking for concentrated elsewhere.
That is not the same as saying category pages are useless, or that every broad page loses every citation. It says their apparent bot popularity did not predict this particular outcome. A separate summary of ChatGPT Search behavior reports that around 85% of retrieved pages in its analysis were never cited. Retrieval itself is already a narrower measure than crawling, and even retrieval is not a citation receipt.
I would still choose a narrow page with a quotable answer over a broad page that makes the reader hunt through several topics. Our three posts give that choice a concrete basis on this site. The category-page chart gives me no reason to buy more of what it measures.
What I would change on a business site
First, report verified answer fetches beside crawl totals, not underneath them as a footnote. I want two ratios each month: verified hits divided by claimed hits, and live-answer fetches divided by verified hits. The homepage’s first ratio was 18.8%; the three category pages recorded zero live-answer fetches. Neither figure alone describes the whole site’s visibility. Together, they make a vendor’s rising bot-traffic chart much harder to mistake for progress. If someone cannot show how they checked bot identity, I would not call the chart verified AI traffic.
Second, build fewer, sharper pages around buyer questions. Say you spend $20k a month and one broad services page tries to cover twelve questions. I would not assume that splitting it will reproduce our results. I would ask whether any one question has a clear answer on that page. Our most-fetched posts put dated, usable details where a reader can find them. My working rule is one buyer question, one URL, a quotable answer in the first 150 words, and the supporting detail underneath. Date material that depends on the year. Name the numbers you can substantiate and cite their source on the page. Do not add a year merely as decoration; make it tell the reader which version of an answer they are getting.
Third, check fetchability before rewriting copy. A polished answer cannot help a retrieval bot that is blocked at the edge or receives blank HTML without JavaScript. Check the bot rules, the response it actually receives, and whether your robots file permits the retrieval crawlers you mean to reach. Before You Write for ChatGPT, Check Whether AI Can Read Your Site lays out that audit before you spend money on new content. There is no point polishing a page the intended fetcher cannot read.
Run the comparison on your own logs
Question: Which pages receive live-answer fetches, and which merely collect claimed bot hits? This is the test I would run prospectively with edge logs or CDN bot analytics. Hold the time window and classification rules constant across paths. Do not treat a change in either as a content result.
- Collect 30 days of requests by path and user agent. Export path, user agent, IP, and timestamp. Keep
robots.txtin the set: its retrieval checks are a useful reminder that a fetch need not mean content demand. - Split claimed traffic from verified traffic. Match OpenAI requests against its published
gptbot.json,searchbot.json, andchatgpt-user.jsonlists; Anthropic requests againstbots.json; and Perplexity against forward-confirmed reverse DNS and its bot files. Leave requests that fail your check in the claimed bucket. - Classify answer retrieval separately. Look for OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, and Perplexity-User, alongside fetch bursts tied to a fresh answer request rather than a steady crawl rhythm. For a shared publisher IP feed, do not infer the exact bot from the IP alone. Keep uncertain requests identifiable rather than silently promoting them to confirmed answer activity.
- Rank your top 10 paths three ways: claimed hits, verified hits, and live-answer fetches. For each path, calculate verified share and live-answer share. Keep the raw counts beside the ratios; a percentage without its denominator is another attractive chart that can waste your afternoon.
My expectation is that the lists will diverge. Claimed traffic may favor a homepage or category pages; answer retrieval may favor dated posts with prices, thresholds, or steps. The mechanism I would test is straightforward: a page that supplies a checkable answer gives a fetcher less to reconstruct than a page that gestures at several answers. That is a hypothesis to compare with your own paths, not a forecast that every site will have our distribution.
If the lists match, inspect the pages before rewriting them; your broad pages may already contain the answers people seek. If answer fetches read zero, check edge rules and rendered output before blaming the copy. Either result changes the next task. Do not commission more content to increase crawling until you know which pages retrieval bots can reach and use.
One site and one month set the boundary
This is one domain, one 30-day window, and a site that publishes dated Google Ads and YouTube guides. A local services site with twelve pages and no dated posts has no reason to expect our 524-fetch pattern. A news site may have a different rhythm again. Treat the homepage’s 18.8% verified share and the three posts’ 524 fetches as observations about groas.com, not targets for another site.
Skip the verified-share test if you cannot obtain edge logs or CDN analytics with IPs. User-agent strings alone cannot run that step. If your firewall blocks retrieval at the edge, a zero in your logs may reflect access policy rather than a failed page. And throughout this analysis, a fetch is evidence of retrieval, not proof of a published citation. That boundary matters most when someone tries to turn a promising log count into a claim about revenue.

The decision the numbers support is narrow: stop buying crawl volume as a goal. On our site, the busiest claimed traffic did not identify the pages getting answer fetches. Three dated, specific posts did that work. If you want to see which buyer questions trigger mentions of your business and where the gaps sit, earned search is built to track that across engines. You can start with the logs: pull the three lists, inspect the pages already getting fetched, and decide what answer each could make clearer. Leave the busy category pages out of the rewrite queue unless the evidence gives you a reason to put them back.
Keep the raw export for the next month. Bot lists and hosts change, and one 30-day slice cannot show whether an edit worked. I keep the three lists frozen by date, then compare the same paths after a fetchability fix or rewrite. If the retrieval pattern moves, there is something to investigate. If only the claimed-hit chart moves, I do not call it a win.




