September 24, 2026
min read

Four Pages Got 91.5% of Our Live-Answer Fetches. Here’s What They Had in Common

Young man with curly hair wearing a black shirt outdoors against green foliage background.


Alexander Perleman
, Head Of Product @ groas
Ex-Goldman Sachs and Stanford Computer Science

alex@groas.ai

LinkedIn
Cover image for: Four Pages Got 91.5% of Our Live-Answer Fetches. Here’s What They Had in Common

Four pages accounted for 91.5% of the live-answer fetches I identified in 30 days of groas.com edge logs. Our broad category pages, despite all the attention they get from crawlers, barely appeared in that narrower set.

 

That is the distinction I want to make here. A bot visiting a page is not the same thing as a page being pulled for an answer, and neither is automatic proof of a citation. I split the logs into raw bot hits, verified crawler fetches and a narrower group I call live-answer fetches. Then I checked what I could against answers in ChatGPT, Perplexity and Google AI Overviews. The pages that kept showing up were tightly scoped, dated analyses. The big, general pages did not.

 

Start with the denominator: a bot hit is not a citation

Most AI-crawler reporting starts with a count of requests carrying names such as GPTBot or PerplexityBot. I used to watch that number too. It answers a limited question: how often did something identifying itself as a crawler knock on the door? It does not tell me whether the requester was genuine, fetched the page HTML or contributed to an answer anyone saw.

 

For this analysis, I used three buckets:

 

Bucket What I counted What it cannot establish
Raw bot hits Requests with an AI crawler user-agent That the requester was the named crawler, or that it read a page
Verified fetches Requests that passed the available crawler-identity checks and fetched page HTML rather than only robots.txt or a sitemap That the page was used in a live answer
Live-answer fetches Verified fetches with the timing, referrer pattern or query-augmented signature I used to identify likely live retrieval That every individual request produced a visible citation

For verification, I checked published IP ranges for OpenAI where reverse DNS was unavailable, and used reverse DNS with forward confirmation where a vendor supported it. A user-agent string alone is a claim, not proof. Anyone can put User-Agent: GPTBot in a request. That makes raw-hit totals a poor basis for deciding what to publish.

 

In these logs, raw hits outnumbered the narrower live-answer set by roughly 40 to 1. Raw hits also outnumbered verified fetches by an order of magnitude; verified fetches still substantially outnumbered the live-answer set. Each filter changed the picture. The practical lesson is not that the narrowest bucket is perfect. It is that the broadest bucket answers the wrong question if the goal is citation.

 

Four dated pages took 401 of 438 likely live-answer fetches

Across the 30-day window, I identified 438 live-answer fetches. Four URLs took 401 of them. The remaining 37 were spread across the rest of the domain:

 

Page group Live-answer fetches Share of total
Dated, news-anchored analyses, including /youtube-ads-frequency-policy-2026 401 across 4 URLs 91.5%
Broad category and service pages: /paid-search, /earned-search and the homepage 3 combined 0.7%
Other blog posts and docs 34 across 19 URLs 7.8%

Source: my 30-day groas.com edge-log analysis. These are requests that met my live-answer-fetch criteria, not a platform-reported count of citations.

 

The split looks different if I stop at raw hits. Crawlers regularly walk navigation, sitemaps and internal links, so broad pages receive plenty of visits. At that level, a busy category page can look like an AI-visibility success. After verification and the live-answer filter, the three broad pages above contributed three fetches combined. The pages crawlers visit most are not necessarily the pages answer systems pull.

 

This is a concentration finding, not a claim that every dated post will perform this way. Four pages dominated this domain during this window. The useful question is what those pages offered that the broader ones did not.

 

The four winners answered a question from that week

Each of the four leading pages targeted a tight question and carried a date in its title: a change to YouTube ad frequency in a specific month, for example, or a policy update for 2026. Each ran roughly 1,200 to 2,000 words and included a table, a dated claim and at least one number that could be quoted directly. None tried to cover an entire head term. They answered something a buyer or media buyer might need to know that week.

 

That combination matters because the likely retrieval task is specific. If the question is what changed with YouTube ad frequency in June 2026, a page naming the change, the date and the relevant figure gives the answer system a self-contained passage. A general guide may mention frequency somewhere deep in the copy, but the reader still has to find which part applies now. The same problem faces a retriever trying to assemble an answer.

 

I used to tell clients to build the big guide first. I was wrong about what that would do for this kind of visibility. The guide can still serve readers and link to narrower work. In these logs, though, the dated memo was the page that kept getting pulled. For a changing question, write the dated answer before expanding the general guide.

 

The broad pages looked healthy until I checked answers

Our /paid-search and /earned-search pages are longer and better linked than the four winning posts, and crawlers visit them constantly. Together with the homepage, they produced only three live-answer fetches in the measured window. That is the number I would miss if I watched a bot-traffic dashboard and stopped there.

 

I also ran 22 prompts across ChatGPT, Perplexity and AI Overviews, asking questions those broad pages purport to address. The prompts included how to run Google Ads without an agency and how YouTube frequency changes affect budgets. None of the resulting answers cited a category page. Eighteen cited one of the four dated posts, usually with a number taken from a table.

 

Manual check Result
Prompts run across the three answer surfaces 22
Answers citing a broad category page 0
Answers citing one of the four dated posts 18

Source: my manual prompt check. It is a check on this set of questions, not a citation rate for every possible query or user.

 

The manual check gives the log pattern some context. It does not turn every request in the table into a proven citation. It does show the same divide from the reader-facing side: the category pages described what groas does; the dated posts stated what had changed. For these prompts, answers favored the second kind.

 

Cartoon of an AI robot librarian passing thick general textbooks to pull a thin dated memo from a shelf

Specificity helps a passage match; freshness gives it a useful date

A live answer to a time-sensitive question needs information tied to that time. A title naming the month, a table carrying the relevant numbers and a sentence stating the change make a page easier to use for that purpose. A broad category page can be longer, more prominent in site navigation and more frequently crawled while still being a weaker match for the question at hand.

 

Consider the illustrative case of a business spending $20k a month on YouTube when frequency costs move 15% in a week. A memo from that week, with the change stated plainly, would address the immediate question better than a 5,000-word guide from last year. I cannot prove a retriever made that exact comparison from an edge log. I can say the observed 401-to-3 split is consistent with answer systems pulling pages that state a recent, scoped change rather than pages that describe a category.

 

Word count did not distinguish the winners here. The four leading posts averaged about 1,600 words; the ignored category pages averaged more than 2,800. The longest post on the site recorded no live-answer fetches in this set. That does not make short writing a ranking tactic. It means adding another 1,000 words would not, by itself, supply the missing dated answer or quotable figure.

 

Freshness also showed up in the fetch curve. For a dated post, the rate peaked in its first 10 to 14 days, then fell by roughly half each following week unless I updated it with a new number or date. I treat these posts as perishable inventory for that reason. When the facts move, I refresh the table and timestamp. When they do not, I do not add a cosmetic date just to make old information look new.

 

That is the publishing tradeoff. A narrow memo can answer a live question quickly, but it also needs maintenance if that question keeps changing. A general page has a different job. It can explain the service and help a human decide what to do after an answer sends them there. I would not judge it by how many time-sensitive answer fetches it earns.

 

Where a 30-day log cannot take us

This is one domain, one niche and one 30-day window. groas.com covers paid and organic search, so the questions that pulled these pages leaned toward ads, benchmarks and policy updates. A site with different topics could have a different distribution. I would not publish the 91.5% figure as a universal target.

 

The measurement has a harder boundary, too. OpenAI runs distinct user-agents with different jobs, and referral information can be stripped before it reaches our logs. Identity checks help separate genuine crawler traffic from a spoofed user-agent; request patterns help narrow that traffic to likely live retrieval. Neither lets me inspect every answer produced after every request. That is why I kept the manual prompt check separate and why I call the narrowest bucket live-answer fetches, not confirmed citations.

 

Line chart showing dated-post fetches peaking and then declining over four weeks

The chart describes the dated posts in this window: an early peak, then a decline unless the underlying information was refreshed. It does not tell me that every topic expires on the same schedule. What it changes in my workflow is simpler: I plan for a follow-up when the facts change instead of treating publication as the end of the job.

 

For finding a question to answer, I use the one-page, one-dated-question approach in our guide to getting cited in Google AI Overviews. I then use groas Earned Search to check visibility across ChatGPT, Perplexity and AI Overviews. The log analysis tells me where to look closely; it does not excuse me from checking what readers actually see.

 

Publish the memo when the question has changed

If I were planning the next batch of pages from these numbers, I would not commission another long category guide just to increase AI-crawler hits. I would look for dated questions we can answer clearly, publish a scoped page while the answer is useful, and watch whether verified fetches and visible citations follow.

 

My working brief is short:

 

  1. Choose a question with a real date attached: a policy update, frequency change, benchmark release or pricing shift in the niche.
  2. State the change early. Put the date, entity and relevant figure together in a sentence a reader can understand without the rest of the page.
  3. Use a table when it makes the before-and-after easier to see. Add the context needed to interpret it; do not pad the page to hit a word count.
  4. Refresh the table and timestamp when the facts move. Watch the fetch curve rather than assuming last month’s interest will persist.

I would skip this approach for a stable, low-news service with no changing numbers. Publishing a stream of dated posts into that void is busywork with nicer formatting. A broader how-to library may serve those readers better.

 

For a site with changing questions to answer, run the test on your own pages. Separate raw hits from verified fetches, inspect the likely live-answer set, and check actual answers. If two or three dated pages take 80% or more of the pulls, that is a publishing decision, not a reason to celebrate a crawler dashboard. Keep making the narrow pages that answer what changed. Let the category pages do their job when a human arrives.