The invoice looks harmless at three domains. At six, the same white-label AEO offer starts eating its own margin; by nine, the agency is reporting bot hits it cannot verify. The failure is not content. It is the order in which pricing, delivery and proof break.

This is a composite, not an incident at one agency. It is the version of the rollout teams run into when they bolt AEO onto SEO retainers without changing how they price, ship or measure the work. The turning points matter because each failure had an earlier decision that could have prevented it.

The agency begins with three retainers at around $1,500 a month each, assuming delivery costs about a third of that. That is familiar reseller math, with little slack. Clients start asking to appear in AI answers, so the agency adds a tracker and a content add-on. The first three domains look fine. Drafts get written, the work fits into spare hours, and the report looks clean.

Then the fourth client signs. Nobody changes the offer.

Failure 1, domains 4 to 6: the prompt bill arrives first

The first sign is an invoice, not a ranking drop. At three domains, the AEO tracker looks cheap. At six, the agency sees what its flat retainer concealed: the meter runs per domain and per prompt. Each new client adds a base cost, while broader prompt coverage calls for another pack.

One breakdown puts Semrush’s AI Visibility Toolkit at $99 a month per domain for 25 daily prompts, with $99 per extra domain and $60 per extra 50 prompts. Its bundle tiers still cap sites and prompts. An agency comparison calls that $99 add-on the agency killer at 15, 30 or 50 clients. The arithmetic starts to hurt well before 50.

Moving to a cheaper tracker does not automatically fix the model. Otterly ranges from Lite at $29 a month for 15 prompts to Premium at $489 for 400 prompts. If the agency priced ten clients as though each needed only a handful of questions tracked, meaningful coverage forces an upgrade it never allowed for in the retainer.

The wrong call was treating tracking as a small, fixed overhead. I used to tell clients tooling was a rounding error. I was wrong about AEO. Here, the questions an agency chooses to monitor are part of the bill, and choosing them takes time. Someone still has to enter prompts and clean up AI-suggested lists that do not mirror real customer behaviour. A white-label pricing guide puts hidden labour at around 90 minutes a month per retainer, or about $88 per account at a $60 loaded rate.

So the agency keeps its $1,500 price, pays for more coverage, and quietly absorbs the prompt work. The earlier decision was to cost delivery per domain per 50 prompts before quoting client five. Waiting until client six signs turns a pricing question into a margin problem.

Failure 2, domains 6 to 8: correct audits sit in developer queues

Next, the work slows down. The audits identify thin titles, missing schema, weak internal links and a robots file blocking the wrong crawler. The agency sends tickets to client developers, as it did for the first three domains.

At six to eight domains, that habit becomes the bottleneck. Every client has a different developer, backlog and approval path. The strategist spends the month following up on changes instead of checking what shipped. The audit can be perfectly sound and still have no effect if its fixes never reach a live page.

The wrong call happened at signup: the agency sold technical execution without securing a way to execute. Its proposal promised progress; its operating model promised tickets.

Edge SEO offers a way around some of those queues. With tools such as Cloudflare Workers, Fastly, Akamai or Lambda@Edge, a team can implement changes to titles, metas, redirects, structured data and robots.txt after the CMS renders HTML but before users or Googlebot see it. The approach removes platform and developer-queue bottlenecks on restrictive CMSs, letting the team ship changes without waiting for a CMS release cycle.

That does not mean access is a detail to sort out later. The agency needs permission, ownership and a clear path for making those changes before it sells the work. Otherwise, it has merely moved the request from one queue to another.

The earlier decision was to agree on technical access at signup. If every AEO fix requires a client developer to act, the agency does not have an execution process. It has a recommendation queue.

Failure 3, domains 8 to 10: approved content still is not live

By domain eight, drafts are sitting in Google Docs with green checkmarks beside them. Approved. On strategy. Correctly briefed. Not published.

Here is where the agency makes the wrong diagnosis. It asks for more content capacity because the calendar is slipping. But writing another draft adds to the pile. Every client runs a different CMS, with different logins, permissions and editors that can break formatting on paste. Someone must format each page, add internal links and schema references, publish it, then check that it actually renders.

At three clients, that handoff can fit into an afternoon. At ten, it becomes a job the agency never scoped. Publishing slips a week, then a month. The client sees no new live URL and has little reason to care that the draft was approved on time.

The tracker stack can widen the gap between knowing what to write and getting it onto the site. One agency comparison found that Profound has no multi-account management: one workspace per account means five clients can mean five logins. The comparison also notes no content-creation layer to carry a prompt gap through to a published page. Lower-tier trackers can leave agencies without the content workflow or white-label reporting they need. Another dashboard does not clear a CMS queue.

The earlier decision was to make publishing part of delivery, with an owner and an agreed route from brief to live URL. Content is not shipped when a client approves the document. It is shipped when the page is live and someone has verified it.

Failure 4: the monthly report starts calling bot noise visibility

By month three, the agency needs something to show for the work. The monthly PDF offers a rising AI visibility score built from server-log hits with AI-bot user-agents. It looks precise. It is not proof that an AI answer cited the brand, or even that a named bot made those requests.

A user-agent is a claim, not an ID. Cloudflare reported Perplexity crawlers masquerading as Chrome on macOS across rotating IPs after its declared bot was blocked; it de-listed Perplexity as a verified bot. The point for this report is simpler: a string in a request header does not establish who sent the request.

Anyone can send GPTBot in a header. A practitioner teardown of edge bot measurement explains the distinction: Cloudflare’s verifiedBotCategory provides the check, while the user-agent is merely what the requester says it is. UA-only counting can file path-probing under a vendor name. Its example shows 36 hits claiming to be Claude-SearchBot hammering /api/uploads/..%2F..%2F.env, all unverified. Put those hits on an upward-sloping chart and the chart still does not mean what the label says.

The wrong call was choosing an easy-to-count proxy before deciding what the report needed to prove. Verified fetches can show verified crawler activity. Live-answer checks can show whether the brand appears in answers. Neither should be silently renamed revenue, but they are more honest than raw bot-user-agent counts. This postmortem on six months of tracking with zero citations shows why the gap between tracking activity and appearing in answers matters.

The earlier decision was to define the report before building it: verified fetches and live-answer citations, clearly separated. If someone can fake a metric with a curl header, it does not belong in a client report.

Failure 5: client nine asks what the agency changed

Client nine reads the report and asks the useful question: what did you change last month, and which change moved the number?

The account manager opens the tracker, the rank tool and the shared document. There are scores, prompts and tickets. What is missing is a timestamped record tying a shipped change to its reason: these titles changed on Tuesday, these pages went live on Thursday, and this is what the team expected each change to affect. I know the irritation from PPC search-term mining at 1am. If the work is not logged where someone can find it, you end up re-arguing work you already did.

The report cannot answer the question by getting louder. A chart does not establish that a page was published, and a ticket marked complete does not explain why the team chose that page. Under pressure, the agency has treated reporting as presentation rather than a record of delivery.

Building its own measurement layer is possible, but it needs the same cost discipline the agency skipped in Failure 1. A practitioner teardown notes that a bot-measurement Worker placed on every request counts every request against Cloudflare’s 100k-a-day free limit. A page view loading 20 files can use 21 requests, putting a site with 10k page views a day beyond that free limit. The cited paid quota is $5 a month for 10M requests, then $0.30 per extra million. Across client domains, homegrown logging is another cost to price, not a free escape from the tracker bill.

The earlier decision was to make an action log part of the offer, with each change, timestamp and reason available under the agency’s brand. Logging every action with its reasoning lets the team answer client nine without reconstructing the month from three tabs and somebody’s memory.

Client report with an unverified visibility graph beside a small verified-fetch count

The pre-mortem: five calls to make while there are still three clients

The agency could have caught this while the first three retainers were comfortable. Not by buying a larger dashboard. By deciding what the offer costs, how work reaches a live site and what the report can honestly prove.

  1. Price prompt coverage. Cost delivery per domain per 50 prompts against the retainer. Check domain add-ons, prompt caps and the time required to choose useful questions before selling the same flat price again.
  2. Secure technical access. Agree on a way to ship title, schema and redirect changes without waiting on each client’s developer. That can include edge execution through a connection that works across platforms and CMSs, with the access settled at signup rather than after the first audit.
  3. Own the route to publication. Put brief-to-live work in scope. Name who handles CMS access, formatting, links and the final check of the rendered page. An approved document is not a published page.
  4. Define proof before drawing charts. Keep verified fetches distinct from live-answer citations in the branded weekly report. Do not relabel unverified bot traffic as visibility because it gives the PDF a nicer line.
  5. Keep the change record. Log what shipped, when and why, so the next client question has an answer that does not depend on the account manager remembering a Tuesday ticket.

These are the same five buyer questions agencies ask before buying another dashboard: white-label setup, prompt caps, content execution, retainer margins and proof of work. Asked before client five, they shape the offer. Asked after client nine, they explain why the offer is in trouble.

Ten client storefronts sharing one jammed delivery tunnel

The order is what makes this postmortem useful. Margin goes first because prompt and domain meters compound behind a flat retainer. Delivery goes next because access and publishing were treated as somebody else’s jobs. Reporting fails last because the agency needs evidence of progress and has not built a trustworthy way to collect it. Content was never the constraint. The business model around content was.

That is why I point agencies toward groas for agencies when they ask what can scale across 10-plus domains. It brings paid, SEO and AI visibility work under the agency’s brand, with edge-level execution and a logged record, rather than handing the team another per-prompt dashboard and leaving it to do the work.

My rule now: do not sell the next domain until you can price, ship and prove its work without borrowing time or margin from the last one.