The question: what did your agency change?
Say you have paid a $4,000 monthly retainer for six months and your cost per acquisition has not budged. The monthly deck still celebrates impression share and click-through rate, while qualified pipeline sits where it did two quarters ago. That is not a messaging disagreement. It is an execution question, and the change log is where I would start.
When you ask what is happening, you may hear about market conditions, aggressive competitors, or the Google Ads learning phase. Any of those could matter. None tells you what the agency actually did. Before you replace the team, renew its contract, or argue through another slide deck, pull the account record.
This is a 30-day test, not a verdict based on a screenshot. You will compare activity across three periods, sort logged changes by commercial consequence, and check the results against what happened to your business. My expectation: when performance has stalled and the log contains little beyond automated or cosmetic activity, the agency will struggle to show how its management earns the retainer. The mechanism is simple. A report can describe work; a change history can show which account edits happened. It cannot show every hour of thinking, so keep that limit in the test.
Setup: fix the window before you inspect the edits
You need standard or administrative access to your Google Ads account. If the agency controls the account and will not give you direct access to its data, resolve that first. You cannot run an independent account audit from a PDF the agency chose to send you.
In the desktop interface, open Campaigns and select Change History. Google’s native interface lets you review a rolling two-year record. Use the date and user filters, inspect the details of individual changes, and export the periods you need. The User field helps distinguish a named operator from the Google Ads system or another tool. For a change whose context is unclear, use the “Go to…” control to inspect the affected campaign or ad group.

Set the test period to 30 completed days ending two days ago. Then pull two comparison windows: the 30 days immediately before it and the first 30 days of the agency’s engagement. Use the same account and the same classification rules for all three. Onboarding usually involves building or restructuring campaigns, so I expect more changes there than in a mature account. I do not expect an onboarding count to serve as a permanent monthly quota.
Before scoring anything, write down the business baseline for each window: spend, CPA, qualified pipeline or closed-won revenue if you have it, and any material change in offer, budget, landing page, or conversion tracking. You are testing management against outcomes, not assuming that every performance shift came from an account edit. Excluding the most recent two days also gives conversions some time to appear; use the same cutoff when comparing performance.
Keep a second boundary in view: the change log is not a complete diary of work. It will not capture every investigation, decision not to edit, or change made outside Google Ads. Conversion-tracking work may require a separate inspection, particularly when it happens in Google Tag Manager. Our 10-point Google Ads audit framework treats conversion integrity as its own checkpoint. For this test, score the account changes you can see, then ask for evidence of consequential work you cannot.
Controls: stop a busy log from passing the test
A raw row count is a bad score. One bulk action can generate many entries; automated changes can fill a month; and a renamed ad group does not become strategy because it has a timestamp. Export the log and give each relevant entry one of three labels:
- Automated activity: a change applied by the Google Ads system, an auto-apply recommendation, or another automation. Note who configured and monitors that automation, but do not count each automatic action as a fresh human decision.
- Cosmetic or clerical activity: naming, labeling, and other housekeeping that does not materially change targeting, auction behavior, the message, or the destination. Some housekeeping is necessary. It is not, by itself, evidence that someone is improving CPA.
- High-impact intervention: a deliberate change with a plausible route to better query quality, bidding, budget allocation, or creative performance. Record what changed and what problem it was intended to solve.

Do not infer automation from a crowded timestamp alone. Bulk edits can be deliberate. Open a sample of the entries, check the user or tool, and ask what decision drove them. Conversely, an agency does not get to call auto-applied recommendations hands-on optimization merely because it enabled them. Unmonitored rules can make consequential changes; the useful question is who reviewed those changes against your economics.
For ambiguous rows, apply one control consistently: what commercial variable could this edit change? If you cannot identify an effect on who sees the ad, what the auction costs, what the visitor reads, or where the visitor lands, put it in the clerical column. Separating vanity updates from structural work matters more than arguing over whether the account looks busy.
What to measure: four places a decision should leave a trace
For every high-impact entry, record its date, campaign, old and new state, apparent purpose, and any supporting performance evidence. Do not award a point because a change sounds technical. An edit counts as a meaningful intervention when you can connect it to a specific problem the agency was trying to address. I would look in four places.
- Queries and exclusions. Check negative keywords, shared lists, campaign-level exclusions, and relevant brand controls. Then read the search term report. If the account is spending on job-seeker queries, competitor names you do not sell, or informational searches that do not fit your goal, ask which exclusions addressed that spend. A month without new negatives is not automatically a failure; a month of visible waste without a response is. Broad match and Performance Max make this check especially important because they can explore adjacent intent.
- Bidding and value signals. Look for considered target CPA or target ROAS changes, value-rule work, and other adjustments tied to conversion volume, lead quality, or margins. Smart Bidding needs inputs and oversight, not random target thrashing. A bid change deserves credit when the agency can explain the signal behind it and what it planned to watch afterward. An unchanged target might also be the right call. Ask for the reason rather than awarding points for motion.
- Budget allocation. Identify shifts between campaigns and compare them with CPA, pipeline quality, and the account’s stated priorities. Money should not remain in a weak test simply because nobody revisited the spreadsheet. Equally, a frozen budget is not proof of neglect if the allocation still matches performance. The test is whether the agency can defend where the next dollar goes.
- Ads and assets. Look for new ad versions, replaced headlines or images, and sitelink updates. Compare them with the creative tests the agency says it ran. Four months without a creative edit is a reason to ask what the team learned about messaging, not a diagnosis of ad fatigue from the log alone.
Do not turn those four buckets into a quota. Account size, campaign type, and business changes affect which interventions make sense. The expectation I would test is narrower: if CPA has stalled and the agency says it is actively improving the account, you should be able to find decisions aimed at that problem or credible evidence of work outside the log. If all you find is clerical movement and automatic recommendations, the activity count is doing the selling instead of the work.
Read the result: price the interventions, then ask why
Count the high-impact interventions in your test window and compare that count with both earlier windows. If a $4,000 retainer bought four identifiable interventions, that is $1,000 per logged intervention. It is a prompt for a hard conversation, not an hourly rate: the log does not tell you how long analysis, implementation, or work outside the interface took. I would not pretend that four entries equal one hour of labor.

Now compare the categories with performance. If onboarding showed sustained work but the current month shows little beyond automated entries, while CPA and qualified pipeline have stalled, the agency has an explanation to provide. If the log is quiet and the account is meeting its CPA target while qualified pipeline grows, you may be looking at restraint rather than abandonment. Ask what the team monitored, what it chose not to change, and whether consequential conversion or offline work happened elsewhere.
I used to give clients a version of the mature account speech when I was juggling too many accounts. It has a real technical point: reckless changes can disrupt a bidding strategy. It also makes an excellent hiding place for a team that checks pacing once a week and calls the absence of fires a strategy. Markets move. Search terms drift. Landing pages change. “We left it alone” is a decision only if someone can show why leaving it alone served the business.
This is where the controls matter. An empty log alone cannot prove nobody worked. A busy log alone cannot prove anybody helped. The combination to challenge is an extended performance plateau, few defensible interventions, and no credible account of what the agency did instead.
Before the next invoice: make the agency explain the record
Book the review before the next billing cycle. Bring the exported log, your three comparison windows, and the business baseline. Keep the conversation concrete:
- Which changes were intended to address the CPA or pipeline plateau? Ask the agency to point to the entries and explain the mechanism.
- Why did the mix of interventions change after onboarding? Less activity may be appropriate; you want the reasoning, not a promise to create more rows next month.
- What work happened outside this log, and how can you verify it? Ask about conversion tracking, offline feedback, and investigations that led to a decision not to edit.
- What happens next? Get the proposed test, the metric it should affect, and when you will review it.
A capable media buyer should be able to walk you through the decisions, acknowledge a lull, or show you meaningful work the Google Ads history missed. If the answer is a tour of Optimization Score and a claim that auto-applied changes count as attentive human management, you have learned something useful about the retainer.
This is the operating-model question behind the audit. Ad auctions keep running between status calls. groas replaces periodic manual execution with specialized AI models that work continuously on bids, targeting, budgets, and creative, while a human strategist sets direction and guardrails. That is a better fit than paying for sporadic oversight dressed up as constant optimization. The change log gives you a way to test the distinction rather than take either pitch on faith.
Run the protocol every quarter, whether an agency, an in-house specialist, or autonomous software manages the account. If the record shows deliberate interventions tied to improving qualified pipeline, keep the operator and give the work room to run. If performance is stalled and the log offers only silence and automated filler, stop renewing the retainer on the strength of a deck. Ask for an answer before the next invoice; if it does not hold up, move the spend to an operating model that will act on what the account needs.




