Google Ads · 8 min read · Oct 1, 2026

How to test AI Max without fooling yourself

Most AI Max tests measure the learning phase, a budget change or a campaign cannibalising itself — not AI Max. Here is how to run one that answers it.

Written by Philipp EndersFact-checked Oct 1, 2026Updated quarterly

Google reports that advertisers who switch on AI Max for Search campaigns typically see 14% more conversions or conversion value at a similar CPA. That is an average across accounts that chose to turn it on. It is not a forecast for your account, and the only way to find out what it does to yours is to test it.

The problem is that most AI Max tests answer a different question than the one that was asked. They measure the learning phase, or a budget change, or a campaign competing against itself. This is a method, not a verdict: how to set a test up so the result means something, and how to read it when it arrives.

Three ways an AI Max test lies to you

Before the setup, the three failures that account for nearly every inconclusive AI Max test we have seen:

  1. The learning phase sits inside the measurement window. Turning search term matching on restarts learning for the campaign. If your “after” period starts on the day you flipped the switch, the first week or two mostly describes an algorithm finding its feet. Campaigns usually look worse during it, which is why the internet is full of AI Max horror stories that are really learning-phase screenshots.
  2. Something else changed at the same time. A budget increase, a new target CPA, a seasonal peak, a landing page release. Any one of them makes attribution impossible, and the temptation to adjust is strongest exactly when the test dips.
  3. The A/B test is two campaigns fighting each other. Duplicating a campaign to run a control and a variant sounds rigorous and is the worst option of the three: both campaigns enter the same auctions for the same queries, so what you measure is the collision. Google’s own AI Max experiments avoid this by splitting traffic inside a single campaign.

First: is there anything to find?

This is the question to answer before you plan a test at all, and it gets skipped almost every time.

AI Max earns its keep by serving on queries you never booked. So the size of the upside is set by one thing: how much relevant demand exists outside your keyword list. That is not a matter of opinion — your own account will tell you in about twenty minutes.

  • How much of your conversion volume comes from brand terms? If it is most of it, the campaign is already harvesting demand that knows you exist. There is little left to discover and a lot to disturb.
  • How tight is your match type mix? A campaign running on exact match in a well-mapped category has a smaller unknown space than one running broad across an exploratory category.
  • Look at the last 90 days of search terms. How many genuinely new converting queries appeared that you had not booked? If the answer is “a handful a quarter”, that is your realistic discovery ceiling.
  • Can your site answer the adjacent intents? This one decides more than people expect. Expansion reaches service enquiries, dealer and location searches, support questions and neighbouring product categories — and sends all of them to the landing page you already use. If that page answers exactly one intent, the new traffic cannot convert, no matter how good the bidding is.

Two shapes follow from this. In campaigns with closed search volume — brand campaigns, tightly mapped exact-match sets — the discovery upside is small and the downside is real. In generic campaigns with open intent, there genuinely is something out there to find, and that is where a test is worth running.

If the honest answer is “there is nothing to discover here”, you have saved yourself six weeks. That is a legitimate outcome of this step.

Five rules for a clean test

Assuming there is something to find, five rules carry the whole method:

  1. Exclude the learning phase. Plan two weeks of learning that are not part of the evaluation. Write the dates down before you start, so the window cannot quietly move once the first numbers look bad.
  2. Freeze the levers. Budget, target CPA or ROAS, bidding strategy, keyword structure and landing pages stay untouched for the full test. If you have to change one, the test ends and a new one begins.
  3. Give it at least four weeks of measurement after the learning period. In a campaign with a handful of conversions a week, longer — a result built on fifteen conversions is a coin toss with extra steps.
  4. Test one thing. Leave automatic text customisation and final URL expansion off for the first test. With all three on you are measuring a package, and when the result is negative you will not know which part caused it.
  5. Write the success criteria down first. Primary metric, the threshold that counts as a win, and what you do in each outcome. Criteria written after the result are not criteria, they are a story.
AI MAX ONBaseline — 6 weeksLearning — 2 weeksEXCLUDEDTest — 4+ weeksMEASUREDBUDGET, TARGET CPA, KEYWORDS AND LANDING PAGES FROZEN THROUGHOUT
A test window that answers the question: six weeks of baseline, then roughly two weeks of learning that never enter the evaluation, then at least four weeks of measurement — with every other lever frozen from start to finish.

On the test mechanism itself: Google’s AI Max experiment splits traffic within one campaign, with a control arm that has AI Max off and a trial arm that has it on, which is cleaner than any duplicate-campaign construction. It is not available everywhere — campaigns using portfolio bidding, shared budgets, text customisation or another running experiment are excluded — and in those cases a frozen before/after comparison is the honest fallback. It is weaker evidence, and worth saying so in the report.

What to measure — and the metric nobody watches

Campaign-level KPIs are where this kind of change goes to hide. Conversions, cost per conversion and CTR for the whole campaign are averages, and an average absorbs a redistribution without showing it.

The mechanism is simple enough to state in one sentence: when a smart bidding strategy takes on extra traffic that converts at a different rate, it rebalances bids to keep hitting its target — which can mean bidding less on the single query that was carrying the campaign.

So alongside the usual campaign metrics, report your most valuable individual query on its own:

  • Clicks and CTR
  • Absolute top impression share
  • Impression share lost to rank
  • Conversions and cost per conversion
CHANGE IN CONVERSIONS, AFTER VS BEFORECampaign level (what the report showed)−9%The single most valuable query−29%BRAND CAMPAIGN TEST, SPRING 2026 · IDENTICAL BUDGET AND TARGET CPA
The same test, two levels of reporting. At campaign level the loss looks like a bad month. The query that carried the campaign lost three times as much — and that one line accounted for the entire drop.

We learned this the uncomfortable way. In a brand campaign test last spring, the campaign-level numbers moved by single digits while the one query that mattered lost 29% of its conversions — the entire loss of the campaign sat in a line nobody was reporting. We published the full test and the numbers on our sister site.

One reporting limit to plan around: a large share of AI Max traffic — in our test roughly 60% — appears in the search terms report only as “other search terms”. Your query-level analysis is therefore a sample, not a census. Treat it as directional, and do not build a case on the visible half alone.

How to read the result

Three questions, in this order:

  • Is the difference bigger than your normal noise? Look at weekly variation in the six weeks before the test. If conversions routinely swing 15% week to week, a 7% difference is not a finding.
  • Did your most valuable query survive? A campaign that gained 5% overall while its best query lost 20% has not improved. It has redistributed, and the redistribution will usually keep going.
  • Does it reverse? The cheapest confirmation available: switch AI Max off, wait out the new learning period, measure three unchanged weeks. If the metrics return, you have a mechanism rather than a coincidence.

Whatever the verdict, one thing is worth keeping: the queries AI Max discovered that converted. Book them as keywords in their own right. Even a failed test usually pays for itself here, and the keywords stay yours after you switch it off.

Be honest about the strength of the evidence, too. A before/after comparison with frozen levers is good practical evidence, not proof — the campaign still ran in a world where competitors, demand and Google’s own systems moved. Say “most plausible explanation” when that is what you have. Everyone reading your report already knows the difference.

When a rollout makes sense

Four conditions, all of which should hold before you widen AI Max beyond the test campaign:

  • The test campaign had genuine undiscovered demand, and the search terms report proves it did — new converting queries, not just more volume on the old ones.
  • Your most valuable queries held their position through the test.
  • Your negative keyword list covers the intents you cannot serve: service, support, careers, dealer and location searches, adjacent product categories.
  • Landing pages exist for the intents the expansion actually reaches. If they do not, build them first — that is a cheaper fix than paying for traffic that lands on the wrong page.

Roll out campaign by campaign, not account-wide in one evening. Each campaign has its own mix of brand and generic demand, and the mechanism that makes AI Max useful in one is the mechanism that makes it expensive in another.

Related reading: the Performance Max bidding change works on the same principle — a target that gets applied across a changed mix of traffic — and CRM conversion imports are what stop a bidding strategy optimising toward leads that never close. If you want the setup checked by someone who runs these tests weekly, that is what our Google Ads work is.

Common questions

PE
About the author
Philipp Enders

Philipp is the Founder and Director of pmax, a performance marketing and AI visibility agency in Calvià, Mallorca. He is also the co-founder of crunchjunkie, an AI visibility tracking platform for monitoring brand citation across ChatGPT, Perplexity, Claude and Gemini.

LinkedIn →