boxoo
Menu
Guide · Audits & optimization

Google Ads experiments: how to A/B test campaigns, ads, and Performance Max

Google Ads experiments run a change on part of a campaign’s traffic beside the original, so you can see whether it worked before you roll it out. Here is how to set one up, read it, and when a simple before/after is enough.

Last checked against Google’s documentation on 5 October 2026.

What are Google Ads experiments?

Google Ads experiments are built-in A/B tests. Instead of changing a campaign and comparing this month with last month, you run the change on a share of the campaign’s traffic and budget at the same time as the original, and Google reports the difference with a measure of statistical confidence. Both arms see the same season, the same competitors and the same news, so the comparison is far fairer than a before/after.

You find them under Experiments in the Campaigns menu. Google’s Experiments page overview lists the types available. This guide covers which one to use, how to set up a custom experiment, how long to run it, how to read the result, and when an experiment is more effort than the question deserves.

Types of experiments in Google Ads

TypeWhat it tests
Custom experimentsA changed copy of a Search, Display, Demand Gen or Video campaign: bid strategy, match types, landing pages, audiences, ad groups
Ad variationsA text change across responsive search ads in one or many Search campaigns
AI Max experimentsA one-click test of AI Max features in a Search campaign
Performance Max experimentsUplift, upgrade and optimization tests for Performance Max
Demand Gen and Video experimentsWhich images or videos perform better
App asset experimentsText and images inside App campaigns

Google’s custom experiments page names Search, Display, Demand Gen and Video as the campaign types custom experiments support, and the setup guide says App and Shopping campaigns can’t use them. If you are weighing AI Max itself, our AI Max for Search guide covers what it changes.

How to set up a custom experiment

Google’s setup steps:

  1. Open Experiments in the Campaigns menu, click the plus button, and choose Custom.
  2. Name the experiment and select the original campaign.
  3. Add a suffix for the trial campaign’s name, such as “_test”.
  4. Choose up to two goals as success metrics.
  5. Make the one change you are testing in the trial campaign.
  6. Set the Experiment split, the split type, and the experiment dates, then save.

Two rules from the setup guide catch people out. Changes you make to the original campaign after the experiment starts are not copied to the experiment, so editing the original mid-test changes what you are comparing. And the original campaign has to stay active; if it is paused, the experiment won’t serve. You can schedule up to five experiments per campaign, but only one runs at a time.

A custom experiment, start to finish
  1. One questionOne change, one primary metric, written down before you start.
  2. Set upTrial copy, 50% split, cookie-based, dates long enough for the volume.
  3. Leave itNo edits to either arm while it runs.
  4. DecideApply, convert to a new campaign, or end it.

Traffic split: cookie-based vs search-based

The Experiment split is the share of the original campaign’s traffic and budget the trial gets. Google recommends 50% for the best comparison. A smaller split protects the original if the change is risky, but the trial collects data more slowly and needs longer to reach a result.

Split typeHow it assigns trafficUse it when
Cookie-based (Google’s recommendation)Each user is randomly assigned to one arm and only sees that versionThe change affects the whole visit or repeat visits, such as a landing page or bid strategy
Search-basedEach search is randomly assigned, so one user can see both versionsYou need data faster and seeing both versions doesn’t bias the result

For Display campaigns Google always uses a cookie split. The custom experiments page also notes that traffic is split by auction: each ad slot is a new auction, so if the trial has different bids or assets, it can win more or fewer auctions than the original even with a 50% split. Uneven impressions between arms aren’t by themselves a sign of a broken test.

A/B testing ads with ad variations

Ad variations are the quickest Google Ads A/B test for ad text, and they are still available. Google’s ad variation setup page puts them under Experiments: click the plus button, then Assets, Assets provided by you, Ad variations. You choose all campaigns or specific ones, select responsive search ads, and describe the change, for example a find-and-replace of “Book now” with “Call now”. You set start and end dates and the experiment split.

The advantage is scale: one variation can change the same phrase across hundreds of ads without editing them one by one. If the variation wins, applying it creates new ads with the change. For tests of whole new ad concepts in one ad group, it is usually simpler to add a second responsive search ad and compare; our CTR guide covers what to compare.

Performance Max experiments

Performance Max doesn’t use custom experiments. Google’s Performance Max experiments page lists three kinds:

Uplift tests answer a question a before/after can’t: whether Performance Max finds new conversions or takes credit for ones your Search campaigns would have won anyway. If you suspect the second, read Performance Max cannibalizing brand search first.

How long to run a Google Ads experiment

Google’s monitoring guide recommends letting an experiment run for 2 to 3 weeks to gather data. If the Results column still shows In progress, Undecided or Unavailable, the Experiments overview suggests at least 4 to 6 weeks, or a larger budget. When the test changes a bid strategy, the custom experiments page advises allowing 7 to 14 days for the trial to stabilize while the strategy learns, so don’t judge that period.

An experiment timeline (bid strategy test)
  1. Week 0
    Write the hypothesis

    The change, the metric that decides it, and the result that would make you apply it.

  2. Weeks 1–2
    Learning

    Check both arms serve and spend. Don’t judge results yet.

  3. Weeks 3–4
    First read

    Look for the asterisk on your primary metric. Conversions from recent clicks may still be arriving.

  4. Weeks 5–6
    Decide

    Apply, end, or extend once if the result is still undecided and the trend is consistent.

Volume matters more than days. An experiment that tests cost per conversion needs enough conversions in each arm to separate a real difference from noise; a campaign with a handful of conversions a week may never reach one. If customers usually take days to convert after the click, end the experiment and wait for the lag before reading the final numbers.

Reading the results and significance

Open the experiment from the Experiments page. The scorecard shows Clicks, CTR, Cost, Impressions and All conversions by default, and you can add the metrics you chose as goals. For each metric, Google shows the difference between the trial and the original as a percentage and a range.

  • Blue asterisk: the difference is statistically significant.
  • Confidence interval: 80% by default. You can choose another level such as 95%, which widens the range and makes significance harder to reach.
  • The range: a result such as [+8%, +12%] means the trial’s true difference is likely somewhere between those values.
Experiment vs original: difference by metricIllustrative example
  • Not significant
  • Significant (asterisk)
Conversions+14%
Cost+9%
Clicks+3%
CTR+1%
View as table
ItemDifferenceGroup
Conversions+14%Significant (asterisk)
Cost+9%Significant (asterisk)
Clicks+3%Not significant
CTR+1%Not significant
For example: conversions rose more than cost, so cost per conversion improved. Clicks and CTR moved within noise. The decision rests on the metric you chose before the test, not on whichever one has an asterisk.

Two reading habits help. Decide on the primary metric before you start, so you don’t pick the winner from whichever metric happened to move. And look at cost alongside conversions: a bid strategy that gets 10% more conversions for 30% more spend may not be a win for your margins, whether you sell products, generate leads or sell subscriptions.

Applying or ending the experiment

Google’s apply guide gives two options when you click Apply:

  • Update your original campaign: the experiment ends and its changes are copied into the original.
  • Convert to a new campaign: the trial becomes a campaign of its own and the original is paused.

Google keeps the performance data of both either way. Updating the original keeps its name and history in one place, which is usually tidier. If the experiment lost, end it; the original never changed. Note that applying a new bid strategy to the original is itself a change, so expect a learning period afterwards; see our Smart Bidding guide.

Common experiment mistakes

  • Testing several changes at once. A new bid strategy plus new ads plus a new page tells you the bundle won, not why.
  • Editing the original mid-test. Changes to the original aren’t copied to the trial, so the arms drift apart.
  • Stopping at the first asterisk. Early results swing. Run the planned duration unless one arm is clearly broken.
  • Calling it during learning. Bid strategy tests need the stabilization period before results mean much.
  • Too little volume. A small split on a low-volume campaign may never reach significance. Use 50%, or test on a bigger campaign.
  • Ignoring conversion lag. Reading results the day the test ends undercounts the trial’s latest conversions.
  • Judging on the wrong metric. CTR and clicks are inputs. Decide on cost per conversion, conversion value, or return on ad spend.

When a before/after with undo is enough

Not every change needs an experiment. Experiments are worth the setup when the change is large, hard to reverse, or likely to shift learning: a new bid strategy, broad match across a campaign, a new landing page, or adding Performance Max. For small, reversible changes, such as adding negative keywords, pausing a keyword that has spent with no conversions, adding sitelinks, or a modest budget increase on a campaign limited by budget, a careful before/after is usually enough: change one thing, note the date and the old value, compare equal periods, and reverse it if it hurt.

That is how Boxoo handles routine changes. The performance and structure agent reads change history and shows the numbers before and after each change, and the budgets and bidding agent checks whether a bid strategy has enough conversions to learn. Each finding arrives as a case with the evidence and a prepared fix. Nothing changes until you press Apply, and most applied changes keep an undo, so a small change can be tried and rolled back without a formal test. The big ones still deserve an experiment; see where they fit in our optimization routine.

Run a free Google Ads audit to find which changes in your account are worth testing first.

Questions

How do I run an A/B test in Google Ads?

Open Experiments in the Campaigns menu, click the plus button and choose the experiment type. For a campaign change such as a new bid strategy or landing page, use a custom experiment: pick the original campaign, make the change in the trial copy, set the split (Google recommends 50%), choose cookie-based or search-based, and set the dates. For ad text, use ad variations.

Which campaigns can use custom experiments?

Google lists Search, Display, Demand Gen and Video campaigns for custom experiments. App and Shopping campaigns can’t use them. Performance Max has its own experiment types, and Demand Gen and Video also have dedicated experiment types for creative.

How long should a Google Ads experiment run?

Google recommends letting an experiment run for 2 to 3 weeks to gather data, and if results still show as In progress, Undecided or Unavailable, 4 to 6 weeks. Low-volume campaigns and long conversion delays need longer. Don’t stop on a date; stop when both arms have enough conversions to compare.

What does the blue asterisk mean in Google Ads experiments?

It marks a metric whose difference between the experiment and the original is statistically significant. The confidence interval defaults to 80%, and you can choose a stricter level such as 95%, which widens the range shown for each metric.

What happens to the original campaign when I apply an experiment?

You choose. “Update your original campaign” merges the experiment’s changes into the original. “Convert to a new campaign” keeps the experiment as a new campaign and pauses the original. Either way Google keeps the performance data of both.

Keep reading

Know which changes are worth testing.

Boxoo’s agents review your Google Ads account every day and file each finding as a case with the evidence, so you test the changes that matter. Nothing changes until you approve a fix, and applied changes keep an undo.