Last checked against Google’s documentation on 5 October 2026.
What are Google Ads experiments?
Google Ads experiments are built-in A/B tests. Instead of changing a campaign and comparing this month with last month, you run the change on a share of the campaign’s traffic and budget at the same time as the original, and Google reports the difference with a measure of statistical confidence. Both arms see the same season, the same competitors and the same news, so the comparison is far fairer than a before/after.
You find them under Experiments in the Campaigns menu. Google’s Experiments page overview lists the types available. This guide covers which one to use, how to set up a custom experiment, how long to run it, how to read the result, and when an experiment is more effort than the question deserves.
Types of experiments in Google Ads
| Type | What it tests |
|---|---|
| Custom experiments | A changed copy of a Search, Display, Demand Gen or Video campaign: bid strategy, match types, landing pages, audiences, ad groups |
| Ad variations | A text change across responsive search ads in one or many Search campaigns |
| AI Max experiments | A one-click test of AI Max features in a Search campaign |
| Performance Max experiments | Uplift, upgrade and optimization tests for Performance Max |
| Demand Gen and Video experiments | Which images or videos perform better |
| App asset experiments | Text and images inside App campaigns |
Google’s custom experiments page names Search, Display, Demand Gen and Video as the campaign types custom experiments support, and the setup guide says App and Shopping campaigns can’t use them. If you are weighing AI Max itself, our AI Max for Search guide covers what it changes.
How to set up a custom experiment
Google’s setup steps:
- Open Experiments in the Campaigns menu, click the plus button, and choose Custom.
- Name the experiment and select the original campaign.
- Add a suffix for the trial campaign’s name, such as “_test”.
- Choose up to two goals as success metrics.
- Make the one change you are testing in the trial campaign.
- Set the Experiment split, the split type, and the experiment dates, then save.
Two rules from the setup guide catch people out. Changes you make to the original campaign after the experiment starts are not copied to the experiment, so editing the original mid-test changes what you are comparing. And the original campaign has to stay active; if it is paused, the experiment won’t serve. You can schedule up to five experiments per campaign, but only one runs at a time.
- One questionOne change, one primary metric, written down before you start.
- Set upTrial copy, 50% split, cookie-based, dates long enough for the volume.
- Leave itNo edits to either arm while it runs.
- DecideApply, convert to a new campaign, or end it.
Traffic split: cookie-based vs search-based
The Experiment split is the share of the original campaign’s traffic and budget the trial gets. Google recommends 50% for the best comparison. A smaller split protects the original if the change is risky, but the trial collects data more slowly and needs longer to reach a result.
| Split type | How it assigns traffic | Use it when |
|---|---|---|
| Cookie-based (Google’s recommendation) | Each user is randomly assigned to one arm and only sees that version | The change affects the whole visit or repeat visits, such as a landing page or bid strategy |
| Search-based | Each search is randomly assigned, so one user can see both versions | You need data faster and seeing both versions doesn’t bias the result |
For Display campaigns Google always uses a cookie split. The custom experiments page also notes that traffic is split by auction: each ad slot is a new auction, so if the trial has different bids or assets, it can win more or fewer auctions than the original even with a 50% split. Uneven impressions between arms aren’t by themselves a sign of a broken test.
A/B testing ads with ad variations
Ad variations are the quickest Google Ads A/B test for ad text, and they are still available. Google’s ad variation setup page puts them under Experiments: click the plus button, then Assets, Assets provided by you, Ad variations. You choose all campaigns or specific ones, select responsive search ads, and describe the change, for example a find-and-replace of “Book now” with “Call now”. You set start and end dates and the experiment split.
The advantage is scale: one variation can change the same phrase across hundreds of ads without editing them one by one. If the variation wins, applying it creates new ads with the change. For tests of whole new ad concepts in one ad group, it is usually simpler to add a second responsive search ad and compare; our CTR guide covers what to compare.
Performance Max experiments
Performance Max doesn’t use custom experiments. Google’s Performance Max experiments page lists three kinds:
- Uplift: measures the extra conversions or conversion value Performance Max brings when added alongside your other campaigns.
- Upgrade: simulates moving budget from an existing campaign into Performance Max, such as a Standard Shopping vs Performance Max test.
- Optimization: tests a change to settings inside a Performance Max campaign, such as text customization and Final URL expansion.
Uplift tests answer a question a before/after can’t: whether Performance Max finds new conversions or takes credit for ones your Search campaigns would have won anyway. If you suspect the second, read Performance Max cannibalizing brand search first.
How long to run a Google Ads experiment
Google’s monitoring guide recommends letting an experiment run for 2 to 3 weeks to gather data. If the Results column still shows In progress, Undecided or Unavailable, the Experiments overview suggests at least 4 to 6 weeks, or a larger budget. When the test changes a bid strategy, the custom experiments page advises allowing 7 to 14 days for the trial to stabilize while the strategy learns, so don’t judge that period.
- Week 0Write the hypothesis
The change, the metric that decides it, and the result that would make you apply it.
- Weeks 1–2Learning
Check both arms serve and spend. Don’t judge results yet.
- Weeks 3–4First read
Look for the asterisk on your primary metric. Conversions from recent clicks may still be arriving.
- Weeks 5–6Decide
Apply, end, or extend once if the result is still undecided and the trend is consistent.
Volume matters more than days. An experiment that tests cost per conversion needs enough conversions in each arm to separate a real difference from noise; a campaign with a handful of conversions a week may never reach one. If customers usually take days to convert after the click, end the experiment and wait for the lag before reading the final numbers.
Reading the results and significance
Open the experiment from the Experiments page. The scorecard shows Clicks, CTR, Cost, Impressions and All conversions by default, and you can add the metrics you chose as goals. For each metric, Google shows the difference between the trial and the original as a percentage and a range.
- Blue asterisk: the difference is statistically significant.
- Confidence interval: 80% by default. You can choose another level such as 95%, which widens the range and makes significance harder to reach.
- The range: a result such as [+8%, +12%] means the trial’s true difference is likely somewhere between those values.
- Not significant
- Significant (asterisk)
View as table
| Item | Difference | Group |
|---|---|---|
| Conversions | +14% | Significant (asterisk) |
| Cost | +9% | Significant (asterisk) |
| Clicks | +3% | Not significant |
| CTR | +1% | Not significant |
Two reading habits help. Decide on the primary metric before you start, so you don’t pick the winner from whichever metric happened to move. And look at cost alongside conversions: a bid strategy that gets 10% more conversions for 30% more spend may not be a win for your margins, whether you sell products, generate leads or sell subscriptions.
Applying or ending the experiment
Google’s apply guide gives two options when you click Apply:
- Update your original campaign: the experiment ends and its changes are copied into the original.
- Convert to a new campaign: the trial becomes a campaign of its own and the original is paused.
Google keeps the performance data of both either way. Updating the original keeps its name and history in one place, which is usually tidier. If the experiment lost, end it; the original never changed. Note that applying a new bid strategy to the original is itself a change, so expect a learning period afterwards; see our Smart Bidding guide.
Common experiment mistakes
- Testing several changes at once. A new bid strategy plus new ads plus a new page tells you the bundle won, not why.
- Editing the original mid-test. Changes to the original aren’t copied to the trial, so the arms drift apart.
- Stopping at the first asterisk. Early results swing. Run the planned duration unless one arm is clearly broken.
- Calling it during learning. Bid strategy tests need the stabilization period before results mean much.
- Too little volume. A small split on a low-volume campaign may never reach significance. Use 50%, or test on a bigger campaign.
- Ignoring conversion lag. Reading results the day the test ends undercounts the trial’s latest conversions.
- Judging on the wrong metric. CTR and clicks are inputs. Decide on cost per conversion, conversion value, or return on ad spend.
When a before/after with undo is enough
Not every change needs an experiment. Experiments are worth the setup when the change is large, hard to reverse, or likely to shift learning: a new bid strategy, broad match across a campaign, a new landing page, or adding Performance Max. For small, reversible changes, such as adding negative keywords, pausing a keyword that has spent with no conversions, adding sitelinks, or a modest budget increase on a campaign limited by budget, a careful before/after is usually enough: change one thing, note the date and the old value, compare equal periods, and reverse it if it hurt.
That is how Boxoo handles routine changes. The performance and structure agent reads change history and shows the numbers before and after each change, and the budgets and bidding agent checks whether a bid strategy has enough conversions to learn. Each finding arrives as a case with the evidence and a prepared fix. Nothing changes until you press Apply, and most applied changes keep an undo, so a small change can be tried and rolled back without a formal test. The big ones still deserve an experiment; see where they fit in our optimization routine.
Run a free Google Ads audit to find which changes in your account are worth testing first.
