Every account we audit has at least one campaign that, according to the report, performs beautifully and that nobody has ever tried switching off. Often it is brand search, sometimes remarketing, increasingly a Performance Max campaign that has claimed credit for sales the site would have made anyway. The report is not lying: it answers, precisely, a different question from the one you care about.

Attribution and incrementality: two different questions

Attribution answers: "of the touchpoints that preceded this conversion, which gets the credit?". The data-driven model does this in a sophisticated way, as we covered in the guide to data-driven attribution across Google Ads and GA4. But it always starts from a conversion that happened and divides the credit among the clicks it observed.

Incrementality answers: "had I not spent this money, how many conversions would I have lost?". It is a counterfactual question, and no attribution model can answer it, because none ever observes the world in which the advertising was absent. The only way to observe that world is to create it: remove the advertising from part of the audience, or part of the territory, and compare.

The textbook case. Someone searches for your shop’s name, clicks the brand ad and buys. Attribution gives the sale to the ad. But without the ad they would have clicked the first organic result, which is your site. The conversion is attributed, not incremental.

The distinction is not academic. Set budgets on attributed CPA and you end up rewarding the campaigns best at intercepting people who were about to buy anyway, and penalising the ones that create new demand but rarely receive the last click.

Google Ads custom experiments create a copy of a campaign, apply a change — a bid strategy, a match type, a landing page — and split traffic between the original and the variant. There are two split options:

  • Cookie-based, the one Google recommends: each user sees only the original or only the variant, so the comparison is cleaner.
  • Search-based: assignment happens on every search, so the same user can see both versions. It reaches significance sooner, but dilutes the effect on behaviour that plays out over several visits.

Google recommends a 50% split for the best comparison, and it is also the split that reaches a readable result fastest. Experiments are the right tool for "does variant B beat A?". They do not answer "is this campaign worth running?", because both arms are advertising. That question needs a control group that sees no advertising at all.

Geo holdout tests

A geo test is the most accessible way to measure a channel’s incrementality, and it needs no special access: split the territory into two groups, keep advertising running in one and switch it off (or reduce it) in the other, then compare total sales, not attributed ones.

How to build a clean test

  1. Choose areas that behave alike. In Italy that usually means regions or provinces. Take at least eight to ten weeks of history and check that the two groups’ sales move together: if the ratio between them is stable over time, you have a sound basis for comparison.
  2. Measure the business metric, not the platform’s. Orders from the back office, qualified leads from the CRM, revenue. Google Ads conversions in the switched-off group are zero by definition and tell you nothing.
  3. Freeze everything else. No promotions in one group only, no price changes, no new campaigns on other channels in a single area.
  4. Fix duration and decision criterion in advance. Write them down before, stick to them after.

Budget for the limits: people move around, boundaries between areas are not clean, and channels that run nationally — television, print, a post that goes viral — can contaminate the comparison. That is why geo tests work best on large volumes and in periods without unusual events.

Conversion Lift: the Google-run version

Conversion Lift is the tool Google uses to measure a campaign’s incremental conversions directly, now grouped in the Experiments section alongside the other lift studies. It comes in two flavours:

  • Geography-based: available across a broad range of campaign types, including Search, Shopping, Performance Max, Demand Gen, Video and Display.
  • User-based: compares exposed users with a held-out group; for Search, Shopping, Display and Performance Max, access goes through the account’s Google representative.

The most notable recent change is the spend threshold: in 2025 Google lowered the minimum for incrementality studies to 5,000 dollars, adopting a Bayesian statistical approach that works with less data and giving a feasibility rating before launch. In plain terms, it is no longer a tool reserved for large advertisers. It remains true that this is Google measuring Google, and the methodology cannot be inspected the way a test you design yourself can; so where budget allows, an independent geo test is a good cross-check.

A complete worked example

An Italian ecommerce business running non-brand Search plus Performance Max. Google Ads attributes around 900 orders every four weeks to the campaigns, on 18,000 € of spend in those areas. Attributed CPA: 20 €. We split the regions into two groups, A and B, of similar volume.

Over the previous ten weeks, group B’s orders were consistently 95% of group A’s, moving between 93% and 97%. We then switch the campaigns off in group B for four weeks.

Group A (live)Group B (off)
Orders during the 4-week test4,0003,420
Expected orders in B with no change (95% of A)—3,800
Orders lost by switching campaigns off—380
Observed B/A ratio85.5% (against 93-97% historically)

In group B the campaigns were really producing 380 orders, not 900. Incrementality is about 42% of what was attributed, and incremental CPA is 18,000 € / 380 = 47 €, not 20 €. With a margin of 60 € per order the campaigns are still profitable, just far less than they looked; with a margin of 35 € you are losing money on every incremental order while the report shows an excellent CPA.

Why the result is credible. The B/A ratio never dropped below 93% in the history; during the test it sat at 85.5%. The gap is much larger than normal week-to-week movement. Had it fallen to 94%, the honest conclusion would have been "no measurable effect", not "a 1% effect".

How long to run it, and how many conversions you need

Duration is decided in conversions, not days. A practical rule of thumb for a two-group comparison at the usual confidence levels: conversions needed per group are roughly 16 divided by the square of the relative difference you want to detect, adjusted for conversion rate. Concretely, at a 3% conversion rate:

  • to detect a 10% difference you need about 1,500 conversions per group;
  • for a 20% difference, about 390;
  • for a 5% difference, over 6,000.

Halving the effect you want to measure quadruples the sample you need. That is why a small account should only test large changes, and why a two-week test on 80 conversions proves nothing in either direction.

Beyond volume, three rules on duration. Cover at least two or three full weekly cycles, because Monday and Saturday are not alike. Add the account’s typical conversion lag, or the live group’s sales will arrive after the test has ended. Avoid unusual periods: a test that runs across Black Friday measures Black Friday.

How to read results without fooling yourself

  • Do not check daily to decide when to stop. Stopping the first time the difference looks significant inflates false positives. Duration is fixed beforehand.
  • Read the interval, not the point estimate. "12% lift, with an interval from 2% to 22%" is a very different result from "lift between 11% and 13%".
  • Do not slice the data afterwards. If the total is not significant, finding a segment — mobile, one region, a time of day — where it is will almost always be noise.
  • A null result is a result. "We see no difference at this volume" is useful: it says the effect, if any, is smaller than you can afford to measure.
  • A test holds for the context it ran in. A channel’s incrementality shifts with seasonality, competition and spend level. Repeat it; do not frame it.

This applies especially to the campaigns reporting tends to over- or under-credit: the Performance Max channel report shows where spend goes, not what it caused, and Demand Gen campaigns nearly always lose out in an attributed-CPA comparison.

Where to start

If you have never run an incrementality test, start with the campaign that combines the lowest attributed CPA with the highest spend: that is where the gap between attributed credit and real contribution can be widest, and where an answer moves the most budget. The rest of the account is easier to optimise once you know that number.

In our free audits we point out which campaigns deserve a test and how to set it up with the volume available. You can see how we work, and the answers to the questions we get most often in the site’s FAQ.