Geo-lift experiments
A geo-lift experiment compares regions where advertising is paused with similar regions where it continues. It measures incremental effect where user-level tracking is not possible and can calibrate a media mix model against observed results.
8 min read · Updated September 7, 2026
What it is
A geo-lift experiment changes advertising in some geographic regions and not in others, then measures the difference in a business outcome. The regions are the unit of randomisation, which is why it works without any user-level data: it needs only sales and spend by region and week.
The most common design is a holdout: pause a channel in the test regions for a fixed period. The most common measurement is a synthetic control, a weighted mix of the untreated regions built to track the test region before the change. The gap between the test region and its synthetic twin during the test is the incremental effect.
Why it matters for a media plan
A media mix model is an observational estimate. It can be well fitted and still credit the wrong channel, because channels move together. An experiment is the only way to measure what a channel causes, and the model’s estimate for that channel can then be compared to it. If they differ by more than a third, the model has an identification problem on that channel and the plan should say so.
A geo test also produces the one kind of number that needs no caveat: measured. In a plan where every figure is tagged, the measured ones are what the others are checked against.
How it works
Design, before anything runs: Fix the KPI, the candidate test regions, the duration and the analysis method. The design step uses the pre-period data to compute, for each candidate design, the smallest effect the test could detect with reasonable power. A test that cannot detect the effect the plan needs to see should not run.
Test regions are chosen for measurability: the regions whose history the others can reproduce. A region that moves unlike the rest makes a poor synthetic twin whatever its size.
Intervention, clean: Switch spend on or off on one date. A gradual ramp-down before the test blurs the boundary between the normal period and the intervention, and the transition weeks end up counted on the wrong side.
Measurement, on the locked design: Fit the synthetic control on the pre-period, project it through the test, read the gap. Then validate: a time placebo, running the same measurement on earlier windows where nothing changed, and a region placebo, running it on untreated regions. Report both. Cross-check with a second method, such as a Bayesian structural time series, as a supporting check.
Readout: the incremental effect on the KPI, its interval, the spend saved or added, and the incremental return that follows. That number then calibrates the model: it becomes a prior on the channel’s coefficient in the next fit.
How Kuwalyst uses it
The Media Planner proposes experiments as recommendations, from the first plan. When a channel’s marginal interval is too wide to size a budget move, the Tune view proposes the geo holdout that would narrow it: the regions, the duration, the expected tightening of the interval. Accepting it puts the experiment on the plan’s timeline; the readout lands in the same workspace and feeds the next fit.
The planner insists on the discipline above. The design is fixed before launch and stored. The validation checks, time and region placebos, pre-period fit, a spend check that the intervention actually happened, are run and reported every time. It is the same discipline Kuwalyst applies when agencies ask it to review their own geo-lift measurement.
Pitfalls
The retrospective test: a spend cut that happened for other reasons can be measured after the fact, and often should be. But it was not designed, its ramp was not clean, and its regions were not chosen for measurability. Report it as what it is.
Seasonal windows: the synthetic control works when the test region moves like the controls. In some categories it does not in summer, or around a national holiday, and the placebo error in those windows is several times larger. Check the placebo error by calendar placement before choosing dates.
Contamination: media bought nationally, or spillover across a region boundary, treats the control regions too. Choose channels and regions where the switch is real.
Too small to see: a holdout of two small regions for three weeks will detect only a very large effect. The design step exists to say so before the money is spent.
Reading one test as a curve: a test measures the effect of one spend change at one level. It anchors the response curve at that point. The shape of the rest of the curve still comes from the model.
See it in the product
