Incrementality: what your media really adds
Incrementality in marketing: what it means, why platform numbers overstate it, and the three ways to measure it: geo-lift, holdouts, calibrated MMM.
Hajime Takeda · Updated September 28, 2026
What incrementality means
Incrementality is the part of a result that advertising caused. Not the sales that happened while the campaign ran, not the sales a platform attributed to its clicks, but the sales that would not have happened without it. Everything else, the orders from customers who were coming anyway, is baseline, and a channel that takes credit for the baseline looks better than it is.
The idea needs a counterfactual: what would have happened without the advertising. Nobody observes it directly, so incrementality in marketing is always an estimate of a difference, between what happened and what would have happened, and the whole discipline is about building that second number honestly. There are two ways to build it. An experiment creates it on purpose, by withholding the advertising from a group that would have behaved like the rest. A model estimates it from history, by separating the media from everything else that moved sales. Both give an incremental effect with an interval, and the interval is part of the answer.
Incrementality is measured on an outcome the business cares about, revenue, orders, sign-ups, store visits, and at a spend level. A channel is not incremental or not; it adds a certain amount at a certain spend, and that amount changes as spend changes. The marginal incremental ROAS is that idea read on the next euro rather than on the average, and it is the number a plan moves budget on.
Why platform attribution overstates incremental revenue
Platform dashboards report attributed sales: the conversions their tracking could tie to an impression or a click, credited by a rule. Attribution overlaps across channels, because several platforms see the same user and each claims the sale; it counts the customers who would have converted anyway, because the rule has no counterfactual; and it stays high on saturated channels, because the average return keeps looking healthy long after the next euro stopped earning it. Add the sales reported by every channel and the total can exceed the sales the business actually made.
Retargeting is the textbook case: it reaches people who have already shown intent, so its attributed conversion rate is high and its incremental revenue can be small. Branded search is another: television creates branded searches, and the search platform reports the resulting sales as its own. In both cases the dashboard is not lying about what it tracked. It is answering a different question from the one a budget needs. Attributed revenue says which touchpoints were on the path. Incremental revenue says what the money changed.
Three ways to measure incrementality: geo-lift, holdouts, calibrated MMM
Geo-lift. A geo-lift experiment changes advertising in some regions and not in others: most often a holdout, where a channel is paused in the test regions for a fixed period. A synthetic control, a weighted mix of the untreated regions built to track the test regions before the change, gives the counterfactual, and the gap during the test is the incremental effect. It works without any user-level data, which is why it is the workhorse for offline channels, for channels that lost their tracking, and for the digital channels whose platform figures are least trusted. The design is fixed before launch: KPI, regions chosen for measurability, duration, analysis method, and the smallest effect the test could detect.
Holdouts. Where the unit is a user or a customer rather than a region, the advertising is withheld from a randomly chosen group. Platform lift studies do this inside one platform, with the platform running the randomisation; a holdout on a customer list does it for email or retargeting; an A/B test does it for a creative or an offer. These tests answer a narrow question: the incremental effect of that channel, on that platform’s audience, at that spend. They cannot see across channels, and a platform measuring its own lift is measuring inside its own audience and rules, so the design and the readout should be checked.
Calibrated MMM. A media mix model estimates the incremental contribution of every channel, every week, from aggregate history, and that is its strength and its weakness: it covers the whole mix, but it is observational, and a well-fitted model can still credit the wrong channel when channels move together. Calibration is the fix. The measured result of a geo test or a holdout becomes a prior on that channel’s coefficient in the next fit, so the model’s curve is anchored at the tested point and the rest of the curve fills in between tests. When the model and the test differ by more than a third, the model has an identification problem on that channel, and the plan should say so.
The three are not rivals. Experiments measure; the model generalises the measurements to every week and every channel; and the next experiment goes where the model is least sure.
How to read an incrementality result: the interval matters
An incrementality test returns an estimate and an interval, and the interval is not decoration. Two illustrative examples: a measured lift of 12% with an interval from 3% to 21% supports a positive effect under the tested conditions and says little about its size; a lift of 4% with an interval from minus 2% to 10% is not distinguishable from zero, and reporting the 4% without its interval overstates what the test showed. The plan should carry the band and say when it crosses zero.
Reading a result also means reading the design. Was the test designed before it ran, or was a spend cut that happened for other reasons measured after the fact? A retrospective test can be worth measuring, but it was not designed, its ramp was not clean, and its regions were not chosen for measurability, so it should be reported as what it is. Did the intervention actually happen, spend on the test regions falling to zero on the date, or did a gradual ramp blur the boundary? Were the placebo checks run, a time placebo on an earlier window where nothing changed, a region placebo on untreated regions, and reported? And was the test large enough to see the effect the plan needed to see? A holdout of two small regions for three weeks detects only a very large effect, and the design step exists to say so before the money is spent. And was the switch real? Media bought nationally, or spillover across a region boundary, treats the control regions too.
Finally, one test measures one spend change at one level. It anchors the response curve at that point; the shape of the rest of the curve still comes from the model. A channel that was incremental at half its budget is not necessarily incremental at twice its budget.
Incremental sales, incremental lift, incremental value: the same idea, three words
The vocabulary varies more than the concept. Incremental sales, or incremental orders, are the units caused by the advertising. Incremental revenue is the same thing in money. Incremental lift is the effect expressed relative to the counterfactual: a 12% lift means the test group’s outcome was 12% above what it would have been. Incremental value is the broader term, used when the outcome is not revenue but sign-ups, visits or lifetime value. Incremental ROAS is incremental revenue divided by the spend that caused it, and marginal incremental ROAS is the same ratio read on the next euro.
Two terms are often confused with these. Lift, on its own, is what platforms call the result of their own lift studies, which is an incrementality measure limited to that platform’s audience and rules. And uplift, in modelling, usually means a model that predicts which individuals respond to a treatment, a different question from how much a channel adds in aggregate. When a report uses one of these words, the useful question is always the same: incremental compared to what counterfactual, measured how, at what spend, with what interval.
Incrementality vs attribution vs MMM
Attribution works at the user level and shares the credit for each observed conversion among the touchpoints it tracked, by a rule. It is granular and fast, and it is the usual tool for optimising inside a channel: which keyword, which audience, which creative. It measures correlation along a path, not causation, and it cannot see what it cannot track, which excludes offline media and a growing share of users.
Incrementality testing measures causation directly, for one channel, at one spend level, over one period. It is the only method that measures what a channel causes, under the conditions of the test. Its limits are cost and scope: one question at a time, a clean intervention, and a result that is a point on the curve rather than the curve.
Marketing mix modeling works at the aggregate level, weekly and by channel, online and offline, and answers the budget question across the whole mix. It is observational, it needs history and variation, and its resolution stops at the channel and the week. Calibrated with experiments, it is the method that turns a handful of measured points into a plan for every channel.
The practical arrangement is a hierarchy: experiments calibrate the model, the model allocates the budget across channels, and attribution optimises inside each channel, with its figures read as attributed, not incremental. The marketing mix modeling pillar goes through the model side of that arrangement in detail.
How the Kuwalyst Media Planner uses incrementality tests
The Media Planner proposes experiments as recommendations, from the first plan. When a channel’s marginal band is too wide to size a budget move, or when its spend never varied enough for the model to identify it, the Tune view proposes the geo holdout that would narrow it: the regions, the duration, and the expected tightening of the interval. Accepting the recommendation puts the experiment on the plan’s timeline. For a brand starting on benchmarks, the plan lists the benchmarks it depends on, ranked by how much the projection would move if they were wrong, and proposes the test or the data connection that would replace each one with a measured value.
The readout lands in the same workspace, tagged measured, next to the model’s figure tagged projected, and the plan shows both. The measured incremental return becomes a prior on the channel’s coefficient in the next fit, and the discipline above is enforced every time: the design is fixed and stored before launch, the placebo checks and the spend check are run and reported, and a lift whose band crosses zero is reported as not distinguishable from flat. It is the same discipline Kuwalyst applies when agencies ask it to review their own geo-lift measurement, and the reason brands can say, at the end, which of their figures were measured and which were modelled.