Skip to content
← All Essays

Retail Media · 23 June 2026

How to Measure Retail Media Incrementality (Without Fooling Yourself)

Here is the uncomfortable arithmetic at the centre of retail media. A network reports 9x return on ad spend. A properly designed holdout test on the same campaign reports 1.4x. Neither number is a lie. They are answers to different questions.

Attributed ROAS answers: how much revenue can be traced back to someone who saw or clicked the ad. Incrementality answers: how much revenue would not have existed without the ad. The gap between those two numbers is the size of the bill you are paying for sales you already had.

Closing that gap is the single highest-leverage thing a retail media team can do, because it converts a channel brands treat as a trade tax into a channel brands fund from growth budget.

Why attributed numbers inflate

Four mechanisms, all mundane.

Selection. Ads are served to people already in the category, already searching, already loyal. The targeting that makes performance look good is the same targeting that guarantees you reach buyers.

Branded search capture. A shopper types your brand name. You bid on it. You “win” the click and book the sale. Removing that placement entirely often changes total sales by a rounding error, because the organic result sat directly beneath it.

Cannibalisation across your own SKUs. The campaign lifts SKU A by 12 percent. SKU B, from the same brand, drops 9 percent. Attributed reporting shows the win and never shows the offset. Always measure at brand or category level, not SKU level.

Window generosity. A 14-day view-through window will attribute a startling share of all category purchases to any campaign with meaningful reach. Long windows are a measurement choice disguised as a technical default.

The four measurement designs, ranked

1. Randomised holdout (the gold standard)

Randomly assign eligible shoppers to test and control before the campaign starts. Control shoppers are suppressed from the campaign. Compare purchase rate and spend per shopper across groups.

Use when: you have logged-in traffic, sufficient volume, and platform support for suppression. Watch for: leakage (control shoppers seeing the ad through another placement), and underpowered tests. If your minimum detectable effect is 15 percent and the true effect is 5 percent, you will conclude “no impact” and be wrong.

2. Geo experiments

Split markets or store clusters into test and control. Turn the campaign on in test geos only. Compare sales trajectories.

Use when: you cannot suppress at user level, or you are measuring offsite and in-store media where user-level control is impossible. Watch for: market imbalance. Match geos on pre-period sales trend, seasonality, store count, and category share, not just size. Use at least 20 geos per arm where you can.

3. Switchback or time-based tests

Alternate the campaign on and off in short blocks (for example, weekly), then compare on-weeks with off-weeks after adjusting for trend.

Use when: geography is not separable and user-level suppression is unavailable. Watch for: carryover. If ad effects persist past the block length, on and off periods contaminate each other. Use blocks at least twice the expected effect duration.

4. Observational models with a control group

Propensity matching, synthetic control, difference-in-differences against unexposed but similar shoppers.

Use when: nothing else is possible, or you need a continuous read between formal tests. Watch for: everything. These methods correct for observable differences only. The whole selection problem lives in the unobservables. Treat the output as a directional estimate, never as a settled number.

A repeatable seven-step test protocol

  1. Define the business question in one sentence. Not “does retail media work” but “does sponsored search on non-branded category terms produce incremental category units for Brand X over eight weeks”. Vague questions produce unfalsifiable answers.
  2. Choose the outcome metric before the test. Category units, brand sales, new-to-brand buyers. Pick one primary metric. Everything else is secondary and reported as such.
  3. Power the test. Calculate minimum detectable effect from your baseline variance, expected lift, and available volume. If MDE exceeds a lift you would consider a success, the test is not worth running. This step kills perhaps a third of proposed tests, which is the point.
  4. Freeze the design in writing. Test and control definition, window length, exclusions, analysis method, success threshold. Signed off by the brand before launch. This is what stops post-hoc window shopping when the result disappoints.
  5. Run clean. No mid-flight budget changes, no creative swaps, no adding placements. A contaminated test is worse than no test because it produces a confident wrong answer.
  6. Analyse the pre-period first. Verify test and control tracked each other before the campaign. If they did not, the design failed and the result is uninterpretable regardless of what it says.
  7. Report the ratio, not just the lift. Incremental revenue divided by spend, alongside a confidence interval. A point estimate without an interval is marketing, not measurement.

Three rules that prevent self-deception

Rule 1: publish the incrementality-to-attribution ratio. For each campaign type, divide measured incremental ROAS by reported attributed ROAS. Branded search might come out at 0.15. Non-branded category search might be 0.6. Offsite prospecting might be 0.8. Once you have these ratios by placement type, you can discount attributed reporting between tests instead of pretending it is truth.

Rule 2: retest annually, per category. Incrementality is not a constant. It moves with competitive intensity, category penetration, and your own ad load. A 2025 result does not license 2027 spend.

Rule 3: let the network fail its own test in public. The fastest way to become a trusted measurement partner is to publish a study showing one of your own placements is barely incremental, and then to reprice or retire it. Brands remember that. It is worth more in renewal negotiations than any deck.

What to do with a bad result

A campaign that measures at 1.2x incremental ROAS is not necessarily a campaign to kill. It is a campaign to reprice, retarget, or reposition. Common fixes, in order of expected value:

  • Shift budget from branded to non-branded and competitor-conquest terms.
  • Cap frequency on offsite, where the tail of impressions is usually pure waste.
  • Move from broad category targeting to new-to-brand and lapsed-buyer audiences, where incremental headroom actually exists.
  • Reduce the window. If your effect is real, it will survive a shorter window. If it disappears, it was never there.

The general principle is the same one that governs any resource decision under uncertainty: fund the thing that changes the outcome, not the thing that correlates with it. That is also the argument in how to allocate resources like a CEO.

The reporting standard to hold yourself to

Every measurement report a network issues should contain: the design used, the test window, sample sizes per arm, the pre-period parallel check, the primary metric, the point estimate, the confidence interval, and an explicit statement of what the test cannot tell you. Nine lines. If a report cannot produce them, it is not a measurement report.

Retail media will keep growing whether or not anyone measures it properly. But networks that measure honestly get repeat budget from brand P&Ls, and networks that do not get squeezed back into trade negotiations every year. That is the whole strategic difference.

Why standardisation is the number one barrier to investment, what separates same-SKU from halo ROAS, and why every network is a walled garden, are set out in The Third Wave of Digital Marketing. For the channel basics, see what a retail media network actually is and onsite versus offsite retail media.