How do you measure campaign incrementality with a control group?
The primary estimand is the absolute outcome difference between assigned groups. Relative lift and incremental conversions are secondary transformations with explicit denominator and population.
Direct answer
Estimate campaign ITT and its interval.
The primary estimand is the absolute outcome difference between assigned groups. Relative lift and incremental conversions are secondary transformations with explicit denominator and population.
01
Direct answer
Evaluate a randomized campaign first through the absolute outcome difference between all units assigned to treatment and control. Relative lift and incremental conversions are secondary; they do not replace the ITT or its interval.
02
Decision question
For the eligible population and prespecified window, what conversion difference does the assignment policy produce, with what precision, and is that range compatible with the declared value threshold?
03
Population, unit and horizon
Population: units eligible before assignment. Unit: the customer, cookie, store or geography actually randomized. Horizon: 14 days in the example. Each unit remains in its assigned arm for analysis.
04
Estimand and design
The primary estimand is the average effect of assignment, ITT = E[Y(1)−Y(0)]. Randomization targets baseline exchangeability; consistency, positivity and no interference remain necessary for interpretation.
05
Required data
Retain unit identifier, eligibility, assignment, outcome, observation status, exposure, window and exclusions decided before the test. Aggregated counts alone cannot diagnose duplicates, attrition or contamination.
06
Identification assumptions
The assignment sequence must be unpredictable and applied at the declared unit; treatment and control must be well defined; follow-up must not depend differentially on arm; one unit must not affect another unit’s outcome without an interference model.
07
Model and symbols
For a binary outcome, p̂1=y1/n1 and p̂0=y0/n0. The primary effect is Δ̂=p̂1−p̂0. Lift is Δ̂/p̂0; incremental test-arm conversions equal n1×Δ̂. Always declare the denominator.
08
Reproducible calculation
Read the CSV, verify two unique arms and denominators, recompute rates from counts, calculate Δ̂, its standard error and interval, then derive lift and incremental volume. No seed is involved: the example contains realized counts.
09
Reference result
Control: 1,600/20,000 = 8.00%. Treatment: 1,900/20,000 = 9.50%. ITT = +1.50 pp; secondary lift = 18.75%; 300 incremental conversions in the treatment arm.
| Arm | n | y | Rate |
|---|---|---|---|
| Control | 20 000 | 1 600 | 8.00% |
| Test | 20 000 | 1 900 | 9.50% |
10
Uncertainty
The normal approximation gives SE=0.2826 pp and 95% CI [0.946, 2.054] pp. With small counts or extreme rates, prefer a score interval; with clustered assignment, estimate variance at the randomization level.
11
ITT, exposure and adherence
The treatment arm has 17,400 exposures for 20,000 assignments. ITT retains all 20,000 units and estimates the assignment-policy effect. Restricting to exposed units changes the estimand and introduces post-randomization selection.
12
Missing outcomes
The reference data have no missing outcomes. In production, report missingness by arm, explain its mechanism and run a bounded sensitivity analysis. Differential deletion after assignment can break randomization’s benefit.
13
Protocol diagnostics
Check expected sample ratio, unit duplicates, pretreatment balance, tracking stability, contamination, interference and window compliance. These diagnose the protocol; they do not replace estimation.
14
Common errors
Typical errors: analyze exposed users as if randomized, exclude nonconverters post hoc, call an absolute difference lift, multiply ITT by the wrong population, stop at the first favorable result or ignore multiplicity.
15
Interpretation
Under the declared protocol, the range is compatible with a positive policy effect for this population and window. It does not describe the effect among exposed users only, persistence beyond 14 days or transport to another campaign.
16
Supported decision
Compare the ITT lower bound with a value threshold fixed before analysis, then choose rollout, another bounded pilot or stop while including cost, risk and capacity. The decision remains human and conditional.
17
Unsupported decision
Do not attribute the effect to exposed users only, to one channel in a multichannel campaign, to demand creation rather than temporal displacement, or to future performance outside the observed population and window.
18
Implementation and deliverable
The minimum deliverable contains frozen protocol, flow diagram, arm counts, absolute ITT, 95% CI, secondary lift, incremental volume, diagnostics, value threshold, limitations and supported/unsupported decision. The CSV and Python script below reproduce the figures.
CSV · CC0
msc-p013-campaign-incrementality.csv ↓Python · MIT
msc-p013-reference.py ↓19
Scientific sources
Hernán and Robins support counterfactual reasoning, randomization and ITT. Fagerland, Lydersen and Laake document intervals for differences in proportions and Wald limitations. These sources prove no commercial performance.
- Hernán & Robins, Causal Inference: What If ↗Full text verified
- Fagerland, Lydersen & Laake (2015) ↗Full text verified
Dataset · Tool
Method connections
