4 experiments to test if a channel drives incremental sales or shifts demand
Many marketers assume every ad exposure creates new revenue, yet campaigns often just shift demand between channels. Deciding whether a channel drives true incremental sales or merely cannibalises other channels is essential to optimise resource allocation and measure real business impact.
This post walks through four practical experiments and diagnostics that isolate incrementality, covering how to define metrics, design unbiased tests, implement sampling and tracking, and measure lift. Follow the steps to diagnose bias, translate lift into action, and confidently decide which channels truly add sales rather than just shifting them.
What is incrementality and how do I calculate lift?
Incrementality measures sales caused by a channel versus a counterfactual without that channel, using a clear primary metric such as incremental orders, revenue, or conversion rate; compute lift as (metric_treatment – metric_control) / metric_control, and report both absolute and relative lift with confidence intervals.
How should I design an experiment to isolate true incremental sales?
Randomise exposure at a unit that prevents contamination, use a clean holdout or matched geographic controls, pre-register the primary metric and power assumptions, stratify on key covariates, and hold creative, pricing, and landing experience constant to avoid confounding.
What tracking and instrumentation are necessary to avoid biased uplift estimates?
Assign persistent identifiers, log exposure context and channel touch type, reconcile client-side and server-side conversion records, run daily consistency checks, and preserve a pure holdout while reporting intent-to-treat as the primary estimate.
How do I diagnose bias in results and turn lift into decisions?
Run diagnostics such as pre-test balance checks, pre-period trend plots, placebo tests, and attrition inspection, apply stratification or regression adjustments if needed, compute cannibalisation and net-new share by cohort, and convert results into predefined scale, optimise, or pause decision rules with planned follow-up tests.

1. Clarify incrementality and define clear, measurable target metrics for campaigns
Define incrementality against a clear counterfactual and select a primary metric, for example incremental orders, revenue or conversion rate. Calculate lift as (treatment metric – control metric) / control metric; for example, if the control conversion is 2.0% and the treatment conversion is 2.4%, incremental lift = (2.4 – 2.0) / 2.0 = 20%. Choose a minimum detectable effect, commit to a target statistical power and a significance level, then translate those choices into the required sample size using a power calculator. Record all assumptions explicitly so stakeholders can assess the test sensitivity.
Select a causal design that suits your business and the likely degree of interference between units. Use randomisation at the user, session or geographic level, and adopt cluster randomisation where assigning individuals risks spillover. Describe how you will minimise audience overlap and how you will detect and manage treatment leakage.
Pre-specify guardrail and secondary metrics, for example total sales across channels, channel-specific sales, customer lifetime value, acquisition cost and churn. Define clear thresholds for each metric that will trigger investigation or a pause to the test.
Document the complete analysis plan. Verify identity stitching and your tracking instrumentation with a trial run on historical data. Include analysis of heterogeneity by new and returning customers, by product category and by geography to understand differing effects across segments.
2. Design experiments to isolate and measure true causal lift
Begin by randomising exposure and keeping a clean holdout group, checking pre-experiment parity so treatment and control are comparable. Use difference-in-differences to isolate incremental conversions rather than relying on attribution alone. Run market or geographic tests using matched or synthetic controls, and validate results with permutation or placebo tests to exclude external shocks. Prevent contamination by avoiding overlap across devices or accounts, and monitor spillovers into neighbouring areas so measured effects reflect local exposure. Pre-register your primary metrics, run power calculations to confirm adequate sample size, and report bootstrap confidence intervals plus corrections for multiple comparisons to quantify uncertainty.
Track attribution at both customer and order level across channels to spot whether gains are genuine or simply shifting demand between routes. Segment results by new and returning customers to reveal true incremental acquisition. Keep creative, pricing and landing experiences constant while measuring, and set exposure thresholds to avoid dilution. Measure assisted conversions, average order value and cohort retention to capture downstream effects. Present variation by audience segment and report lifetime value or repeat rates so stakeholders can judge whether short-term uplifts translate into durable growth.

3. Implement effective sampling, tracking and control measures for reliable results
Choose a randomisation unit that reduces contamination, for example individuals, households or geographic clusters. Assign treatments before any targeting and stratify on key covariates such as prior spend, visit frequency and product mix. Check balance with standardised mean differences and available pre-experiment metrics, and adjust your stratification if you find imbalance. Turn the smallest commercially meaningful uplift into an effect size you can detect, and include expected compliance and attrition when you do the power calculation. Decide whether a single large random sample or stratified samples are better for any planned subgroup analysis. Monitor realised base rates and variance early, and be prepared to increase sample size or re-stratify if conversion or variance deviates from your assumptions.
– Assign persistent unique identifiers to instrument exposures and conversions end to end. Log exposure context, channel touch type and attribution window, and capture conversions with the same identifier.
– Reconcile client-side events with server-side records. Run daily consistency checks and flag missing joins to avoid biased uplift estimates.
– Maintain a pure holdout that receives no channel touches, or use geographically or temporally isolated holdouts when individual-level randomisation risks spillover.
– Implement placebo exposures as falsification checks. Report intent-to-treat as the primary estimate and provide per-protocol results where compliance is imperfect.
– Pre-register the primary metric, attribution window and hypothesis. Calculate incremental lift with confidence intervals and run falsification, heterogeneity and carryover diagnostics.
– Present absolute and relative uplift together with uncertainty and sensitivity to attribution choices.

4. Measure impact, identify bias and act on results
Define lift around a single, clear primary outcome, for example incremental revenue per exposed user, conversion rate lift or number of net new purchasers. Report both absolute and relative lift and include uncertainty measures, such as confidence intervals or bootstrapped distributions, so readers can see the range of plausible effects.
Before you trust lift estimates, run targeted diagnostics. Check pre-test balance on key covariates, plot pre-period trends, run placebo tests and inspect for differential attrition or contamination. If diagnostics reveal issues, apply stratification, regression adjustment or re-randomisation as appropriate.
Address measurement and implementation problems explicitly. Quantify exposure misclassification and incomplete compliance, instrument for actual exposure when assignment and delivery diverge, and cluster standard errors when randomisation occurs at the group level.
Break down lift by cohort, product category, channel and lifetime-value band to show where sales are genuinely incremental versus shifted. Compute cannibalisation rates and the share of gross sales that are net new to demonstrate business impact. Report conservative estimates alongside naive estimates so stakeholders can assess robustness. Document every analytical choice and sensitivity check so others can replicate the analysis. Translate findings into predefined decision rules and operational levers: set thresholds for scale, optimise or pause, and adjust targeting, creative or frequency based on results. Plan follow-up tests to confirm persistence, and communicate both the estimated lift and its uncertainty so teams can act with measured confidence.