Ghost ads vs. PSA holdouts: why your control group probably lies
Incrementality experiments live or die on a clean control. Two methods promise one; only one delivers an apples-to-apples comparison.
The contamination problem
To measure lift you compare exposed users against users who were not exposed. But the people an ad system would have shown your ad to are not random — they were selected by the algorithm as high-intent. Comparing them to a random unexposed crowd inflates apparent lift massively. This is selection bias, the deepest pitfall in observational attribution.
PSA holdouts
The old fix: show the control group an unrelated public-service ad. Now both groups "won" the auction, so they are comparable. But you pay for those wasted PSA impressions, and the PSA itself can have a small effect.
Ghost ads
The better design, formalized by Johnson, Lewis, and Nubbemeyer: the system logs which control users would have been served your ad without actually serving anything. No wasted spend, no PSA artifact, and the comparison is between genuinely matched populations — those who won the auction in test vs. their counterparts in control.
The catch
Ghost ads require platform instrumentation you don't control. Outside Google's Conversion Lift and a few walled gardens, you cannot run them yourself.
Bottom line for practitioners: If a platform offers ghost-ad-based lift (Google, Meta Conversion Lift), prefer it over any self-built exposed-vs-unexposed comparison. If only PSA holdouts are available, accept the wasted spend — it buys you a valid counterfactual. Never compare exposed users to a random unexposed sample and call it incrementality.
Credit Where Due
@CreditWhereDue
Ghost ads vs. PSA holdouts: why your control group probably lies
Этот пост опубликован в Telegram-канале Credit Where Due. Подписаться можно по ссылке: @CreditWhereDue.