Why did a winning test result shrink by half three months after launch?
A product team shipped a feature that A/B tested to a 12% conversion lift, then watched the gain erode to 6% over the following quarter. The question: was the original test wrong, or did the effect itself change?
The phenomenon. This is novelty-and-primacy decay, documented across experimentation literature (including published work from large-scale experimentation platforms). A randomized test measures the effect during the test window — but if the lift partly came from users reacting to newness, the true long-run effect is smaller. The experiment was causally valid; the causal effect was simply time-varying.
What the studies show. Analyses of large experiment portfolios find that a meaningful share of "winning" tests show effect decay, and that short tests systematically over-estimate long-run impact when novelty is present. Documented cases report initial lifts halving within weeks once the novelty faded; conversely, some learning-curve features under-estimate because users improve over time (primacy).
The nuance. Even a clean randomized experiment can mislead if you assume the measured effect is stationary. Correlation-versus-causation isn't the only trap — a real causal effect that decays will fool you if your test window is too short to see the steady state.
Bottom line for practitioners: a valid A/B result tells you the effect during the test, not forever. For features prone to novelty, run longer or add a holdback cohort that keeps the old experience for months, then re-measure. The 12% that became 6% was both numbers being true — at different times.
Credit Where Due
@CreditWhereDue
Why did a winning test result shrink by half three months after launch?
Этот пост опубликован в Telegram-канале Credit Where Due. Подписаться можно по ссылке: @CreditWhereDue.