Deep dive: a pre-launch checklist so your A/B test produces a real answer
Most landing-page A/B tests don't fail to find a winner — they fail to be valid, and the team ships a phantom. The statistics literature on online experiments (the Kohavi/Microsoft body of work is the standard reference) catalogs the recurring sins: peeking, underpowering, and stopping at the first "significant" blip. A pre-launch checklist prevents the most expensive ones.
The checklist, run before traffic flows:
— Step 1: State one primary metric. Picking a winner across five metrics guarantees a false positive eventually (the multiple-comparisons trap). Choose the one that maps to revenue and commit.
— Step 2: Compute sample size in advance. Use your baseline conversion rate and the smallest lift worth caring about; the calculator tells you how many visitors per arm you need. No number means no stopping rule.
— Step 3: Set the duration to cover full business cycles. Run at least one or two complete weeks so weekday/weekend and payday effects average out. Mid-week stops bake in day-of-week bias.
— Step 4: Forbid peeking-and-stopping. Looking is fine; stopping the moment it crosses significance inflates false positives badly. Decide the end date up front and honor it.
— Step 5: Check for sample-ratio mismatch. If your 50/50 split arrives 54/46, the randomization or tracking is broken and the result is untrustworthy — investigate before believing anything.
— Step 6: Pre-register the hypothesis. Write down what you expect and why. It blocks the after-the-fact rationalization that turns noise into a "learning."
Mechanism: an experiment converts a guess into knowledge only if its statistical guarantees hold. Break the assumptions — peek, underpower, multi-test — and you've spent traffic to manufacture confidence in nothing.
The honest caveat: most landers don't get enough traffic to detect small lifts in reasonable time. If the math says you'd need months, don't run an underpowered test that will mislead you — test bigger swings, or rely on principled design instead of pretending you measured.
TL;DR
— One primary metric, sample size computed in advance, full-cycle duration.
— Don't peek-and-stop; check sample-ratio mismatch; pre-register the hypothesis.
— Low traffic? Test bold changes or trust principled design — never an underpowered test.
Above Fold Lab
@AboveFoldLab
Deep dive: a pre-launch checklist so your A/B test produces a real answer
Этот пост опубликован в Telegram-канале Above Fold Lab. Подписаться можно по ссылке: @AboveFoldLab.