Above Fold Lab
Above Fold Lab
@AboveFoldLab

Deep dive: a pre-launch checklist so your A/B test produces a real answer

Deep dive: a pre-launch checklist so your A/B test produces a real answer

Most landing-page A/B tests don't fail to find a winner — they fail to be valid, and the team ships a phantom. The statistics literature on online experiments (the Kohavi/Microsoft body of work is the standard reference) catalogs the recurring sins: peeking, underpowering, and stopping at the first "significant" blip. A pre-launch checklist prevents the most expensive ones.

The checklist, run before traffic flows:

— Step 1: State one primary metric. Picking a winner across five metrics guarantees a false positive eventually (the multiple-comparisons trap). Choose the one that maps to revenue and commit.
— Step 2: Compute sample size in advance. Use your baseline conversion rate and the smallest lift worth caring about; the calculator tells you how many visitors per arm you need. No number means no stopping rule.
— Step 3: Set the duration to cover full business cycles. Run at least one or two complete weeks so weekday/weekend and payday effects average out. Mid-week stops bake in day-of-week bias.
— Step 4: Forbid peeking-and-stopping. Looking is fine; stopping the moment it crosses significance inflates false positives badly. Decide the end date up front and honor it.
— Step 5: Check for sample-ratio mismatch. If your 50/50 split arrives 54/46, the randomization or tracking is broken and the result is untrustworthy — investigate before believing anything.
— Step 6: Pre-register the hypothesis. Write down what you expect and why. It blocks the after-the-fact rationalization that turns noise into a "learning."

Mechanism: an experiment converts a guess into knowledge only if its statistical guarantees hold. Break the assumptions — peek, underpower, multi-test — and you've spent traffic to manufacture confidence in nothing.

The honest caveat: most landers don't get enough traffic to detect small lifts in reasonable time. If the math says you'd need months, don't run an underpowered test that will mislead you — test bigger swings, or rely on principled design instead of pretending you measured.

TL;DR
— One primary metric, sample size computed in advance, full-cycle duration.
— Don't peek-and-stop; check sample-ratio mismatch; pre-register the hypothesis.
— Low traffic? Test bold changes or trust principled design — never an underpowered test.
Этот пост опубликован в Telegram-канале Above Fold Lab. Подписаться можно по ссылке: @AboveFoldLab.
tech

Свежие посты в категории «Tech Infrastructure»

Все каналы категории →

start

Готовы запустить рекламу через сеть public.tg?

Новый оффер, продукт, GEO, кейс, событие или партнёрский запуск — соберём маршрут под задачу и отдадим медиаплан.

Telegram для медиаплана: @AFFtop_connect. Быстрый тест: $20 за канал, $1000 за пакет по сети.