Statistically significant. Practically worthless.
Huge test, 200k visitors. Variant won, p < 0.001. Rock solid significant.
The lift? 0.4%.
With enough traffic ANY difference becomes significant. Significance just means "real," not "big enough to care about." That 0.4% wouldn't pay for the dev time to ship it.
P-value answers "is it real." It never answers "is it worth it."
— Set a minimum lift threshold BEFORE testing (e.g. "must beat +3%")
— Big traffic makes tiny effects significant, ignore them
— Look at the confidence interval, not just the p-value
Go set a "don't bother shipping under X%" rule with your team. Report back what number you pick.
Split Test Street
@SplitTestStreet
Statistically significant. Practically worthless.
Этот пост опубликован в Telegram-канале Split Test Street. Подписаться можно по ссылке: @SplitTestStreet.