Will widening my check interval really cut false alarms, or just hide outages?
Q: I run 30-second checks and get woken up constantly. If I go to 60 seconds, am I just blinding myself?
A: One reader's case is the clearest answer I've seen. They ran a single-region 30s HTTP check on a Node API and logged 41 pages over 30 days. When they pulled the raw history, 33 of those 41 were single-failed-checks that recovered on the very next poll — transient TCP resets, not real downtime.
They switched to a confirm-on-failure rule: 60s interval, but only alert after 2 consecutive fails. Pages dropped from 41 to 6 in the next month. Crucially, all 6 were genuine outages, and mean detection time only slipped from 30s to about 110s — a price worth paying.
The follow-up you're thinking: yes, keep 30s for your payment endpoint where 80 extra seconds matters. Tune per-endpoint, not globally.
Got a question? Drop it in the comments.
Pingback Clinic
@PingbackClinic
Will widening my check interval really cut false alarms, or just hide outages?
Этот пост опубликован в Telegram-канале Pingback Clinic. Подписаться можно по ссылке: @PingbackClinic.