Q: A site is either up or down. Why complicate it?
A: Because the binary view hides the failures that hurt most. Plenty of real incidents are partial: the homepage loads but checkout 500s, the API works but is 8 seconds slow, one region is broken while others are fine. Your simple up/down monitor calls all of that up.
The myth is that availability is one switch. In practice it's per-endpoint, per-region, and per-latency-threshold.
Recommendation: monitor your critical user journeys separately (login, search, checkout), and treat slow as a form of down by setting a response-time threshold that pages, not just a connection check.
Anticipating the follow-up: pick the latency limit from real user tolerance, often 3-5 seconds for a page, much tighter for an API.
Got a question? Drop it in the comments.
Pingback Clinic
@PingbackClinic
Q: A site is either up or down. Why complicate it?
Этот пост опубликован в Telegram-канале Pingback Clinic. Подписаться можно по ссылке: @PingbackClinic.