Our alerts tell us something broke but not what to do. Does that matter?
Q: We get a page that just says 'API down.' Then we scramble. Is there a measurable benefit to adding more to the alert?
A: There is, and a reader quantified it. Their bare alerts forced on-call to rediscover context every time — which dashboard, which logs, who owns it. Median time-to-resolution was 47 minutes, and a chunk of that was just orientation.
They enriched each alert with three things: a direct link to the relevant dashboard, the last deploy that touched the service, and a one-line runbook link. No new monitoring, just better alert payloads.
Median time-to-resolution dropped from 47 to 19 minutes over the next quarter. Most of the saving was eliminating the 'where do I even look' phase.
The follow-up: keep runbooks short and current. A linked runbook that's six months stale is worse than none, because people trust it and follow a dead path. Review them when the service changes.
Got a question? Drop it in the comments.
Pingback Clinic
@PingbackClinic
Our alerts tell us something broke but not what to do. Does that matter?
Этот пост опубликован в Telegram-канале Pingback Clinic. Подписаться можно по ссылке: @PingbackClinic.