Q: We just survived an outage. What monitoring changes should the post-mortem include?
A: Every incident is free tuning data. Run this section in your post-mortem, no blame, just gaps.
— Time-to-detect: how long from real failure to first alert? If it was slow, tighten the interval or fix the check.
— Did the right alert fire? If the homepage check passed while checkout was broken, you need a deeper transaction check.
— Did the right person get paged, fast? If not, the escalation chain had a gap.
— Was the status page updated, and how late? Bake the first update into the runbook.
— Add one new monitor for the specific failure mode you just hit so it never surprises you twice.
Follow-up: schedule the change. A post-mortem with action items nobody dates is a post-mortem you'll repeat.
Got a question? Drop it in the comments.
Pingback Clinic
@PingbackClinic
Q: We just survived an outage. What monitoring changes should the post-mortem include?
Этот пост опубликован в Telegram-канале Pingback Clinic. Подписаться можно по ссылке: @PingbackClinic.