SSL certificates always expire on the worst day
An expired certificate takes your whole site down in a way no code change can fix fast, and it is entirely predictable. Why cert expiry is the outage you can see coming.
More articles on incident management, monitoring, and on-call.
An expired certificate takes your whole site down in a way no code change can fix fast, and it is entirely predictable. Why cert expiry is the outage you can see coming.
A status page is not a dashboard for you. It is a trust tool for your customers, and it only works if it tells the truth when the truth is inconvenient.
When DNS breaks, your servers are fine, your usual checks may be fine, and your users cannot reach you at all. Why DNS failures are so easy to miss and how to catch them.
Most runbooks are written once, never opened, and useless by the time you need them. What separates a runbook that helps mid-incident from one that just exists.
Severity levels exist to tell people how hard to run. When every incident is a P1, they tell people nothing. How to keep severity meaningful.
Most incidents that fall through the cracks fall through at the handoff between shifts. How to hand off on-call so the context travels with the pager.