Your uptime check passed while the service was down
A 200 on your homepage proves almost nothing. Why partial outages slip past basic uptime checks, and how to monitor the path customers actually use.
Articles tagged "incidents".
A 200 on your homepage proves almost nothing. Why partial outages slip past basic uptime checks, and how to monitor the path customers actually use.
A fishbone diagram maps every contributing cause of an incident, not just one. What it is, how it beats the 5 Whys, and how to run one in a postmortem.
Mean time to recovery is easy to measure and easy to game. What to track after an incident instead, so your postmortems actually change something.
Alert fatigue is a trust problem, not a volume problem. Why false positives erode your team, and how confirming failures across locations fixes it.
When your service is down, silence is worse than the outage. How to communicate during an incident in a way that keeps customer trust.
In an outage, restarting the unhealthy box is often the fix. In a breach, it is how you destroy the evidence you needed. Why your outage instincts betray you in a security incident.