Pick the right SLO before you worry about the error budget
An error budget only works when the SLO beneath it reflects what users feel. Why copied reliability targets fail, and how to choose an SLI worth measuring.
Articles tagged "best-practices".
An error budget only works when the SLO beneath it reflects what users feel. Why copied reliability targets fail, and how to choose an SLI worth measuring.
A 200 on your homepage proves almost nothing. Why partial outages slip past basic uptime checks, and how to monitor the path customers actually use.
A fishbone diagram maps every contributing cause of an incident, not just one. What it is, how it beats the 5 Whys, and how to run one in a postmortem.
Mean time to recovery is easy to measure and easy to game. What to track after an incident instead, so your postmortems actually change something.
Alert fatigue is a trust problem, not a volume problem. Why false positives erode your team, and how confirming failures across locations fixes it.
When your service is down, silence is worse than the outage. How to communicate during an incident in a way that keeps customer trust.