Your uptime check passed while the service was down
A 200 on your homepage proves almost nothing. Why partial outages slip past basic uptime checks, and how to monitor the path customers actually use.
Articles tagged "monitoring".
A 200 on your homepage proves almost nothing. Why partial outages slip past basic uptime checks, and how to monitor the path customers actually use.
If your uptime checks run from one location, you are seeing one network path, not your users. Why single-node monitoring misses real outages and invents fake ones.
An expired certificate takes your whole site down in a way no code change can fix fast, and it is entirely predictable. Why cert expiry is the outage you can see coming.
When DNS breaks, your servers are fine, your usual checks may be fine, and your users cannot reach you at all. Why DNS failures are so easy to miss and how to catch them.
Your app can be perfectly healthy and still be down because something it relies on failed. Why you should monitor your dependencies, not only yourself.
Knowing something broke is half the job. Knowing it came back, and how long it was down, is the other half. Why recovery notifications matter.