Metrics and percentilesmedium3-5 years

A dashboard shows average latency for an endpoint holding steady at 45 ms, and an alert is configured to page when that average exceeds 100 ms. Users are complaining about slow page loads. What's likely wrong with the alert, and what should it watch instead?

An average blends every request into one number, and a healthy majority can hide a broken minority almost completely — the lesson's own measured example has a mean of 45.6 ms while almost all requests took about 20 ms and roughly one in a hundred took nearly two seconds; no request actually took 45.6 ms, it's just where the math lands when you mix fast and slow. A page built from ten backend calls means most page loads include at least one call from that slow tail, so users feel the p99, not the average, even while the average dashboard looks completely fine. The fix is alerting on a percentile — p95 or p99 — and, even better, on the share of requests that finished within a specific target, because that's the number that actually reflects what a user experienced on a given page load.

The lesson behind it →
More on Metrics and percentiles