A static alert (`error rate > 1% over 5 minutes`) paged the on-call rotation 20 times in a month, 11 of which corresponded to no real incident, and 7 of which were actually the same slow-burning incident firing and clearing repeatedly. What's wrong with the rule's design, and how does a symptom-based burn-rate alert avoid both failure modes at once?
A single static threshold over one short window has no sense of scale — six errors out of forty requests at 4am is "more than 1%" exactly the same as six errors out of six hundred requests at peak traffic, even though the first is statistical noise from a tiny sample and the second might be a real problem, and it has no way to distinguish a genuinely bad five minutes from ordinary variance bouncing just above and below a fixed line. Both of those show up in the lesson's own measured month: 11 pages for nothing, mostly at low-traffic hours where a handful of errors trips the percentage; and 7 separate pages for one real slow-burn incident that hovered near the 1% line and kept crossing it back and forth, which looks like flapping noise rather than the single ongoing problem it actually was. A burn-rate alert fixes both by asking a scale-aware question — at this pace, how much of the error budget will be spent — using a long window for statistical significance and a short window for fast recovery, which is exactly why it paged only 3 times in the same month, once per real incident, and never for nothing.