Alerting and SLOshard8+ years
A team's SLO period is 7 days, and they want a 'page' rule that fires once a sustained rate would spend 1% of that week's error budget within a 30-minute window. Derive the burn-rate threshold this rule should use, and explain why it isn't the same as the standard 14.4×-over-1-hour rule even though both are labeled 'page'.
The formula is budget consumed = burn rate × (window ÷ period), so solving for the burn rate: n = 0.01 × (10,080 minutes ÷ 30 minutes) = 0.01 × 336 = 3.36. A sustained burn rate of about 3.36× over a 30-minute window spends 1% of a 7-day budget. That's a genuinely different rule from 14.4× over 1 hour, not the same threshold restated — the standard rule's 14.4 is the answer to a different question entirely (spending 2% of a 30-day budget in 1 hour); changing the period, the window length, or the budget-consumption target all independently change what number comes out, so two 'page' rules only mean the same thing operationally if all three of those choices match, and here none of them do.
PreviousA platform team sets head sampling at 1% to control tracing storage cost. During an incident with a 0.1% failure rate, an engineer complains that almost none of the failed requests have a trace. Is that a bug, a misconfiguration, or expected behavior of head sampling — and what would tail sampling have done differently, at what cost?Next A static alert (`error rate > 1% over 5 minutes`) paged the on-call rotation 20 times in a month, 11 of which corresponded to no real incident, and 7 of which were actually the same slow-burning incident firing and clearing repeatedly. What's wrong with the rule's design, and how does a symptom-based burn-rate alert avoid both failure modes at once?