A platform team sets head sampling at 1% to control tracing storage cost. During an incident with a 0.1% failure rate, an engineer complains that almost none of the failed requests have a trace. Is that a bug, a misconfiguration, or expected behavior of head sampling — and what would tail sampling have done differently, at what cost?
It's expected behavior, not a bug — head sampling decides whether to keep a trace at the very start of the request, before anything about that request's outcome is known, so a 1% keep rate applies uniformly regardless of whether the request will eventually succeed or fail. With a 0.1% failure rate, failed requests are rare to begin with, and a uniform 1% sample of all traffic keeps roughly 1% of those failures too — a tiny fraction, exactly matching the lesson's own measured result of 13 kept traces out of 1,003 actual failures at 1% head sampling. Tail sampling decides after the trace completes, so it can keep every trace that had an error and every trace slower than a threshold regardless of the overall sampling rate, at the operational cost of buffering every in-flight trace's spans until each one finishes — a collector tier sized for your traffic and your longest trace, rather than a stateless per-request coin flip.