Failure detectionhard5-8 years

Two nodes have identical 2,200ms silences: node A's heartbeats normally arrive every 1,000ms ± 20ms, node B's arrive every 1,000ms ± 400ms. A fixed 2-second timeout has already fired identically for both. Explain exactly what a phi-accrual failure detector computes differently for them, and why treating them the same would be the actual bug.

A fixed timeout has exactly one bit of output — "not yet suspect" or "suspect" — so at 2,200ms of silence it has to treat both nodes identically, even though the same silence means very different things for each: it's an almost-unprecedented gap for A's tight, predictable history and an entirely ordinary late arrival for B's wide, jittery one. Phi accrual replaces that one bit with a continuous number, φ, computed from each node's own recent heartbeat-interval history: given how long the current silence has lasted, it estimates the probability that a heartbeat would still arrive this late under that node's normal pattern, and reports φ as roughly "how many powers of ten unlikely is this silence" — each unit of φ is a tenfold drop in that probability. At 2,200ms, A's φ rises fast (this silence is many standard deviations outside A's history) while B's rises slowly (B has plenty of gaps this size in its own history already) — which is exactly the correct, adaptive behaviour: A's silence is stronger evidence of a real failure, B's is comparatively weak evidence the network is merely doing what it always does.

The lesson behind it →