Consistency and reliabilitysenior8+ years

A cache entry costing 300 ms to rebuild is protected from a stampede by probabilistic early expiration (XFetch) instead of a Redis rebuild lock. Walk through the actual formula being evaluated on every read, and explain precisely why a read close to the entry's expiry is far more likely to trigger an early rebuild than one early in the entry's life — without the formula needing to know explicitly how much time is left.

The check, run on every read, shifts the current time forward by a random amount before comparing it to the entry's real expiry time: now − (delta × beta × ln(random())), where delta is how long the last rebuild took, beta is a tuning knob usually left at 1, and random() draws a fresh value between 0 and 1 on every single read. Because the comparison target — the real expiry time — is fixed, a read that happens close to expiry only needs a small random shift to cross it, while a read far from expiry needs a large one, which almost never happens; the formula never explicitly checks "how much time is left" at all, that behaviour falls out naturally from comparing an ever-larger random shift against a fixed, unmoving target. delta scaling the shift means an expensive rebuild gets nudged to start earlier than a cheap one, purely because its typical shift is larger.

The lesson behind it →
More on Consistency and reliability