Distributed systemshard5-8 years

Three service instances coordinate through a Redis lock (`SET lock:invoice-1041 <owner> NX PX 1000`) before sending an invoice, so only one should send it. Worker A acquires the lock, then a GC pause holds it inside its critical section for 1.5 seconds — longer than the lock's 1-second expiry. Walk through what actually happens, and what a plain `DEL` on release makes worse.

Once worker A's pause outlasts the lock's 1-second expiry, the lock is simply gone from Redis's point of view — nothing tells worker A this happened, because nothing can reach into a paused process. Worker B, watching the key become free, acquires it legitimately and starts its own work believing it's the sole holder, which by that point it correctly is. When worker A wakes up, it still believes it holds the lock, finishes its work, and calls a plain DEL — which doesn't check whose lock it's deleting, so it deletes B's currently-valid lock, letting a third worker in while B is still mid-flight. The fix for that specific danger is to store a unique value at acquire time and release only if the stored value still matches — atomically, since even a GET followed by a DEL is itself a race — but that only stops one lock holder from deleting another's lock; it does nothing about worker A still writing its own stale result after the pause, which is a separate, deeper problem.

The lesson behind it →
More on Distributed systems