A home page built from a 300 ms query is cached for a minute. When its entry expires, 200 concurrent requests all miss the cache at once and all run the expensive query. Walk through why a per-instance single-flight lock only partly fixes this across four application instances, and why a Redis-based rebuild lock introduced a second bug on top of the first fix.
A per-instance lock (a lock inside one JVM) only coordinates threads within that one process — across four application instances, each instance's own lock lets through one rebuild, so four instances means four database loads instead of 200, a real improvement but not "exactly one." A Redis-based rebuild lock coordinates across every instance, correctly bringing it down to one — but the lesson's own experiment hit a second, more subtle bug: a waiting request can check the cache right before the rebuilder finishes writing the new value, find it still empty, then acquire the now-released lock and rebuild a second time. The fix for that is checking the cache again after acquiring the lock — the same double-checked-locking pattern used in ordinary concurrent Java code.