Resiliencesenior8+ years

A cache node restarts. Every request now misses, hits the database, the database slows under the new load, requests time out, clients retry, and the cache can never refill because the database is too slow to answer the fills either. An hour later, restarting the cache node again doesn't help — the system stays overloaded at a load level it handled fine yesterday. What's actually going on, and why doesn't removing the original trigger fix it?

This is a metastable state: a system that's stable under normal load enters, through a trigger (here, the cache restart), a state where a separate sustaining effect keeps it overloaded even after the trigger itself is long gone — in this case, the retries. The cache restart was a single, momentary event; the reason the system is still overloaded an hour later has nothing to do with that event anymore and everything to do with the feedback loop it kicked off: every timed-out request gets retried, every retry adds load to an already-overloaded database, the added load makes the database slower, slower responses cause more timeouts, and more timeouts produce more retries. Restarting the cache node again doesn't fix anything because the cache was never the ongoing problem — the retry loop is, and it's self-sustaining independent of the cache's state at this point. The fix is always to break the feedback loop directly: stop or throttle the retries, shed load, warm the cache from a replica before letting traffic back in, or restart with a fraction of normal traffic — not to keep re-addressing the original trigger.

The lesson behind it →
More on Resilience