A service adds a 100ms timeout to a call that used to hang for two seconds under fault injection, and the chaos experiment confirms p50 latency is back to normal. A teammate says the resilience story is done. What does a bare timeout still leave unaddressed once the dependency has a sustained outage, not just a transient fault, and what does a circuit breaker add?
A timeout answers "how long does one call wait before giving up" — it stops a single hung call from holding a thread forever, which is exactly what the lesson's experiment proves. But if the dependency stays down for ten minutes, every single request in that window still pays the full 100ms timeout before failing, because nothing remembers that the last hundred calls all failed the same way. A circuit breaker adds that memory: once it sees enough recent failures, it trips open and stops calling the dependency at all, failing immediately with no wait — which is what actually saves the accumulated cost of "100ms times every request, times ten minutes" that a bare timeout keeps paying. It then periodically lets a few trial calls through (half-open) to find out, safely, whether the dependency has recovered, rather than either staying open forever or slamming it with full traffic the instant it looks better.