A payment provider's latency rose from 200ms to 6s for thirty seconds during their own failover. The checkout service had a 10-second read timeout and three retries with no backoff. Walk through why this became a two-hour outage, and name every configuration mistake.
Each checkout request could hold a worker thread for up to roughly 30–40 seconds (a 10-second read timeout, up to three attempts, with no backoff spacing them out). The 200-thread pool filled within a minute, so every checkout — including ones with nothing to do with payments — queued and timed out at the load balancer too. The clients (web and mobile) then retried their own failed requests immediately, three times each, with no jitter. By the time the provider's own failover finished at second 30, it faced roughly nine times its normal request rate arriving all at once, went into its own overload, and started shedding load with 503s — which checkout retried. The two systems held each other down for two hours. Every piece was a configuration mistake: a read timeout ten times the healthy p99, no backoff or jitter, retries at more than one layer, and no circuit breaker to stop calling a dependency that was clearly unhealthy.