Rate limitinghard5-8 years
A caller runs 12 instances, each configured with a Resilience4j `RateLimiter` of 50 requests/second, to stay under a dependency's 600 req/s capacity. The caller then autoscales to 20 instances under load. What breaks, and what's the actual fix?
Resilience4j's RateLimiter is local to each instance, so "12 instances × 50/s = 600" was only correct by coincidence of the instance count at the time it was set. At 20 instances the same per-instance limit allows 20 × 50 = 1,000 requests/second — well past the dependency's 600 req/s capacity — with nobody having touched the rate-limiter configuration at all. The real fix is a shared bucket that holds at 600 regardless of how many instances are running, typically via Bucket4j or Spring Cloud Gateway's Redis-backed limiter, or deriving the per-instance number dynamically from the current replica count if a shared store isn't available.
PreviousA payment provider's latency rose from 200ms to 6s for thirty seconds during their own failover. The checkout service had a 10-second read timeout and three retries with no backoff. Walk through why this became a two-hour outage, and name every configuration mistake.Next Three years into a platform, the shared API gateway has sixty-one custom filters, one of which runs a blocking JDBC query against the orders table to check whether an order is cancellable before proxying the request. What's wrong here beyond "it's slow", and how should this actually be structured?