Rate limitinghard5-8 years

A caller runs 12 instances, each configured with a Resilience4j `RateLimiter` of 50 requests/second, to stay under a dependency's 600 req/s capacity. The caller then autoscales to 20 instances under load. What breaks, and what's the actual fix?

Resilience4j's RateLimiter is local to each instance, so "12 instances × 50/s = 600" was only correct by coincidence of the instance count at the time it was set. At 20 instances the same per-instance limit allows 20 × 50 = 1,000 requests/second — well past the dependency's 600 req/s capacity — with nobody having touched the rate-limiter configuration at all. The real fix is a shared bucket that holds at 600 regardless of how many instances are running, typically via Bucket4j or Spring Cloud Gateway's Redis-backed limiter, or deriving the per-instance number dynamically from the current replica count if a shared store isn't available.

The lesson behind it →