Capacity estimationmedium0-2 years

A design doc says "we handle 500 requests/second at peak, so a thread pool of 50 threads is plenty of headroom." The doc doesn't mention request duration. What's missing from that sizing argument, and how do you fix it?

Requests per second alone doesn't say how many requests are in flight at once — that also depends on how long each one takes. Little's Law, L = λ × W, gives the actual number: λ is the arrival rate (500/s here), W is the average time a request spends in the system, and L is the average number of concurrent in-flight requests, which is what a thread pool actually needs to hold. Fifty threads is exactly enough only if the average request takes 100ms (50 = 500 × 0.1). If a request actually takes 300ms — plausible with a database write plus a downstream call — L becomes 150, three times the pool size, and the excess queues up and the service falls over on latency, not CPU.

The lesson behind it →