A load test shows p99 latency at 14ms at 80% utilisation and 110ms at 90% utilisation — twelve percent more traffic, eight times the tail latency. Is this evidence of something specific to this service, or would you expect the same shape in any queue? Explain why, using Little's Law.
This shape is not specific to this service at all — it's what queueing theory predicts for any stable queue, whether the servers are HTTP threads, database connections, or people at a supermarket till. Little's Law (L = λW) says the average number of requests in the system equals the arrival rate times the average time each one spends in the system, and it holds regardless of what's actually being queued. For a simple single-server queue, the average wait time scales roughly as utilisation ÷ (1 − utilisation) times the service time — and as utilisation approaches 1, that factor doesn't grow gently, it heads toward infinity, because the denominator is heading toward zero. That's why latency barely moves from 50% to 80% utilisation and then explodes from 80% to 95%: it's the same curve every queue produces, not a quirk of this particular service's code.