Capacity planningeasy0-2 years

A load test claims p99 latency of 50ms at the target throughput, using a tool that waits for each response before sending the next request. Production, at the same throughput, sees p99 latencies over a second during a brief slowdown. What's wrong with the load test's methodology?

This is coordinated omission: a load generator that waits for each response before sending the next one automatically slows itself down exactly when the system it's testing slows down, which means it stops sending requests during the very window it should be measuring most carefully. Real users don't behave that way — they keep arriving at whatever rate they arrive at, whether or not the system is currently struggling, so a slowdown in production means requests pile up and wait, which is exactly the tail latency a load test needs to capture and this kind of tool structurally cannot. The fix is a load generator that sends at a fixed rate regardless of how fast responses come back — an "open model" — which tools like wrk2, Gatling and k6 support specifically for this reason.

The lesson behind it →
More on Capacity planning