Load balancingmedium3-5 years

Four identical-looking servers behind a load balancer; one is having a bad day (GC storm) and answers three times slower than the others. Why does round robin make this dramatically worse than it needs to be, and what actually fixes it?

Round robin sends every server an equal share of requests regardless of how fast it's actually answering — the slow server still gets a full quarter of the traffic, its queue backs up because it can't clear requests as fast as they arrive, and the p99 for requests that happen to land there blows out (measured: 556ms, against 6ms p50). A load-aware algorithm fixes it by routing around the backlog instead of ignoring it. Least outstanding requests sends each request to whichever server currently has the fewest in-flight requests, and it noticed the slow server backing up and cut its share to under 6% — a 17× improvement in p99 with no change to any server. Power of two choices gets most of that benefit by checking only two random servers per request, which matters when many load balancers each see only part of the picture.

The lesson behind it →