Discoveryhard5-8 years
A caller uses a keep-alive HTTP client against a Kubernetes Service with five ready pods behind it, and one pod ends up receiving 80% of the traffic. Why, mechanically, and what fixes it at each layer?
Kubernetes Services load-balance per TCP connection, not per request: kube-proxy's rule picks a pod at random only when a new connection is opened, so a client that keeps a handful of connections alive keeps sending most of its traffic through whichever pod those connections happened to land on. Fixes exist at two layers: at the client, shorten idle-connection lifetime so connections churn and get rebalanced periodically (at some handshake cost); at the platform, either a service-mesh sidecar that balances per request instead of per connection, or a headless Service that hands the client the pod IPs directly so it can balance client-side across them itself.
PreviousAn orders service calls a slow, non-critical reporting service and a fast, critical payments service, both from the same shared thread pool. Why does a bulkhead around each matter, and which Resilience4j bulkhead type — semaphore or thread-pool — fits which?Next A payment provider's latency rose from 200ms to 6s for thirty seconds during their own failover. The checkout service had a 10-second read timeout and three retries with no backoff. Walk through why this became a two-hour outage, and name every configuration mistake.