Scalinghard5-8 years

A Deployment runs 4 replicas with a CPU request of 500m each, targeting 60% average utilisation. The metrics API reports the pods averaging 82%. What does the HPA compute as the desired replica count, and why does `kubectl get hpa` sometimes show a target below the configured one for longer than seems reasonable?

Plugging the numbers into the HPA's own formula — ceil(currentReplicas × currentMetricValue / desiredMetricValue) — gives ceil(4 × 82/60) = ceil(5.47) = 6 replicas. The apparent lag people notice isn't a stuck controller: the HPA polls on a fixed sync period (15 seconds by default, not continuous), ignores small deviations inside a tolerance band (10% by default) so it doesn't chase noise, and — when scaling behavior windows are configured — looks back over recent recommendations and picks the most conservative one from that window rather than acting on the latest reading alone. A target that looks stuck below the configured value for longer than expected is almost always one of those three deliberate delays doing its job, not a controller that's failed to notice a real change.

The lesson behind it →