nginxhard5-8 years

During a traffic spike, users report a mix of errors: some requests get an instant failure, others hang for exactly 60 seconds before failing. The on-call engineer sees both 502s and 504s in the nginx access log at the same time. Walk through what's different about these two failure modes and where you'd look for each.

A 502 means nginx could not get a valid response from the JVM at all — the upstream wasn't listening, refused the connection outright, closed it mid-response, or sent back something nginx couldn't parse — and that failure is essentially instant, because nginx doesn't wait around for a response that was never coming. A 504 means the opposite: nginx successfully reached the JVM, which accepted the request and is alive, but didn't answer within proxy_read_timeout (60 seconds by default) — the JVM is up and working, just too slow, which is why those requests hang for the full timeout before failing rather than failing instantly. Seeing both at the same time during a spike is consistent with a JVM under real load: some requests are timing out because the JVM is genuinely too slow to answer within 60 seconds (504), while others are hitting a JVM that has, for a different reason — maybe the OOM killer, maybe a restart, maybe connection saturation — actually gone away or refused the connection outright (502). The 502s point at systemctl status, ss -tlnp | grep 8082 and the kernel log; the 504s point at what the JVM is doing while it's alive but slow — a thread dump, GC pauses, a slow query.

The lesson behind it →
More on nginx