Java in a containersenior8+ years

A team migrates a service onto a base image built on `eclipse-temurin:8u181-jre` — a build that predates container awareness — inside a container with a 512 MiB memory limit, on a node with 64 GiB of physical RAM. No `-Xmx` or `-XX:MaxRAMPercentage` is set. It runs fine for days, then dies with `OOMKilled`. Why does it take days rather than failing immediately, and what's actually different from the same scenario on a current JDK?

Before container awareness, the JVM sized its default heap by reading /proc/meminfo — the host's total physical memory — because that was the only notion of "the machine" a userspace process had, and it was a correct assumption everywhere except inside a container with a tighter cgroup limit than the host actually has. On a 64 GiB host, a pre-aware JVM's default heap comes out on the order of 16 GiB (roughly a quarter of what it read), entirely independent of the container's real 512 MiB ceiling. It doesn't fail immediately because the heap grows lazily as the application actually allocates objects, not all at once at startup — the process runs normally for as long as its real memory usage stays under 512 MiB, which can be days of ordinary traffic before enough has accumulated to cross the cgroup's actual limit. The moment it does, the kernel's cgroup enforcement kills the process outright with no Java exception and nothing in the application log, because the JVM was never told a tighter limit than the host's existed in the first place. This became the default fix around Java 10, backported to 8u191 — worth hedging the exact point-release precision on, since that detail is easy to misstate — after which the JVM reads the cgroup's actual limit and sizes against it instead.

The lesson behind it →
More on Java in a container