Testcontainershard5-8 years

A shared CI runner keeps running out of disk, and `docker ps -a` shows hundreds of old `postgres:16-alpine` containers. The pipeline sets `TESTCONTAINERS_RYUK_DISABLED=true` because 'Ryuk kept failing on this runner'. Explain how Testcontainers talks to Docker, what Ryuk does and why it is not a shutdown hook, and how you would fix this properly.

Testcontainers does not call the docker CLI; it speaks the Docker Engine HTTP API over DOCKER_HOST or the default socket. When the first container is requested, it also starts Ryuk, a small container of its own, and holds a connection to it for the life of the test JVM. When that connection drops, Ryuk removes every container, network and volume labelled with that session's id. Because it is a separate process watching a connection, it works even after kill -9 or a CI timeout, which skip JVM shutdown hooks. Disabling it means every run that does not stop its containers cleanly (timeouts, crashes, cancelled jobs, containers left to 'end with the JVM') leaves them behind, which is the disk problem. The fix is to make Ryuk work on the runner, not to turn it off.

The lesson behind it →
More on Testcontainers