Configuration and probesmedium3-5 years

A teammate wires the Postgres connection check into `/actuator/health/liveness` so "the pod restarts if the database is unreachable." The database has a rough five minutes one afternoon. What happens to the service, and what should the probe wiring actually look like?

It turns a database incident into a restart storm on top of it. Liveness failing means the kubelet restarts the container — the pod is torn down and started fresh — which does nothing to fix an unreachable database and instead throws away every in-flight request the pod was actually handling, cycles the JVM's warmup, and does this repeatedly for as long as the database stays down, since a fresh container immediately fails the same liveness check again. What should have happened is readiness failing, not liveness: readiness removes the pod from the Service's endpoint list so no new traffic reaches it while the dependency is down, without touching the running process at all, and the pod goes straight back to serving traffic the moment the check passes again — no restart, no lost state, no warmup. The fix is putting the database check behind /actuator/health/readiness and keeping liveness limited to whether the JVM itself is wedged.

The lesson behind it →
More on Configuration and probes