Flaky testssenior8+ years

A CI guard asserts that a practice problem's deliberately broken starter code (an unsynchronised `count++` across a thread pool) **fails** its tests. It passed five deploys in a day, then failed a deploy that changed only markdown, reporting that the starter passes every case. Locally it passes six out of six runs. How do you handle this, what mechanism explains it, and under what conditions would you accept a retry?

Diagnose, do not re-run. First, the change could not be the cause: it touched no Java. Second, it does not reproduce locally, so the difference is the environment. Third, find a mechanism that explains both: a lost update needs one thread to be preempted between the read and the write of count++. On a laptop with real parallelism that happens constantly, so the broken code fails as intended. On a small CI VM the pool's threads rarely overlap, and the JIT may compile a loop over a non-volatile field into something close to one read, one add and one write, so the broken code produces the correct total and the guard complains. Only then decide: this assertion is genuinely probabilistic (a race may legitimately not show in a run), so a bounded retry on the failing path, without weakening the assertion and written down, is honest here and nowhere else.

The lesson behind it →