On-call and incidentsmedium3-5 years

A deploy fails with `deploy failed before the switch — removing /srv/code10x/releases/20260912-204250`, and the on-call engineer's first instinct is to roll back. Why would that be the wrong move here, and what does a good runbook do to prevent it?

Rolling back here would be actively harmful, not just unnecessary: the message says the deploy failed before the switch, meaning the live site was never touched and is still serving the previous, working release — there is nothing to roll back from. Rolling back in this state means switching to some other, potentially older release for no reason, on a site that was never broken. The instinct to "do something" during a failure is natural, but the runbook's job is precisely to stop the wrong action as much as to prompt the right one — a good runbook entry states explicitly, in the SEVERITY or CHECK line, that nothing was switched and no user impact occurred, so a tired engineer at 3am reads that line before reaching for the rollback command rather than after.

The lesson behind it →