A rolling deploy ships a migration that renames `users.email` to `users.email_address` in the same release as the code that reads the new name. Midway through the roll, requests start failing. Why, and what's the actual fix — not "don't rename columns," but the mechanism that makes a rename safe under a rolling deploy?
A rolling deploy means, for the entire duration of the roll, some instances are running the old code and some the new — and both sets of instances share the same database at the same moment. Renaming a column in the same deploy as the code that expects the new name breaks that: the moment the migration runs, the old code (still serving traffic on instances not yet replaced) tries to read users.email, which no longer exists, and every request hitting one of those old instances fails, for as long as the roll takes. The fix is expand-contract: split the change into three separate deploys spread over time — first expand (add the new column, both old and new code coexist, the old one just ignores the new column), then migrate (deploy code that writes both columns and reads the new one, backfill the old data in batches), then contract, days later after the rollback window has closed, remove the old column once nothing reads it anymore. At every single deploy in that sequence, both the old and new code that could be running at that moment can read and write correctly — nothing is ever both required and missing at the same time.