How would you prove a schema change's recovery path works before an incident forces you to use it?
answer
- the step that never runs until it must
- apply, write rows, undo, re-apply
- assert data, not only shape
- corrective units must be idempotent
- old build against recovered schema
basics
~20 sExecute it in a test rather than assuming it: apply the change, write rows through the new shape, run the undo or corrective unit, then assert both the schema and the surviving data. A recovery step that has never run is untested code.
solid answer
~50 sThe forward path is exercised by every environment; the recovery path is usually written once and first run during an incident. Rehearse it instead. For an undo step, loop in the test job: apply the chain to head, write representative rows, run the undo, assert the schema returned to its previous shape, then re-apply to prove the pair is repeatable. That immediately exposes undo steps that were auto-generated and never parsed, are empty, or only work on an empty table. For a corrective forward unit, start from a copy that already has the bad change and malformed rows, apply the fix, and assert malformed rows were repaired, good rows untouched, and running it twice is safe. Assert on data, not just shape — an undo that drops a column discards everything written to it — and check the previous application build still runs against the recovered schema.
go deeper
Recognise that a recovery step is code like any other: if it has never been executed, nobody knows whether it works. The test can simply run it and check the schema afterwards.
Explain the apply, write rows, undo, re-apply loop and what each stage catches, and why an undo tested on an empty table proves much less than one tested with rows present.
Show that you assert on surviving data, not only shape, that corrective units are checked for idempotence, and that the previous application build is booted against the recovered schema before you call the path proven.
Decide which changes get a rehearsed undo at all and record the answer. 'No undo, forward fix only' is a legitimate position; discovering it mid-incident is not, so make the recovery cost explicit alongside the change.
## The step nobody runs until it is an emergency A change chain's forward path is exercised constantly: every developer machine, every test run, every environment applies it. The recovery path — whatever you would do if the change turns out to be wrong — is typically written once, never executed, and first invoked by a stressed person during an incident. That makes it the least trustworthy code in the deployment, and it is trivially testable in advance. Two shapes of recovery exist. One is an **undo step** attached to the change, which returns the schema to its previous state. The other is a **forward fix**: leave the bad change in place and ship a new change unit that corrects it. Which of the two your team relies on is a policy decision made elsewhere; this is about proving that whichever one you rely on actually executes. ## Rehearsing an undo step The mechanical rehearsal is a loop, run in the same job that replays the chain: 1. Apply the chain up to and including the new change. 2. Write rows through the new shape, so the database is not empty when recovery runs. 3. Run the undo step. 4. Assert the schema is back to the previous shape — the added column is gone, the tightened constraint is loosened, the renamed object answers to its old name. 5. Apply the change again, to prove the pair is repeatable rather than a one-way trip. Three defects fall out of this loop almost immediately: an undo step that was auto-generated and never syntax-checked; an undo step that is empty because the tooling could not infer one; and an undo that only works on an empty table, because it drops a column that now has a not-null dependent or violates a constraint that new rows already satisfy. ## Assert on data, not only on shape A schema-shape assertion is the easy half and the less important one. The question that matters during an incident is what happens to the rows written since the bad change went out. | Change | What the undo does to data | |---|---| | Added a column | Dropping it discards everything written to it — irreversible | | Widened a type | Narrowing it back can fail or truncate values already stored | | Split one column into two | The reverse must merge them, and the merge rule is business logic | | Added a table | Dropping it discards its rows | | Tightened a constraint | Loosening it is safe; the rows already conform | So the rehearsal should write representative rows before recovering and then assert what survived. If the honest answer is "the last twenty minutes of writes are gone", that is a finding worth having on paper **before** the incident, not a discovery made at 3 a.m. ## Rehearsing a forward fix A forward fix is rehearsed differently, because its starting point is not a clean database — it is a database that already has the bad change applied and rows written under it. The rehearsal is: 1. Restore a snapshot, or replay the chain, to the exact state the production database is in, including the bad change. 2. Generate rows in the states the defect produces, including the malformed ones. 3. Apply the corrective unit. 4. Assert the malformed rows were repaired and the well-formed ones were untouched, and that the unit is safe to run twice — corrective units are re-run more often than any other kind. Idempotence is worth singling out. A repair statement that is safe to apply once and doubles a value on the second run is a worse incident than the one it was written to end. ## What the rehearsal still cannot tell you - **Time under real volume.** An undo of a large change can be as slow as the change; if you never rehearse the recovery against a realistically sized copy, you know it is correct but not that it is usable inside an incident window. - **The version straddle.** Recovery is rarely schema-only: the previous build of the application has to run against the recovered schema. Add that to the rehearsal — boot the previous build against the post-recovery schema and run its data-access tests — or you will restore a schema no deployed code can serve. - **Anything already read by downstream consumers.** Undoing a schema change does not un-send the messages, exports or cached views produced while it was live. ## Making it routine Keep the apply-undo-reapply loop in the ordinary test job for every change that ships an undo step, so the recovery path is exercised on every commit rather than annually. Keep the snapshot-based forward-fix rehearsal for the changes that touch data, because that is where the correctness question lives. And write down, next to the change, which recovery it supports and what it costs — a change whose honest answer is "no undo, forward fix only" is fine, as long as that is a decision on record rather than a discovery.
- Why does the rehearsal insist on writing rows before running the undo?Because an undo against an empty table is the easy case. Rows are what make it fail or lose data: dropping a column discards what was written to it, narrowing a type can truncate stored values, and loosening or dropping a constraint behaves differently once rows depend on it.
- What does an apply-undo-reapply loop catch that a one-way undo test does not?That the pair is repeatable. A re-apply after undo fails when the undo left residue — a leftover index, a sequence not reset, a partially dropped object — which is exactly the state an incident produces when someone recovers and then retries the fixed change.
- Why must a corrective forward unit be safe to run twice?Because corrective units are re-run more often than any other kind: an incident interrupts them, someone reruns the deploy, or the same unit ships to several databases at different states. A repair that doubles a value or re-applies an offset on the second pass turns one incident into a worse one.
- The undo restores the schema correctly, but the service still fails after recovery. What was not rehearsed?The version straddle. Recovery is rarely schema-only — the previously deployed application build has to run against the recovered schema. Include booting that build and running its data-access tests against the post-recovery schema, otherwise you restore a schema no deployed code can serve.
saying these in an interview costs you the question
- The undo step was generated automatically, so it must work.
- Undoing a schema change puts the data back the way it was.
- Asserting the schema shape after the undo is enough.
- Corrective units run once, so idempotence does not matter.
- Recovery only concerns the database, not the deployed application build.