skip to content

What must be true of a document shape change so you can safely roll the application release back after deploying it?

level: seniorimportance: should knowfreq 40%

answer

  1. Rolling back code does not roll back data
  2. Which release reads what the other wrote?
  3. Tolerance has to ship first
  4. Additive is reversible, destructive is not
  5. One step in the plan is one-way

basics

~20 s

The previous release must still be able to read documents the new release wrote. That means shipping the tolerant reader before the new writer, keeping the old fields populated while both releases can run, and delaying any destructive rewrite until rollback is off the table.

solid answer

~50 s

Rolling back a release does not roll back the documents it wrote. Those stay in the collection in the new shape, and the restored older code will meet them. So rollback safety is a property of the *write* path, decided before you deploy. Three rules. **Reader first**: deploy code that understands both generations, and only in a later release start writing the new one — the version you might roll back to must already tolerate what the new one produces. **Write both while both can run**: keep populating the fields the old release reads, so a document written by the new code is still fully interpretable by the old. **Defer the destructive step**: dropping a field, or a backfill that overwrites the old representation, is the point of no return, because it changes documents the previous release depended on and no rollback restores them. Until that last step, rollback is a deploy. After it, rollback is a restore.

code

json · 8 lines
json
// during the transition: both representations present
{
  "_id": 7,
  "schemaVersion": 2,
  "firstName": "Ada",
  "lastName": "Lovelace",
  "name": "Ada Lovelace"
}

go deeper

for a junior

Remember the core fact: rolling back code does not undo the documents that code already wrote, so the older release still has to be able to read them.

for a middle

Explain the ordering — tolerant reader in one release, new writer in the next — and why keeping the old fields populated is what lets both releases run against the same documents.

for a senior

Walk a real sequence end to end and name the single irreversible step, the checks that gate it, and how you bound the dual-write phase so it does not become permanent model debt.

for a principal

Own the rollback contract as policy: how long a rollback window stays open, who signs off on the destructive step, and how consumers outside your service are represented in that decision.

## What rollback actually means here Rolling back a deployment restores the previous binary. It does not restore the database. Every document the new release wrote during its time in production stays exactly as written, in the new shape, and the restored older code now has to read them. If it cannot, the rollback — the thing you were relying on as your safety net — is itself the outage. That reframes the question. Rollback safety is not something you evaluate during an incident; it is a property you build into the *order* of your changes, days earlier. The test to apply to any shape change before you ship it is blunt: **if I restore the previous release right now, can it correctly read every document the new release has written?** ## Reader before writer The single most important sequencing rule is that tolerance ships before production. Release N teaches the reader to handle both the old and the new generation, but keeps writing the old. It changes no data, so it is trivially rollback-safe. Release N+1 flips the writer to the new shape. If N+1 has to be rolled back, the restored N already understands what N+1 wrote. Teams that skip this and ship reader and writer in one release create a window where the only code that understands the new documents is the code they are trying to remove. Rolling back is then strictly worse than staying broken, which is a terrible position to discover at three in the morning. The same rule governs the version number itself. Decide early what an older reader does when it meets a version higher than it knows about. Failing loudly is usually right for a semantic change, and silently ignoring the unknown is right for a purely additive one — but it must be a decision made in an earlier release, not a behaviour you discover. ## Writing both representations while both releases can run While two releases could plausibly be running — during the rollout, during a canary, and for however long the rollback window stays open — a document must be readable by both. The mechanism is to keep populating the old representation alongside the new: the new code writes both `name` and the split `firstName`/`lastName`, so the old code finds what it expects and the new code finds what it prefers. This costs a little storage and a modest amount of write-path complexity, and it introduces a genuine hazard: two representations of the same fact can disagree. Two disciplines contain it. Nominate one representation as authoritative — the new one — and derive the other from it in a single function on the write path, so they cannot drift independently. And bound the period: dual writing is a phase with an owner and an end date, not a permanent feature of the model. Dual-write code that outlives its purpose is how a codebase ends up with three fields that all sort of mean the same thing. Note also that dual writing does not protect the old release from documents *created* by the new one under a new version number. If the old code branches on the version and does not know the new value, populating the old fields will not help it. That is why the reader-first release matters more than dual writing does. ## The step you cannot undo Every migration has exactly one irreversible moment: the destructive rewrite. Removing a field from documents, or running a backfill that replaces the old representation rather than adding beside it, changes data the previous release depends on. No deploy restores it; only a restore from a backup does, and that costs whatever writes happened since. So treat it as a separate, explicitly gated step, not as the tail end of the feature. The gates that make it safe are the ones you can check: the reader-first release is fully rolled out everywhere, the rollback window for the writer release has expired, a count shows no documents remain on the old generation, and no consumer outside your service — reports, exports, sibling services, cached client copies — still reads the field you are about to delete. Doing the removal in its own small release also means that if something is still reading the field, you find out from one tiny deploy rather than from a large one you also wanted to roll back for unrelated reasons. ## Keeping an escape hatch Where the transformation is lossy — splitting, truncating, reinterpreting units, collapsing a structure — consider preserving the original value under a separate field for the duration of the window, so a mistaken transformation can be recomputed from the source rather than reconstructed from a backup. Delete that holding field in the same contract step as everything else. It is cheap insurance precisely when the conversion is the part most likely to be wrong. ## The summary an interviewer wants Additive changes are reversible; destructive ones are not. Order your releases so that everything reversible happens first, keep both representations valid while two releases could be running, and put the single irreversible step last, behind checks you can actually evaluate.

  • Why does deploying the tolerant reader and the new writer in the same release make rollback dangerous?
    Because the only code that understands the new documents is the code you would be removing. Documents written during the release stay in the new shape, and the restored previous version meets a shape it was never taught to read. Splitting the change means the release you roll back to already tolerates everything the newer one produced, so the rollback is a deploy rather than a data-recovery exercise.
  • How do you keep two written representations of the same fact from drifting apart?
    Make one authoritative and derive the other from it in a single function on the write path, so no code path can update one without the other. Then bound the phase with an owner and an end date, and verify with a periodic check that samples documents and compares the two. Long-lived dual writes that nobody owns are where the disagreements start.
  • What must you confirm before running the destructive step that removes the old field?
    That the tolerant-reader release is fully rolled out, that the rollback window on the writer release has closed, that a count shows no documents on the old generation, and that no consumer outside your service still reads the field — reports, exports, sibling services and cached client copies. Ship the removal as its own small release so a surprise costs one tiny rollback.

saying these in an interview costs you the question

  • Assumes rolling back the release also reverts the documents it wrote
  • Ships the new reader and the new writer in one release
  • Drops the old field in the same change that adds the new one
  • Treats a completed backfill as reversible
  • Leaves dual-write code in place permanently with no owner

context