How do you verify that saved user data survives both an upgrade and a rollback to the previous release?
answer
- Two directions, not one
- Rollback is the untested direction
- A corpus written by old releases
- Make the new build write before rolling back
- Version stamp plus a rule for unknown content
basics
~20 sKeep saved data written by every supported release, upgrade each one and assert the content afterwards, then reverse it: install the previous release over the upgraded data and check the older build reads it without discarding what it cannot interpret.
solid answer
~50 sTwo properties are being tested and they are not the same. **Backward compatibility** is the new build reading data written by an old one — covered by keeping saved files and stores produced by each supported release and upgrading every one, asserting content rather than just "it started". **Forward compatibility** is the old build reading data written by the new one, and it is what makes a rollback survivable: if the previous release drops fields it does not recognise, or refuses to open a newer version stamp, a rollback destroys data instead of restoring service. So the drill runs both directions on real saved data, including the skipped-version paths people take. Add a version stamp in the stored data, tolerance for unknown content, migrations that are idempotent and resumable after a kill, and a checked pre-migration copy. Time the migration on the largest realistic dataset — a slow upgrade is an outage.
go deeper
Be ready to explain in plain terms that a new version has to open data saved by an old one, and that testing it needs real saved data from the older version rather than data you just created.
An interviewer expects both directions named and distinguished, the idea of a version stamp in the stored data, and the upgrade paths worth covering — previous, oldest supported, and a skipped-version jump.
Describe the drill you would actually run, including making the new build write before rolling back, asserting content rather than start-up, timing the migration, and handling an interrupted migration.
Own the release policy: the transition window where both builds tolerate the format, add-tolerate-require-remove sequencing, what the rollback guarantee is, and how snapshot age and migration duration are budgeted as release criteria.
### The two directions Stored data outlives the build that wrote it, so a release has to be checked in both directions. - **Backward compatibility** — the *new* build opens data written by an *older* one. This is the upgrade path, and most teams test at least the happy version of it. - **Forward compatibility** — the *older* build opens data written by the *newer* one. This is the rollback path, and it is the one that gets skipped, because nobody plans to roll back. Yet rollback is the standard incident response, and a rollback that loses data is worse than the incident. When an interviewer asks about upgrade testing, naming the second direction unprompted is most of the answer. ### Building the corpus The hardest input is realistic saved data. A file written by the current build with a script is not evidence, because it lacks everything old builds used to produce. Keep a **versioned corpus**: for every release still in the support matrix, saved files, local databases and preference stores produced by that release doing real work — including the ugly cases: a half-finished record, a store left behind by a crash, data written in a non-default locale, a store at realistic size rather than three rows. Freeze each artefact and keep it under version control with the release that produced it. When a release drops out of the matrix, its corpus entry retires with it. ### The upgrade matrix Users do not upgrade one version at a time. The paths worth running are: the immediately previous release; the oldest release still supported; at least one skipped-version jump; and an upgrade interrupted partway and then resumed. For each, assert **content** — record counts, specific field values, referential integrity, ordering, and anything derived — not merely that the application starts. A migration that silently drops a column passes a start-up check every time. Also time it. Migration duration on the largest realistic dataset is a release-gating number, because during a migration the product is unavailable. ### The rollback drill The rollback drill is the one to describe in detail: 1. Take a snapshot of the stored data before the upgrade. 2. Upgrade, then exercise the product so the new build **writes** — this is essential, because forward compatibility is only tested against data the new build produced. 3. Install the previous release over that data. 4. Assert what the old build does: does it open at all, does it discard fields it does not recognise, does it round-trip data it cannot interpret, does it refuse on a version stamp? The usual designs that make step 4 survivable are a **version stamp** in the stored data plus a rule about what an older reader does with unknown content — either preserve it untouched on rewrite, or refuse to write at all and stay read-only. Both are defensible; silently dropping it is not. Where the new format genuinely cannot be read by the old build, the fallback is the pre-upgrade snapshot, and then the test is that restoring it actually works and its age is acceptable. ### A worked example A hotel booking channel manager ships an on-premises connector that keeps a local store of rate plans and pending updates, peaking near 1,200 requests per minute during promotions. Release N added a per-plan restriction field to the store. A property hit an unrelated problem and rolled back to N-1 the same evening; N-1 read the store, ignored the restriction field, and rewrote every record without it. The connector then pushed the pre-existing restriction-free rates from its now-authoritative local store — a **stale-cache read** promoted into a write — and 41 rooms went on sale at an unrestricted rate for about 3 hours 20 minutes. Nothing about that was a code defect in the usual sense: the upgrade worked, the rollback worked, and the data was destroyed between them. The corrective work was a preserve-unknown-content rule for the store's reader, a version stamp read by both builds, a pre-upgrade snapshot taken automatically, and a rollback drill added to the release checklist that exercises exactly the sequence above. ### What to say about coordination When the same data is written by more than one participant during a rolling upgrade, the honest answer includes a transition window: for one release both builds must be able to read and write the format, so a rollback never meets data it cannot handle. That means format changes land at least one release before the code that depends on them — add-and-tolerate first, require later, remove later still. It is slower, and it is what makes rollback a routine action instead of a gamble.
- How would you generate the corpus of old saved data if nobody kept any?Reconstruct it: check out each release still in the matrix, run it in a clean environment through the flows that write — including partial and interrupted ones — and freeze the resulting stores as fixtures alongside the release tag. Supplement with anonymised copies of real customer data where policy allows, because real data carries shapes nobody thinks to create. Then keep the practice going by capturing a fresh artefact from every release as it ships.
- What makes a data migration safe to retry after it is killed halfway?Idempotence and recorded progress. The migration should be decomposed into steps that can each be applied twice without changing the result, with a persisted marker of the last completed step so a rerun resumes rather than restarts. Combine that with a pre-migration copy and an explicit failure path that leaves the store in the old shape rather than a half-new one, and a killed run becomes a retry instead of a restore.
- When is refusing to open newer data the right behaviour for an older build?When silently continuing risks a wrong write. If the older build cannot represent what it is reading, preserving unknown content untouched is best; if it cannot even do that, refusing and telling the operator to restore the snapshot is far better than rewriting the store into a lossy shape. The decision turns on blast radius: a read-only client can afford to degrade, a component whose local store is authoritative for anything downstream cannot.
It is the difference between checking that a new reader opens last year's documents and checking that last year's reader opens today's — only the second one tells you whether retreating is safe.
saying these in an interview costs you the question
- Tests only the upgrade direction and never the rollback
- Asserts the application starts rather than checking content
- Builds old data with the current build's own writer
- Assumes users upgrade one version at a time
- Never times the migration on realistic data volume
- Lets an older reader silently drop fields it does not recognise