skip to content

Schemaless Migrations & Document Versioning

Without ALTER TABLE, changing shape means every collection holds several generations of document at once. Interviewers ask how you ship such a change with zero downtime — the good answer names a version field and a lazy migration path.

on this pageshow

questions

5

In a document database, why do teams store a schemaVersion field in each document, and how does the read path use it?

level: middleimportance: must knowfreq 65%

answer

  1. No ALTER TABLE, so old shapes survive
  2. Several generations live in one collection
  3. Read code must know which one it holds
  4. One small integer written by every writer
  5. Branch once, in an upcasting layer

basics

~20 s

It records which shape generation a document was written with. With no ALTER TABLE, one collection holds several generations at once, so the stored version lets read code branch deterministically instead of inferring shape from which fields happen to be present.

solid answer

~50 s

A document store has no schema-wide rewrite: writing a new shape changes one document and leaves every older document exactly as it was. So a collection permanently holds several generations, and the application is the component that has to reconcile them. A small integer field — commonly `schemaVersion` — written by whoever creates or rewrites the document says which generation it is. The read path uses it in one place: an upcasting layer at the boundary where raw documents become domain objects, applying v1→v2, then v2→v3 transformations until the object matches what current code expects. Everything above that layer sees one shape and never checks a version. The alternative, sniffing which fields are present, breaks as soon as a field is legitimately optional, and it gives you no way to count how many documents are still on the old shape — which is exactly the number that tells you when the migration is finished.

code

json · 3 lines
json
// two generations coexisting in one collection
{ "_id": 1, "schemaVersion": 1, "name": "Ada Lovelace" }
{ "_id": 2, "schemaVersion": 2, "firstName": "Ada", "lastName": "Lovelace" }

go deeper

for a junior

Be ready to say that a document database cannot rewrite every record at once, so old and new document shapes coexist, and that a stored version number is how code tells them apart.

for a middle

Explain the mechanics: who writes the field, what a single upcasting layer does with it, and why sniffing for a field's presence fails once that field is optional. Know what does and does not deserve a version bump.

for a senior

Show you have operated this. Name the write paths that must set the field, what happens when a wholesale document replace drops it, and how counting documents per generation is what eventually lets you delete branch code.

for a principal

Own the convention across services: who may write the collection, what the policy is when a reader meets a version it does not know, and how you keep the number of supported generations bounded rather than accumulating forever.

## Why a version field exists at all In a relational database, changing a table's shape is one privileged operation. The engine rewrites or re-describes every row, and once it commits there is exactly one shape in the table. A document store offers no equivalent. Writing a document with a new shape changes that document and nothing else. Every document written before the change keeps its old shape until something rewrites it — which, for records nobody touches again, may be never. That is the fact interviewers are probing: **a collection holding several generations of document simultaneously is the normal steady state, not a temporary glitch during a deploy.** Since the store will not reconcile the shapes, the application must, and to do that it needs a reliable answer to one question: which generation is this document? ## What the field is A version field is a small integer stored on every document, set by the writer to the generation of the shape it wrote. Generation 1 is the shape at the moment you introduce versioning; each subsequent breaking change increments it. It is an attribute of the document's shape, not of the application release, not of the record's business state, and not a timestamp. Storing it costs a few bytes per document and one line in the write path, and it pays for itself the first time you need to answer "how much of this collection is still on the old shape?" — a question you cannot answer at all without it. ## Why not just look at the data The tempting alternative is to infer the generation: if `lastName` is present, treat the document as the new shape. This holds only while every changed field is mandatory. The moment a field is legitimately optional, a new document that simply omits it is indistinguishable from an old document that never had it, and the read path silently picks the wrong branch. With three or four generations the presence tests start interacting, and a chain of heuristics grows in every place that reads the collection. Inferring also destroys observability. A stored integer can be counted, grouped and indexed; a heuristic cannot. "Zero documents remain below generation 3" is the evidence that lets you delete the generation-2 branch. Without the field, deleting old branch code is always an act of faith. ## Where the branching belongs The durable structure is a single upcasting layer at the boundary where a raw document becomes a domain object. Read the document, read its version, then apply small, pure, one-step transformations in sequence — v1→v2, v2→v3 — until the object matches what current code expects, and hand that up. Written as a chain of one-step functions rather than one large conditional, adding a fourth generation means writing one function and touching nothing else; the transformations are also trivially unit-testable against a fixture document of each generation. Crucially, business logic above that layer never sees a version number. When version checks leak into query builders, HTTP handlers and report jobs, every future shape change becomes a multi-file archaeology exercise, and the branches never get removed because nobody can prove they are all dead. ## What actually deserves a bump Not every change is a new generation. A purely additive optional field that old readers ignore and new readers default when absent needs no bump — the old documents are still correctly interpretable. Bump when existing data changes meaning or location: a renamed field, a scalar becoming an array or subdocument, an embedded copy moving out to a reference, a unit change from dollars to cents. The rule of thumb: bump when current code cannot correctly interpret an old document without being told it is old. Silently changing units without a bump is the classic disaster, because nothing errors — the numbers are merely wrong by a factor of a hundred. ## Forward compatibility and the write path The field also protects readers running older code, which may encounter a version higher than they know about. Decide that policy deliberately: for a change where misreading is harmless, ignoring the unknown field is fine; where it is not, failing loudly beats silently misinterpreting. Shipping the version-tolerant reader before the writer is what makes that policy real. Two write-path details bite teams in practice. First, every path that creates a document must set the version, including importers, seed scripts and other services writing the same collection — one unversioned writer reintroduces the ambiguity you bought the field to remove. Second, any update that replaces a document wholesale rather than modifying specific fields must carry the version forward, or it will quietly reset documents to an unknown generation.

  • When is a shape change additive enough that you should not bump the version?
    When old documents remain correctly interpretable without being told they are old — typically adding a new optional field that old readers ignore and new readers treat as absent-means-default. Bump only when existing data changes meaning or location: a rename, a scalar becoming an array, a unit change. Bumping for every additive tweak inflates the branch matrix and trains the team to ignore versions.
  • Where in the codebase should the version-to-shape translation live?
    In one upcasting layer at the boundary where raw documents become domain objects, written as a chain of small one-step transformations. Everything above it works with a single current shape and never reads a version number. If version checks appear in handlers, query builders or reporting jobs, each future change becomes a multi-file hunt and old branches can never be safely removed.
  • What breaks if some writers set the version field and others do not?
    You lose the one guarantee the field buys: that the stored number is authoritative. An importer or sibling service that omits it produces documents indistinguishable from pre-versioning legacy records, so migration progress counts become meaningless and the read path has to fall back on shape sniffing anyway. Every write path — including seed scripts and bulk replaces that overwrite whole documents — must set or preserve it.

It is the print run marked inside a book. The publisher cannot recall the copies already on shelves, so every copy states which edition it is and the reader adjusts for the differences.

saying these in an interview costs you the question

  • Says the database migrates old documents automatically on read
  • Detects the shape by checking whether a field is present
  • Bumps the version for purely additive, backward-compatible changes
  • Stores the application release number instead of a shape generation
  • Scatters version checks through handlers instead of one upcast layer

context

open as a page

In a document store, how do you choose between migrating documents lazily on read and running a background backfill?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Lazy migration converts a document when application code touches it, so cold records never convert and branch code lives forever. A backfill sweeps the whole collection and gives a finish line. Most real migrations ship the tolerant reader first, then backfill, then delete the branch.

open as a page

When you introduce a schemaVersion field into a live collection, how do you handle documents written before it existed?

level: middleimportance: should knowfreq 42%

basics

~20 s

Treat an absent version field as the oldest known generation by convention, rather than rewriting every document just to stamp a number on it. Read code normalizes missing to 1, and migration progress is counted as documents that are missing the field or below the current generation.

open as a page

What must be true of a document shape change so you can safely roll the application release back after deploying it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The previous release must still be able to read documents the new release wrote. That means shipping the tolerant reader before the new writer, keeping the old fields populated while both releases can run, and delaying any destructive rewrite until rollback is off the table.

open as a page

Your read path branches across four document generations at once. How do you decide what to backfill, retire, and enforce going forward?

level: principalimportance: should knowfreq 28%

basics

~20 s

Measure how many documents sit on each generation, price each surviving branch as permanent code and test surface, then set a supported-generations floor. Fund a backfill to clear everything below it, verify a zero count including consumers outside your service, and enforce the floor with monitoring so the count cannot creep back up.

open as a page