skip to content

In a document store, how do you choose between migrating documents lazily on read and running a background backfill?

level: seniorimportance: must knowfreq 58%

answer

  1. One converges, the other never finishes
  2. Cold documents are the problem
  3. Who else reads this collection directly?
  4. A finish line is what deletes branch code
  5. Ship the tolerant reader before either

basics

~20 s

Lazy migration converts a document when application code touches it, so cold records never convert and branch code lives forever. A backfill sweeps the whole collection and gives a finish line. Most real migrations ship the tolerant reader first, then backfill, then delete the branch.

solid answer

~50 s

They answer different questions. **Migrate-on-read** transforms the document in the application after loading it, optionally writing the converted form back. It needs no bulk job and spreads cost across normal traffic, but documents nobody reads are never converted, so you can never prove the old generation is gone and the branch code becomes permanent. It also turns reads into writes on a hot path, and any consumer that reads the store directly — reporting, exports, another service — bypasses your read path entirely and still sees the old shape. A **background backfill** rewrites documents in bounded, resumable batches until none remain on the old generation. That gives you the count that justifies deleting the branch, at the cost of a job you must throttle, restart safely, and write so it cannot clobber concurrent application writes. In practice you do both: deploy a reader that tolerates both generations, run the backfill behind it, verify zero stragglers, then remove the branch.

code

javascript · 16 lines
javascript
const CURRENT = 2;
const BATCH = 500;

async function backfill(store) {
  for (;;) {
    const batch = await store.find({ generationBelow: CURRENT }, { limit: BATCH });
    if (batch.length === 0) return;
    for (const doc of batch) {
      const next = upcast(doc);
      // conditional on the document still being old: a concurrent
      // application write wins instead of being clobbered
      await store.updateFields(doc._id, next, { ifGeneration: doc.schemaVersion ?? 1 });
    }
    await sleep(throttleMs); // keep headroom for live traffic
  }
}

go deeper

for a junior

Know the two options by name: convert a document when code reads it, or run a job that converts them all. Be able to say why the first one never quite finishes.

for a middle

Explain the mechanics of each: what write-back on read costs, and what a batched, resumable job needs so it can be killed and restarted without damage.

for a senior

Demonstrate the production sequence — tolerant reader first, throttled backfill behind it, verify a zero count, then contract — and name the consumers outside your read path that a lazy-only strategy leaves stranded.

for a principal

Own the policy: which collections get funded backfills, how migration work is budgeted alongside the feature that caused it, and how you stop the codebase accumulating permanent branches nobody is allowed to delete.

## The two strategies Once a collection holds more than one document generation, something eventually has to converge them — or the read-path branch is permanent. There are exactly two mechanisms for doing the conversion, and mature teams treat them as complementary rather than as alternatives. **Migrate on read (lazy).** When the application loads a document and finds an older generation, it transforms it in memory to the current shape. Optionally it writes the converted document back, so the next read finds it already current. **Backfill (eager).** A background job walks the collection, finds documents below the current generation, rewrites each one, and stops when none are left. ## What lazy migration buys and costs The attraction is that it requires no separate job, no capacity planning and no operational supervision. Conversion cost is spread across traffic you were serving anyway, and it is proportional to what is actually used: in a collection where ninety percent of documents are never read again, you do a tenth of the work. The costs are real and they are structural. *It never finishes.* Cold documents stay on the old shape indefinitely. You cannot answer "is anything still on generation 1?" with anything except "probably", so the generation-1 branch, its tests and its fixtures stay in the codebase forever, and every subsequent change must consider it. *Write-back puts writes on the read path.* A read that rewrites the document needs write access and consumes write capacity in whatever path is hottest, often precisely the path you were trying not to disturb. Under a read burst, that amplification arrives exactly when you have least headroom. It can also produce surprising contention when many readers touch the same popular documents at once. *It only protects your read path.* Direct readers — analytics extracts, a partner service reading the same collection, an ad-hoc operations query, a search indexer — do not pass through your upcasting layer. They still meet old documents, and they need their own tolerance. ## What a backfill buys and costs The backfill buys a finish line. Once a count of documents below the current generation returns zero, the old branch is provably dead and can be deleted; the codebase converges back to one shape. It also converts documents that no online path would ever touch, which is what makes the collection uniform for every consumer, not just your service. The cost is that a backfill on a large collection is a genuine piece of production engineering. Four properties matter: *Batched and throttled.* Rewriting every document generates write load and, in a replicated deployment, replication traffic. Batches must be small enough to keep latency for real traffic acceptable, with a rate control you can turn down while it runs. *Resumable.* The job will be killed — a deploy, a node restart, a bad afternoon. Selecting the next batch by "still below the current generation" makes restart trivially correct: already-converted documents no longer match. Progress is the shrinking count of remaining documents, not a cursor you must persist. *Idempotent and non-clobbering.* The application is writing concurrently. Make the update conditional on the document still being at the old generation, and transform field-by-field rather than replacing the whole document, so a concurrent business write is not silently reverted to a stale snapshot the job read seconds earlier. *Observable.* Emit remaining count, rate and error count. A backfill you cannot watch is a backfill nobody dares run at full speed. ## Why the answer is usually "both" The sequence that works is not a choice between the two: 1. **Ship the tolerant reader first.** Deploy code that reads both generations and writes the new one. Nothing has migrated yet; the system is simply safe with either shape. 2. **Run the backfill behind it.** Traffic is unaffected because the reader already copes; the job simply reduces the number of old documents from N to zero at whatever pace you choose. 3. **Verify.** Count documents below the current generation. Not "the job reported success" — the count. Include any secondary copies: search indexes, caches, derived collections, exports other teams hold. 4. **Contract.** Remove the old branch, its fixtures and its tests, and — only now — stop writing any legacy field kept for compatibility. Lazy conversion still earns its place inside this sequence as the reader's behaviour, and as the whole strategy for collections where a bulk rewrite is not worth it: small lookup collections, or data with a natural expiry where every record is rewritten or aged out within a known window anyway. There the finish line arrives on its own, and you can put a date on it. ## When lazy alone is defensible Choose lazy-only deliberately, not by default, and say why: the collection is enormous and mostly cold, the change is cosmetic enough that a permanent branch is cheap, or records naturally turn over within a bounded period. Write that reasoning down next to the branch, because the person who inherits it needs to know whether the branch is waiting for something or waiting for nothing.

  • What makes a backfill safe to kill and restart at any moment?
    Selecting each batch by the condition "still below the current generation" and making the update conditional on that same state. Already-converted documents stop matching, so a restart naturally skips them and no external cursor has to survive the crash. Combine it with field-level updates rather than whole-document replacement so a concurrent application write is not overwritten by a stale in-memory copy.
  • Why is 'the backfill job finished successfully' not enough evidence to delete the old read branch?
    A job can exit cleanly having skipped documents it never selected — a filter that missed unversioned records, documents written during the run by an older release, or a shard or partition the job did not cover. The evidence that matters is a count query showing zero documents below the current generation, repeated after the run, plus a check of derived copies such as search indexes and exports.
  • What is the risk of writing converted documents back during a read?
    It puts write load and write permissions on the hottest path, so a read spike becomes a write spike exactly when headroom is lowest, and popular documents can attract concurrent rewrite contention. It also converts only what is read, so it still leaves cold documents behind and gives you no completion signal.

saying these in an interview costs you the question

  • Says lazy migration eventually converts the whole collection
  • Rewrites documents in one unbounded pass with no throttle
  • Replaces whole documents, clobbering concurrent application writes
  • Deletes the old branch because the job exited successfully
  • Forgets that exports and other services read the store directly

context