skip to content

A field is duplicated into millions of documents across six collections and must now become editable. How do you decide whether to keep duplicating it?

level: principalimportance: should knowfreq 36%

answer

  1. make it two measurable sides
  2. the cost that grows with the team, not the data
  3. decide per collection, not globally
  4. can you even find the copies?
  5. retire reads before retiring the field

basics

~20 s

Measure what the duplication actually buys on the read path, then count the maintenance surface: how many writers can create these documents, how often the field changes, and whether the copies are even findable. Keep duplication only where the read win survives an honest edit and repair cost.

solid answer

~50 s

Treat it as a cost comparison you can defend with numbers rather than a principle. On the benefit side: which read paths depend on the copy, what their latency would be if they read the owning document instead, and whether that difference is user-visible. On the cost side: the edit frequency once the field is editable, the fan-out per edit, the number of independent write paths that can create documents containing the copy — each one is a place a future engineer forgets — and whether the copies can be located by an indexed reference at all. Then choose per collection, not globally. Usually two or three of the six genuinely need the read locality and the rest are copying out of habit; those go back to reading the source. Where duplication stays, consolidate it behind one owning write path, store the source version so drift is detectable, and fund the repair sweep as part of the decision, not as a follow-up.

go deeper

for a junior

Know that a duplicated field becoming editable changes everything: the copy was cheap mainly because it never changed, and now every edit has to reach every copy.

for a middle

Be able to lay out both sides — what the read actually saves versus edit frequency, fan-out and whether the copies are findable by an indexed reference.

for a senior

Sequence the change safely: move reads to the source while the copy is still maintained, verify under real traffic, then stop propagating and remove in resumable batches with a reconciliation sweep as the safety net.

for a principal

Own the decision per collection and the rule going forward: duplication requires a named owner, a stated staleness budget, one creating write path and a funded repair sweep, or it is not approved.

## Reframe the question "Is this duplication still worth it?" is unanswerable in the abstract, and the abstract version is what usually gets argued. Turn it into two measurable sides and decide per collection. ## The benefit side: what the copy actually buys For each collection holding the copy, identify the read paths that use it and ask what they would cost without it. Often the honest answer is *nothing*, because the read already loads the owning document for other reasons, or the query returns a page of twenty items and fetching the referenced records is one additional targeted lookup. Duplication earns its place where the alternative is a per-item lookup across a large result set, where the read path must stay fast under a strict latency budget, or where the copy is what makes a query filterable or sortable at all — sorting by a name you do not store is not something you can do at read time cheaply. Be suspicious of "it avoids a join" as a stand-alone justification. The question is whether the avoided work is on a hot path and whether the difference is visible to a user or to capacity. ## The cost side: the maintenance surface Four numbers describe the ongoing cost, and the third is the one teams overlook: 1. **Edit frequency.** A field that was immutable and is now editable changes the whole calculation; the copy's cost was near zero precisely because nothing ever changed. 2. **Fan-out per edit.** How many documents one edit must rewrite, and whether that number is bounded. Unbounded fan-out is a future incident with a date attached. 3. **Number of independent writers.** Every service, job, import and admin tool that can create a document containing the copy must know to populate it, and every future one must remember. This is the term that grows with the organisation rather than with the data, and it is why duplication that was fine at three writers becomes unmanageable at fifteen. 4. **Findability.** Can you select the documents holding a copy of a given source record by an indexed reference? If the copy was stored without the source id, propagation and repair both degrade to full scans, and the duplication is already unmaintainable regardless of the other numbers. ## Symptoms that the answer is already "stop" Some signals settle the argument without further analysis: nobody can enumerate where the field is copied; copies of copies exist, so a repair must run in a specific order; incidents involving disagreement between screens recur; the reconciliation sweep repairs a growing number of documents each run; a schema change to the copied shape now requires coordinating several teams. Each of these says the surface has outgrown the ownership. ## The options, roughly in order of preference **Narrow the copy.** Keep duplicating only the fields that are genuinely stable and read-critical — an identifier and a display label — and read everything volatile from the source. This usually removes most of the pain for a fraction of the effort of full removal. **Drop the copy where the read does not need it.** Per collection, not globally. The collections that lose the copy read the owning document instead, which is a change to a handful of read paths and permanently retires a slice of the maintenance surface. **Consolidate the writers.** Where duplication stays, make exactly one code path responsible for creating those documents, so the rule that populates the copy exists in one place instead of being re-implemented by each new caller. This converts "every future engineer must remember" into "one function enforces it". **Make the copy a frozen snapshot.** If the business meaning permits — the value as it was at the time of the event — rename it and stop propagating entirely. This eliminates the obligation rather than managing it, and it is surprisingly often the correct reading of the domain. **Keep duplicating and fund the machinery.** A legitimate choice when the read win is real: asynchronous propagation with monitored lag, a stored source version, and a reconciliation sweep with an alarm on its repair count. The point is that this is a funded, staffed decision, not an assumption. ## Executing the change safely Removal is a migration, not a deploy. Read paths move to the source first while the copy is still maintained, so the copy becomes unused but correct; then propagation is switched off and the field's absence is verified against real traffic; only then is it removed from the documents. Backfills run in resumable batches, and the reconciliation sweep is the safety net throughout. Doing this while the field is becoming editable means sequencing carefully: make the read path independent of the copy *before* the first edit lands, so the first rename does not become the first incident. ## How to answer Refuse the global yes/no, give the two sides as measurable quantities, name the writer-count term as the one that scales with the organisation, decide per collection, and finish with the migration order that lets you retire the copy without a flag day.

  • Which cost term grows with the organisation rather than with the data?
    The number of independent write paths that can create documents containing the copy. Each service, import and admin tool must populate it correctly, and every future one must remember, so the risk scales with headcount and team count rather than with document count. Consolidating creation behind one owning code path is the direct mitigation; unbounded writers is the strongest argument for dropping the copy.
  • How would you retire a duplicated field without a flag day?
    Move the read paths onto the owning document first, while the copy is still being maintained, so the copy becomes unused but correct and any regression is a latency change rather than a wrong answer. Verify against real traffic, then stop propagating, then remove the field in resumable batches. Keep the reconciliation sweep running throughout so a missed read path shows up as drift, not as a user report.
  • When is keeping the duplication the right call despite the pain?
    When the read win is real and measured: a hot path over large result sets, a strict latency budget, or a filter or sort that is only possible because the value is stored locally. Then keep it, but fund the machinery — one owning write path, a stored source version, monitored propagation lag and a repair sweep with an alarm on its repair count — and record the staleness the field promises.
  • What would make you narrow the copy rather than remove it?
    When the read path needs only a small stable subset — an id and a display label — while the expensive obligation comes from volatile fields copied alongside them. Narrowing keeps the read locality that mattered and removes most of the propagation traffic and most of the drift surface, at the cost of one extra lookup on the rarer paths that need the volatile data.

saying these in an interview costs you the question

  • Answers with a blanket rule instead of measuring the read benefit
  • Ignores how many independent services can write the copied field
  • Removes the field before moving the read paths off it
  • Plans one big migration script with no resumability
  • Keeps duplicating without funding propagation monitoring or repair

context