skip to content

How do you decide how stale a duplicated field in your documents is allowed to be, and how do you enforce that bound?

level: seniorimportance: should knowfreq 44%

answer

  1. ask what a wrong read costs
  2. some fields get no staleness at all
  3. the mechanism follows the number
  4. typical lag is not a bound
  5. count what the repair job fixes

basics

~20 s

Classify each duplicated field by what a wrong read costs: display-only fields tolerate seconds or minutes, while money, permissions and inventory tolerate nothing and must be read from the source. Then pick a propagation mechanism whose measured lag meets that bound, and add a sweep that caps the worst case.

solid answer

~50 s

Start from consequence, not from convenience. For every duplicated field, ask what happens if a reader sees the old value: a slightly outdated display name is harmless, an outdated permission or price is an incident. Anything in the second group should not be duplicated at all — it is read from the owning document at decision time, so it is correct by construction. For the rest, state a number: this copy may be up to N seconds behind. That number then chooses the mechanism. Updating copies in the same request gives near-zero lag at the cost of write latency and coupling; propagating asynchronously from a change feed or an outbox gives lag equal to consumer delay, which you must measure and alarm on rather than assume; a scheduled backfill gives lag equal to the schedule. Whatever you choose, a periodic reconciliation sweep is what turns 'usually fresh' into an actual upper bound, because the propagation path will fail sometimes.

go deeper

for a junior

Understand that a duplicated field is only as fresh as whatever process updates it, and that some values — permissions, prices, stock — should simply be read from the record that owns them.

for a middle

Explain how the tolerated staleness picks the mechanism: same-operation updates for tiny fan-out, asynchronous propagation for larger, a scheduled job for slowly changing data.

for a senior

Show the enforcement layer: a stored source version so drift is detectable, monitored propagation lag, and a reconciliation sweep whose interval is the real worst case and whose repair count is a health metric.

for a principal

Own the inventory: every duplicated field has a named owner, a written staleness budget and a detection method, and any field that cannot supply all three does not get duplicated.

## Staleness is a product decision before it is a technical one Every duplicated field implicitly promises something about freshness, and most teams never write that promise down. The result is that nobody can say whether a copy that is three hours old is a bug or normal. The fix is to make the promise explicit per field: this copy may be at most N behind, and here is what happens if it exceeds that. The number comes from consequence, not from how easy propagation is. Ask what a reader does with the value: - **Displayed and forgotten** — an avatar, a display name, a category label. A minute or an hour of staleness costs nothing; the worst outcome is a user seeing their own rename late. - **Used to make a reversible decision** — a search facet, a sort order, a recommendation input. Minutes are usually fine because the decision can be redone. - **Used to authorise, charge, or allocate** — a permission, an account status, a price at checkout, remaining stock. Zero staleness is tolerable. These must not be duplicated for read convenience; they are read from the owning document at the moment of the decision, so the check and the value cannot disagree. That last category is the one interviews probe. A candidate who says "copy the role into the session document so authorisation is fast" has just made revocation take as long as the propagation lag, which is exactly how a fired employee keeps access. ## Mechanism follows the bound Once a field has a number, the mechanism is largely determined: **Same-operation update.** The write that changes the source also rewrites the copies. Lag is effectively zero, but write latency now scales with the number of copies, and the writer becomes coupled to every document shape that holds one. Viable when the fan-out is small and known. **Asynchronous propagation.** The change is published — via a change feed the store itself offers, or via a record written alongside the source change and read by a consumer — and copies are rewritten shortly after. Lag equals consumer lag, which is a real, measurable quantity that must be monitored and alarmed on, not assumed. This is the default for medium and large fan-out because it keeps the user-facing write fast. **Scheduled backfill.** A job periodically rewrites copies. Lag equals the schedule interval plus the run time. Cheap and simple, appropriate for slowly changing dimensions, useless for anything a user expects to see change promptly. A subtlety worth stating: asynchronous propagation makes the write acknowledgement and the read correctness two different moments. If the product needs the editor of a field to immediately see it everywhere, either update the copy that the editor's own view reads synchronously, or have that view read the source. ## Enforcement is a sweep, not a hope A propagation mechanism gives you a *typical* lag. It does not give you a bound, because consumers stall, jobs fail, and code paths get added that bypass the propagation entirely. The bound comes from reconciliation: a periodic sweep that compares copies against the source and repairs the mismatches. Its interval is the true worst-case staleness, and the count of documents it repairs is the health metric — a stable trickle is background noise, a spike means something upstream broke or a new writer skipped the rules. To make the sweep cheap and targeted, store the source's version or update timestamp with the copy. Then the sweep can select copies whose recorded version is behind the source instead of comparing every field of every document, and a reader can even decide at read time that a copy is too old to use. ## Documenting the promise The practical artefact is small: for each duplicated field, record which document owns the source, who propagates it, the maximum staleness, and how divergence is detected. Fields that cannot answer all four are the ones that will cause an incident. This inventory is also what lets a team say no to the next copy — duplication is not free, and the cost is exactly this row of the table. ## How to answer Lead with consequence-based classification, name the category that must never be duplicated, map the remaining bounds onto same-operation, asynchronous and scheduled mechanisms, and finish with reconciliation as the thing that turns a typical lag into a guaranteed one. Mentioning that the repair count is a monitored signal shows you have actually run one of these.

  • Which fields would you refuse to duplicate no matter how much it speeds up reads?
    Anything that authorises or allocates: permissions and roles, account suspension status, remaining inventory, the price used at checkout. Duplicating them makes revocation and correction take as long as the propagation lag, which is how a revoked user keeps access or an item oversells. These are read from the owning document at the moment of the decision so the check and the value cannot disagree.
  • Why is a measured propagation lag not the same as a staleness bound?
    A measured lag describes the healthy case. Consumers stall, jobs fail, deploys pause propagation, and new write paths bypass it entirely, so the tail is unbounded unless something else caps it. The reconciliation sweep provides that cap: its interval is the real worst case, and the number of documents it repairs tells you when the primary path stopped working.
  • What do you store alongside a copy to make drift cheap to detect?
    The source's version number or last-updated timestamp as of when the copy was written. A sweep can then select only the copies whose recorded version is behind the current source, instead of comparing every field of every document, and a read path can decide that a copy older than its budget should trigger a live lookup. Without such a marker, auditing means re-reading everything.

saying these in an interview costs you the question

  • Copies roles or permissions into documents to speed up authorization checks
  • Treats the observed propagation lag as a guaranteed bound
  • Sets one staleness policy for every duplicated field
  • Has no way to detect that a copy diverged
  • Assumes the write acknowledgement means all copies are already updated

context