skip to content

Why is the price paid stored on an order document not the same kind of duplication as a copied product name?

level: middleimportance: should knowfreq 48%

answer

  1. ask what happens when the source changes
  2. one is a record of an event
  3. an invoice must reproduce it years later
  4. the field name should carry the intent
  5. classify every field into one of three buckets

basics

~20 s

The price paid is a distinct historical fact the order owns, so it must never track the product's current price. A copied product name is a cache of a mutable fact elsewhere and carries a synchronisation obligation. Only the second is duplication debt.

solid answer

~50 s

Two values can look identical on disk and mean completely different things. The price the customer paid is a fact about the order: it was true at checkout, an invoice must reproduce it years later, and updating it when the catalogue price changes would be a data-corruption bug, not a fix. A product name copied in for display is a mirror of a value that still lives somewhere else and is still allowed to change, so someone owes it an update policy. The test is a single question: if the source changes, must this value change too? If yes, it is a cache and needs an owner and a staleness bound. If no, it is not really duplication at all — it is a separate fact that deserves its own name, like `unitPriceAtPurchase`, and the reference id should sit beside it so the current record is still reachable.

code

json · 12 lines
json
{
  "_id": "order-9001",
  "productId": "prod-42",
  "snapshot": {
    "nameOnReceipt": "Blue Widget",
    "unitPriceAtPurchase": { "amount": 1299, "currency": "EUR" },
    "priceListVersion": "pl-2026-07",
    "taxRateApplied": 0.19
  },
  "customerTierForDisplay": "gold",
  "status": "shipped"
}

go deeper

for a junior

Remember the one-line test: if the source value changes, must this value change too? An order's price paid must not change; a copied display name probably should.

for a middle

Be able to classify every field of a document as owned, frozen snapshot, or cache of a mutable fact, and explain why only the last one creates ongoing maintenance work.

for a senior

Show both failure directions: silent drift from a mis-classified cache, and a destructive backfill that overwrites what a customer was actually charged. Argue for naming and structure that make the second impossible.

for a principal

Set the convention across services: snapshot fields are named and grouped for their intent, versioned sources are referenced by version id, and every cached field is registered with an owner and a staleness bound.

## Two values that look the same An order document might contain `productName` and `unitPrice`, both copied from the product document at checkout. On disk they are indistinguishable — two scalars that were read from somewhere else. Semantically they can be opposites, and confusing them is one of the most common modelling mistakes in document databases. ## The deciding question Ask: *if the source value changes, must this value change too?* If the answer is **yes**, the field is a cache of a mutable fact whose home is another document. It has a synchronisation obligation: an owner who updates it, a bound on how stale it may get, and a way to find and repair copies that drifted. This is the field people mean when they talk about denormalization debt. If the answer is **no**, the field is not a copy at all in the semantic sense. It is a distinct fact that this document owns, which merely happened to be derived from another document at a moment in time. The order's price paid, the tax rate applied, the shipping address used, the terms version accepted, the product name as it appeared on the receipt — all of these are facts about the *event*, not about the current state of the referenced entity. Rewriting them to match today's catalogue would falsify history and, in regulated domains, produce an invoice that no longer matches what the customer agreed to. ## Why the distinction matters operationally Get it wrong in the cache direction and you ship silent divergence: a product name that never updates while the product team assumes it does. Get it wrong in the snapshot direction and you ship something worse — a well-intentioned backfill that "repairs" historical orders to the current price, destroying the record of what was actually charged and producing an accounting discrepancy nobody can reconstruct. Backfills are hard to undo because the original value is gone. This is why naming carries real weight here. `unitPrice` on an order is ambiguous and invites a future maintainer to sync it. `unitPriceAtPurchase`, `priceCharged`, `nameOnReceipt` state the intent in the schema itself, where the next reader will see it. A comment in a wiki does not travel with the document; the field name does. ## Snapshot fields still need the reference A point-in-time snapshot does not remove the need for the source identifier. The order still stores `productId` so the current product page can be reached, restocking can be computed, and reporting can group orders by product across renames. The snapshot answers "what did the customer see and agree to"; the reference answers "what is this thing now". Storing only one of the two is a design smell in both directions. ## Mixed documents are normal A single document routinely contains both kinds. An order may hold the price paid (frozen), the shipping address used (frozen), the customer's current loyalty tier for display (a cache that may drift), and the order status (owned outright by the order). A useful discipline is to be able to classify every field of a document into exactly one of three buckets: owned by this document, frozen snapshot of an external fact, or cache of a still-mutable external fact. Only the third bucket generates ongoing maintenance work, and it should be the smallest. ## Versioned sources as a middle path Where the source is genuinely versioned — pricing schedules, terms of service, tax tables — a strong pattern is to store the *identifier of the version* rather than only the value: the order records which price list version applied. That gives an unambiguous, auditable link back to a record that itself never changes, so the value can be recomputed and verified rather than trusted blindly, while still being immune to later edits of the current version. ## How to answer Lead with the test question, classify the two example fields, then make the naming point and note that the reference id stays regardless. Adding the failure mode in each direction — silent drift for a mis-classified cache, a destructive backfill for a mis-classified snapshot — shows you have seen both go wrong.

  • How would you make the frozen fields of an order harder to corrupt by accident?
    Name them for their intent (`unitPriceAtPurchase`), keep them in a clearly delimited sub-document such as `snapshot`, and treat the order as append-only for those fields so the update path never touches them. Where the source is versioned, store the version identifier alongside so the value can be re-derived and verified from an immutable record instead of trusted.
  • Does a frozen snapshot still need the source's identifier?
    Yes. The snapshot answers what the customer saw and agreed to; the identifier answers what the entity is now. Without it you cannot link an order to the current product page, group historical orders across a rename, or compute anything from live catalogue data. Storing the value without the reference makes the order a dead end for every query that needs current state.
  • A field is genuinely a cache — what minimum must accompany it?
    An owner that updates it, a stated bound on how stale it may be, the source identifier indexed so copies are findable, and something that makes drift detectable, such as the source's version or update timestamp stored with the copy. Without those four, the copy is not a design decision, it is an unowned liability that surfaces as two screens disagreeing.

saying these in an interview costs you the question

  • Backfills historical order prices to match the current catalogue
  • Calls every copied value 'denormalization' without asking if it can change
  • Names a frozen field identically to its mutable source
  • Stores the snapshot value but drops the source identifier
  • Assumes an immutable snapshot needs a sync job like any other copy

context