skip to content

When is a shallow comparison enough at a seam, and what does a deep comparison at a frequently checked seam cost?

level: seniorimportance: should knowfreq 56%

answer

  1. one level, by reference, per property
  2. immutable updates summarise deep changes upward
  3. cost lands on checks, not writes
  4. the equal case is the full walk
  5. normalise at the boundary instead

basics

~20 s

Shallow comparison checks each top-level property by reference, which suffices under immutable updates, because a nested change forces new references up the path. A deep comparison walks the graph on every check, so its cost scales with the data.

solid answer

~50 s

A shallow comparison walks one level: for each top-level property it asks whether the two values are the same, without descending. It suffices when the data was produced by immutable updates, because a nested change has already forced a new reference at every level of the path, so the difference is visible at the top. It fails when the value is rebuilt from outside on each arrival — parsed afresh from a polled payload, say: contents match but every nested object is newly allocated, so the verdict is "changed" every time. A deep comparison fixes that, and the bill arrives at every check rather than at every write: cost proportional to the data, the *equal* case being the full walk, trouble with cycles and with values a generic walk cannot interpret. Usually the better move is upstream — normalise once where data enters.

go deeper

for a junior

Learn the distinction: shallow looks at each top-level property by reference; deep descends through everything. Neither reads anything if the two values are already the same object.

for a middle

Explain why immutable updates make shallow sufficient — a nested change forces new references all the way up the path — and name the case that defeats it, a value re-allocated by its producer.

for a senior

Argue the cost honestly: the walk runs per check rather than per write, equality is its worst case, and a payload that grows turns a free seam into a slow one. Prefer normalising at the boundary.

for a principal

Own it as policy: comparison level is a per-seam decision with a written reason, and a house rule of deep comparison spreads a data-sized cost across the codebase while hiding which seams were actually at risk.

## The three levels of comparison | comparison | what it examines | cost | typical use | |---|---|---|---| | reference | whether both names point at one object | one check | the default at almost every seam | | shallow | each top-level property, by reference, no descent | proportional to the property count | comparing a small set of inputs, or a flat options value | | deep | every leaf, descending through nested values | proportional to the whole graph | values whose producer you do not control | Shallow comparison is the interesting middle. It is not "a bit of deep comparison"; it is a reference comparison applied once per property. That is why it pairs so exactly with immutable updates. ## Why immutable updates make shallow enough When a change is applied by rebuilding the path from the root, a change anywhere below the top surfaces as a **new reference on the path**. The field two levels down changed, so its owner is new, so *its* owner is new, and so on up to the root. A comparison that looks at the top-level properties therefore sees the difference without descending — the producer's discipline has already summarised it. It follows that shallow comparison is a contract with the producer, not a property of the comparison. Shallow is enough when: - the value was produced by path copying and sharing; - the properties are primitives, which compare by value anyway; - unchanged branches are deliberately shared, so their references genuinely mean "unchanged". ## Where shallow stops working The broken case is a value whose producer allocates afresh: - a payload polled on an interval and parsed into new objects each time — identical content, new identities at every level; - a message deserialised on arrival; - a value assembled per pass by mapping or filtering, which allocates a new container and often new elements; - anything crossing a boundary that copies rather than shares. Here shallow comparison reports a change on every check, and so does a reference comparison. Deep comparison is the answer that works at the seam — and it is worth asking first whether the seam is the right place to fix it. ## What a deep comparison actually costs - **It runs on the check, not on the write.** Writes are rare; checks happen on every pass, for every dependent. A cost moved from the rare event to the frequent one usually grows rather than shrinks. - **The equal case is the expensive case.** A walk can stop early when it finds a difference, but proving equality requires visiting every leaf — so the cheapest outcome for the runtime is the most expensive for the comparison. - **It is only as correct as its walker.** Cycles need explicit handling or the walk does not terminate; functions cannot be compared meaningfully; keyed collections, dates and host objects each need dedicated handling, and a naive walker silently treats two different values as equal or two equal values as different. - **A serialise-and-compare shortcut is not deep equality.** Comparing two values by converting them to text and comparing strings depends on property order, drops or mangles anything the text format cannot represent, and allocates two strings per check. - **Equal contents do not restore identity.** A deep comparison can tell one seam that nothing changed, but the value is still a new object, so any other place keyed on identity still sees churn. - **The cost is invisible until the data grows.** A seam that compares a ten-field payload is free in development and a problem when the same seam sees a thousand rows in production. ## The judgment a lead is asked for 1. **Fix the producer first.** Normalise or intern the value where it enters the system: accept the payload, diff or merge it once against what you already hold, and hand on a value whose unchanged branches keep their identity. Every seam downstream then stays on a single reference check. 2. **If you must compare contents, bound the shape.** Compare a small, known set of fields rather than an arbitrary graph — that is a shallow comparison over a chosen projection, and its cost is fixed and reviewable. 3. **Measure the thing you are avoiding.** The needless work a deep walk prevents is sometimes cheaper than the walk. "Compare harder" is a plausible fix that has to earn its place with a measurement, not a hunch. 4. **Make the level a property of the seam, not a house style.** A blanket policy of deep comparison spreads a payload-sized cost everywhere and hides which seams were genuinely at risk. Choose deliberately, seam by seam, and write down why. The honest summary is that shallow comparison is not a weaker deep comparison — it is the comparison that matches immutable production, and reaching past it is a signal that something upstream is allocating where it could be sharing.

  • Why is the equal case the worst case for a deep comparison?
    Because a difference lets the walk stop as soon as it finds one, while equality can only be established by visiting every leaf. So the verdict that lets the runtime skip work is the verdict that cost the most to obtain — the opposite of the bargain a reference check offers.
  • A polled payload arrives with identical contents each time. What is the better fix than a deep comparison at every seam?
    Merge it once at the boundary: compare the arrival against the value you already hold, and keep the existing objects for branches that did not change. One comparison at entry replaces one comparison per seam per pass, and every downstream consumer goes back to a single reference check.
  • Does a deep comparison reporting "equal" remove the churn entirely?
    Only at that seam. The value is still a freshly allocated object, so anything else comparing by identity — another child, a dependency list, a cache key — still sees a change. That is why stabilising the producer fixes the class of problem while a deep comparison fixes one instance.

saying these in an interview costs you the question

  • Calls shallow comparison a cheap approximation of deep comparison.
  • Adds deep comparison at a hot seam without measuring what it saves.
  • Thinks a deep walk exits early even when the values are equal.
  • Treats text serialisation as a correct deep equality check.
  • Expects an equal verdict to stop identity churn at other seams.
  • Assumes a generic walk handles cycles and keyed collections correctly.