skip to content

Two units of work over the same rows are open at once in one process; what stale-object problems follow and how do you contain them?

level: seniorimportance: should knowfreq 50%

answer

  1. two sets, two maps
  2. each instance frozen at its own load
  3. structure first, refresh second
  4. version check makes the collision fail

basics

~20 s

Each set has its own identity map, so one row becomes two drifting objects: edits in one are invisible in the other, and the later write can overwrite the earlier. One set per operation, plus a version check, contains it.

solid answer

~50 s

The one-object-per-row rule is scoped to a single tracked set, so two open sets mean two objects for one row, each frozen at its own load time. Three problems follow: edits made in one set are invisible in the other until they are written and the other set re-reads; a comparison by reference across the two says 'different row' when it is the same row; and whichever set writes second can overwrite the first set's columns with the values it has been holding. Containment is mostly structural — **one tracked set per logical operation**, and never hand a tracked object to code running under a different set; pass an identifier or a plain transfer shape. Where a set is long-lived, refresh the objects a decision depends on, or discard them so the next read is real. Then make the overlap fail loudly rather than silently, with a version column checked in the update's where clause.

go deeper

for a junior

Take away the scope rule: the one-object-per-row promise only covers a single tracked set. If two are open, the same row is two separate objects and they can disagree.

for a middle

Explain the three consequences — invisible edits, failed reference comparison, blind overwrite — and the difference between re-reading an object and discarding it from the set.

for a senior

Diagnose it from symptoms rather than theory: a value moving backwards, an intermittent membership failure. Fix by collapsing to one set per operation and refreshing only what a decision depends on.

for a principal

Treat it as a boundary design question: where sets may be opened, what may cross between them, and which read paths should not be tracked at all. Then require a version check so any remaining overlap fails loudly.

## What actually goes wrong An identity map belongs to one tracked set. Open a second set over the same rows — a background job alongside a request, a nested service that opens its own, a long-lived set plus a short one for a side query — and the guarantee that made your code safe quietly stops applying: - **Two instances, two states.** Each set materialised the row at a different moment and has been accumulating its own edits. Neither knows the other exists. - **Cross-set comparison fails.** Reference equality between the two objects is false even though the row is the same, so membership tests, set arithmetic and identity-based caches all misbehave at the boundary. - **Blind overwrite at write time.** If both sets write, the second write carries the full state it holds, including columns it never intended to change, and lands on top of the first. - **Cascading staleness.** An object in one set may carry associations resolved earlier; an operation reading through it can act on a shape of the data that no longer exists. ## Why it is hard to see Inside each set, everything looks perfectly consistent — that is the whole point of instance reuse. The inconsistency is only visible from outside, which is why it usually first appears as an intermittent report: a value that "goes backwards" after two operations run close together, a membership check that fails for an object the collection visibly contains, an audit column reverting. ## Containment, in the order worth applying it 1. **One tracked set per logical operation.** Most of these failures are two sets that had no business being open together. Open at the operation boundary, close at the end, do not nest a second one inside. 2. **Do not pass tracked objects across a set boundary.** Hand over the identifier, or a plain transfer shape carrying just the fields the other side needs, and let it load or work with its own. Passing the object across is what turns a scoping mistake into an equality and staleness bug. Exactly what happens to an object when it leaves its set — and how it is taken back in — is the detachment topic's material. 3. **Refresh what a decision depends on.** In a set that must live longer than one short operation, re-read the specific objects whose current values a decision turns on, immediately before deciding. A blanket refresh of everything held is expensive and still races. 4. **Discard rather than refresh when you want a clean read.** Discarding removes the object from the set so the next read materialises it again from the database. It also drops pending edits on that object, which is why it is the honest choice when you know your copy is worthless, and the wrong one when it is not. 5. **Make the collision fail.** A version column carried on the object and checked in the update's where clause turns "second writer silently wins" into a failed update the caller can retry. Without it, none of the steps above eliminate the race, they only shrink the window. ## Refresh versus discard versus a new set | Action | Cost | Pending edits on the object | What you hold afterwards | |---|---|---|---| | Re-read the object | One statement | Dropped | The same instance, updated from the row | | Discard the object from the set | None now; a read on next use | Dropped | Nothing; the object becomes detached | | Clear the whole set | None now; reads on next use | All dropped | An empty set, everything detached | | Close and open a fresh set | Set-up cost | All dropped | A clean map that re-reads on demand | Clearing the whole set is the blunt instrument that is reached for too often: it throws away every unwritten change in the operation, including ones nobody was worried about. ## Read-only paths deserve a different answer If a set exists only to serve reads, tracking buys nothing and costs freshness. Reading those paths as untracked projections removes the second map entirely: no held instances, no reconciliation, no cross-set equality question. That is often a cleaner containment than tuning refresh calls, and it is worth proposing when a system has grown a long-lived set mainly to render things. ## What interviewers are listening for That you locate the problem in the *scope* of the guarantee rather than in the mapper being unreliable; that you reach for structure first and refresh calls second; and that you name the version check, because a candidate who only proposes refreshing has described a smaller race rather than a correct system.

  • Why is clearing the whole tracked set a poor first response to a staleness report?
    Because it discards every unwritten change in the operation, not just the object you were worried about, and it forces a re-read of everything on next use. It also hides the real defect, which is usually that two sets were open over the same rows at once. Refresh or discard the specific object instead.
  • Two sets each hold the same row and both write. What stops the second write silently overwriting the first?
    A version column carried on the object and checked in the update's where clause: the second write matches zero rows and fails, so the caller can re-read and retry instead of quietly losing the first change. Without such a check the last writer wins with whatever state it happened to hold.
  • Is a second tracked set ever legitimate rather than a mistake?
    Yes — a short independent set is a normal way to write something that must persist regardless of the main operation's outcome, such as an audit or a failure record. The rule is that it should touch different rows, and that objects must not be shared between it and the main set.

saying these in an interview costs you the question

  • Blames the mapper instead of the two overlapping sets
  • Clears the whole set as the standard fix for staleness
  • Passes tracked objects between code running under different sets
  • Thinks refreshing before writing removes the race
  • Assumes both sets see each other's unwritten edits
  • Relies on last-writer-wins with no version check