skip to content

After a set-based update runs through a data-access layer, why can objects already loaded in the unit of work be wrong?

level: middleimportance: must knowfreq 64%

answer

  1. the layer never saw those rows change
  2. the identity map answers before the database does
  3. dirty checking compares against a stale snapshot
  4. bulk first, or detach afterwards

basics

~20 s

The statement changes rows in the database, but the unit of work never sees those rows. Its tracked objects still hold the earlier values, and the layer keeps serving those stale copies — or flushes them back over the change.

solid answer

~50 s

A tracked object is a copy: loaded once, held in the unit of work's identity map, and returned for every later lookup of that identity. A set-based update never passes through that machinery — the database changes the rows directly — so the layer has no way to know a copy is out of date. Two things then go wrong with the copies themselves. Reads inside the same unit of work keep returning the pre-statement values, because the identity map answers before the database is consulted. Worse, if such an object is dirty, dirty checking compares it against its **load-time snapshot** and flushes the old field values back, silently reverting part of the bulk change. The fixes are ordering and eviction: run the bulk statement before you load anything affected, or detach the affected objects (or clear the unit of work) immediately after, then reload what you still need.

go deeper

for a junior

Remember the shape of the bug: a bulk statement changes rows behind the layer's back, so objects loaded before it still show the old values.

for a middle

Explain the two mechanisms by name — the identity map serving the stale copy, and the load-time snapshot making dirty checking overwrite the change — and give the bulk-before-load ordering fix.

for a senior

Show how you would keep this out of a codebase: bulk statements in their own unit of work, targeted eviction where they cannot be, and a test that loads first and asserts the value after the statement rather than in a fresh session.

for a principal

Consider whether the codebase should allow a bulk path and an object path to touch the same rows in the same request at all, and what convention or review rule makes the safe ordering the default rather than a thing to remember.

## Why the copies exist at all A unit of work keeps an **identity map**: at most one object per row identity, for as long as the unit of work lives. That map is what makes repeated lookups cheap and what makes two references to the same row the same object. Each tracked object also carries a **snapshot** of the values it was loaded with, so that at flush time the layer can compare current fields against the snapshot — **dirty checking** — and write only what changed. Both mechanisms assume one thing: that changes to a row go through the object. A set-based write breaks that assumption on purpose. ## What actually goes wrong Consider a unit of work that has loaded some order objects, then issues one statement that archives every closed order. ```sql UPDATE orders SET status = 'ARCHIVED' WHERE status = 'CLOSED' ``` Three distinct failures follow, and they are worth separating because they have different fixes. 1. **Stale reads.** A later lookup by identity is answered from the identity map without touching the database, so it returns `status = 'CLOSED'` — a value that no longer exists in any row. Code that branches on that value takes the wrong branch. 2. **Lost updates on flush.** If a tracked object was also modified in memory, dirty checking compares it against its load-time snapshot. Fields the bulk statement changed are not part of that comparison at all — the object never knew about them — so the flush writes the object's own, pre-statement values back and quietly undoes the bulk change for those rows. 3. **Ordering surprises.** Many layers flush pending changes before executing a statement that could be affected by them, but nothing flushes *after*. So a bulk statement can see your pending object changes while your objects never see the bulk statement's effect — the influence runs one way only. Anything cached beyond the unit of work is stale for the same reason, and re-populating it is a separate concern from the tracked set. ## The remedies, in order of preference - **Order the work so nothing is loaded yet.** Run the bulk statement at the start of the unit of work, before any query has materialised an affected object. Nothing can be stale if nothing was loaded. - **Evict what the statement touched.** Detach the affected objects right after the statement, or clear the unit of work entirely if the affected set is broad. Subsequent lookups then go back to the database. Clearing is blunt: it detaches everything, including objects you meant to keep working with, and any pending unflushed changes to them are lost. - **Reload explicitly.** Where you must keep working with specific objects, refresh them from the database after the statement so their snapshot matches reality. - **Do not mix in one unit of work.** The cleanest structure is a unit of work that does bulk statements *or* object work, not both. Keeping them apart removes the whole class of bug rather than managing it. ## Why this bites in production and not in tests A test typically issues the statement in a fresh unit of work and asserts against a fresh read, so the identity map is empty and everything looks correct. A real request handler loads objects for validation or authorisation first, then does the bulk change, then continues to work with what it loaded — the exact sequence that produces the stale read. The symptom in production is usually not an exception: it is a value that is right in the database and wrong on the screen, or a row that flips back to its old state minutes after a mass update ran. ## What to say in an interview Name the mechanism, not just the symptom: the identity map serves the stale object, and the snapshot makes the flush overwrite the bulk change. Then give the ordering fix first — bulk before load — and eviction as the fallback, noting that clearing the unit of work is coarse and discards pending changes too. That sequence shows you understand *why* the layer cannot detect the change, which is the real content of the question.

  • Why can a stale tracked object actively revert part of the bulk change rather than merely read wrong?
    Because dirty checking compares the object's current fields against the snapshot taken when it was loaded, and writes the differences. The bulk statement's effect appears in neither the object nor its snapshot, so the flush sends the object's pre-statement values for those columns and overwrites what the statement wrote.
  • When is clearing the whole unit of work the wrong reaction to a bulk statement?
    When other objects in it carry pending changes you still intend to flush: clearing detaches everything and those changes are lost. Prefer detaching only the objects the statement's predicate could have matched, or restructure so the bulk statement runs before anything is loaded.
  • Does running the bulk statement first guarantee later reads are correct?
    Within that unit of work, yes for the rows it changed — nothing was loaded, so nothing is stale, and later queries see the new values inside the same transaction. It says nothing about caches outside the unit of work, or about concurrent transactions whose own snapshots may predate the statement.

It is like editing the master schedule on the wall while everyone is holding a photocopy taken this morning. The wall is right, every copy is wrong, and anyone who writes their copy back onto the wall erases the update.

saying these in an interview costs you the question

  • Thinks the layer detects a bulk statement and refreshes affected objects
  • Believes a stale object only causes wrong reads, never a wrong write
  • Says a later query in the same unit of work always re-reads the row
  • Treats clearing the unit of work as free rather than as discarding pending changes
  • Assumes a stale copy is a concurrency problem rather than a self-inflicted one