skip to content

Unit of Work Mechanics

The set a data-access layer holds between load and commit: how edits are noticed, when they turn into statements, how long the set lives. Interviewers probe it because a mapper's surprises start here.

on this pageshow

questions

29

When a data-access layer loads a row into a tracked object, what snapshot does it keep, and what is that snapshot for?

level: juniorimportance: must knowfreq 72%

answer

  1. a second copy of the row
  2. taken at load, read at flush
  3. compare, do not be told
  4. no difference, no statement

basics

~20 s

A tracking layer copies each loaded row's column values into a private snapshot beside the object. At flush it compares the live object against that snapshot; fields that differ become the UPDATE, and an object with no difference produces none.

solid answer

~50 s

When a data-access layer that tracks loaded objects reads a row, it materialises the object the caller uses and privately records the values exactly as they arrived. That record is the snapshot, sometimes called the original or loaded values, and it belongs to the unit of work rather than to the object. Application code assigns fields normally; nothing is written at that moment. At flush the layer walks its tracked objects and compares each mapped field against the snapshot — this comparison is what `dirty checking` means. No differing field means no statement at all; one or more differences produce an UPDATE, and after it succeeds the snapshot is replaced with the written values so an unedited object never writes twice. The price is a second copy of every loaded row and a compare over the whole tracked set.

go deeper

for a junior

Recall the shape: the layer keeps the loaded values, you edit the object, and at flush it compares the two and writes only what differs. No explicit save call is needed for an object that was loaded.

for a middle

Be able to explain the mechanics: when the snapshot is captured, that the compare runs over the whole tracked set, that it is replaced after a successful write, and that an edit-and-revert produces nothing.

for a senior

Show that you cost it. Two copies per loaded row and a full-set compare per flush is the price of tracking, and you should know when to read untracked or project instead of loading mapped objects.

for a principal

Frame the trade: comparison asks nothing of the model but pays memory and flush time, while instrumented notification is cheap at flush and couples the model to the layer. Decide which the codebase can live with.

## What a snapshot is A data-access layer that **tracks** the objects it loads does two things with every row it reads. It materialises the object the application will hold, and it privately records the column values exactly as they arrived — a flat set of **original values** usually called the *snapshot* or *loaded state*. The snapshot belongs to the unit of work, not to the object: application code never receives it, cannot assign to it, and normally has no way to reach it. Its whole purpose is to let the layer answer one question — *what changed since load?* — without the object having to cooperate in any way. The caller assigns fields the way it would on any object in memory. Nothing signals the layer, nothing is queued, and no statement is issued at the moment of assignment. ## What the layer does when the unit of work flushes 1. It iterates the objects it is tracking. 2. For each one it walks the mapped fields and compares the current value against the value held in the snapshot. That comparison is exactly what the term **dirty checking** names. 3. An object with no differing field yields no statement — loading a row and only reading it costs nothing at write time. 4. An object with at least one difference yields an UPDATE. *Which* columns the statement carries is a layer decision: some emit only the differing columns, others emit every mapped column with its current value. 5. Once the statement succeeds the layer replaces the snapshot with the values just written, so the same untouched object flushed again produces nothing. *When* that flush happens — on an explicit call, before a query that would be affected, or at commit — is a separate question from how the change was noticed; this one is only about the noticing. | moment | what the layer does | what it costs | |---|---|---| | load | captures original values beside the object | a second copy of the row's values | | edit | nothing at all | nothing | | flush | compares live values against the snapshot | fields compared x objects tracked | | after write | replaces the snapshot with what was written | nothing beyond the write | | object leaves the set | drops the snapshot | detection stops for that object | ## Noticing is not writing The gap between assignment and statement is deliberate and is usually called **write-behind**: changes accumulate in memory and are turned into statements later, in one go. That is why a method can assign three fields on two objects and produce two statements rather than six, and why an object edited and then edited back to its loaded value produces none. It also means the database sees nothing until the flush, so a query issued from raw SQL in the middle of the method may not see the pending edit. ## Why compare, rather than be told Comparison is the mechanism that asks the least of the model. The object does not have to inherit from a base type, does not have to raise an event on assignment, and does not have to be built by the layer at all — a plain object with plain fields works. That generality is bought with two costs: - **Memory.** Every tracked row is held roughly twice: once as the object, once as the original values. - **Flush time.** The compare visits every tracked object, not just the edited ones, because until it has compared them the layer does not know which ones differ. - **Reads pay both.** A query that loads ten thousand rows nobody intends to modify still snapshots ten thousand rows and still scans them at flush. Other layers make the opposite trade: they have the object report its own writes through generated or intercepted members, keeping a dirty set incrementally and skipping both the copy and the scan — at the cost of requiring the model to be instrumented. Others again keep no tracked set whatsoever and expect the code to name the row and the columns it wants written. ## How to opt out when you do not need it - Read **untracked** (often offered as a read-only read) when a path will never write: the layer materialises objects but keeps no original values, so there is nothing to compare and nothing to spend. - Select just the columns you need into a plain transfer shape rather than a mapped object, so no tracking applies in the first place. - Keep write paths narrow, so the tracked set stays small where flushes actually happen. The one thing to remember about untracked reads is that they are silent: the objects look ordinary, so assigning a field on one compiles, runs, and quietly changes nothing in the database.

  • Does keeping a snapshot really mean two copies of every loaded row are in memory?
    Roughly, yes. The object graph is one copy and the flat array of loaded values is another, so a tracked read of a wide row costs about twice what the same read costs untracked. That doubling is the main reason layers offer a read-only or untracked read mode for large result sets.
  • If code assigns a field the same value it was loaded with, is an UPDATE produced?
    Under snapshot comparison, no: the compare sees equal values on both sides and nothing is written. A layer that instead marks the object dirty when a member is written may treat it as changed and emit a statement, unless it re-verifies the value before flushing. The mechanism, not the assignment, decides.
  • Where does the snapshot come from for an object the code created rather than loaded?
    There is none. A newly created object has no loaded state to compare against, so the layer treats it as an insert rather than diffing it. Its snapshot is established from the values written when the insert goes out, after which ordinary comparison applies to it like any other tracked object.

It is like photocopying a form before handing someone the original to fill in: at the end you lay the two side by side and only the boxes that differ get typed up.

saying these in an interview costs you the question

  • Believes an already-loaded tracked object needs a second save call before it will update
  • Thinks the snapshot is the object itself, so edits change both sides
  • Says the snapshot is written back to the database at commit
  • Assumes merely reading a tracked row produces an UPDATE
  • Expects change detection to keep working after the object leaves the tracked set
  • Thinks tracking is free because the layer is told about every assignment
open as a page

In a data-access layer with a tracked set, what happens to an object's unwritten edits and unresolved links when it is detached?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Detaching removes the object from change tracking: later edits are no longer noticed or written, edits not yet flushed are usually lost, and links left as unresolved stand-ins have no set behind them to load through.

open as a page

In a data-access layer that tracks loaded objects, what does a flush do, and how is it different from a commit?

level: juniorimportance: must knowfreq 82%

basics

~20 s

A flush turns the unit of work's accumulated in-memory changes into SQL statements sent on the already-open transaction. A commit ends that transaction, making the changes durable and visible to others. Flushed-but-uncommitted work is still discarded by a rollback.

open as a page

In a data-access layer with an identity map, what does loading the same row twice inside one tracked set give you?

level: juniorimportance: must knowfreq 68%

basics

~20 s

You get the same object both times. A tracked set's identity map, keyed by type plus identifier, makes one row exactly one object for the life of that set; a by-key lookup it already holds needs no statement.

open as a page

In a data-access layer that tracks loaded objects, what do the transient, managed, detached and removed states mean?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Transient means never stored and not watched. Managed means the tracked set holds it and will write its edits. Detached means it once was managed but its set has ended. Removed means a delete is planned.

open as a page

When a data-access layer opens and closes a unit of work, what happens to the underlying database connection?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Opening a unit of work borrows a pooled connection; every statement it runs travels over that connection; closing hands it back, still established. A unit that is never closed keeps holding it, and the pool drains.

open as a page

How do snapshot comparison, change notification, and explicit marking differ as ways a data-access layer notices edits?

level: middleimportance: must knowfreq 62%

basics

~20 s

Snapshot comparison keeps original values and diffs them at flush: general, but costs a second copy and a full scan. Notification has instrumented members report each write, keeping a dirty set incrementally. Explicit marking makes the code name the write.

open as a page

When a detached object rejoins a tracked set, how does reattaching that instance differ from copying its state onto a freshly loaded one?

level: middleimportance: must knowfreq 66%

basics

~10 s

Reattaching makes the very instance you hold tracked again. A copy-in instead resolves the row's own instance, assigns your values onto it and returns that one, leaving the object you passed in detached forever.

open as a page

Why does a mapper with a tracked set flush automatically before running a query, and when does that not help?

level: middleimportance: must knowfreq 66%

basics

~20 s

A query is evaluated by the database, which knows nothing about unwritten in-memory edits, so its filters ignore them. Many layers therefore flush pending changes before an affected query; that cannot help statements the layer never sees.

open as a page

When a query returns a row the tracked set already holds with an unsaved edit, which values does the caller get back?

level: middleimportance: must knowfreq 58%

basics

~20 s

Normally the held object, edit intact - not the columns the database just sent. The layer reconciles each returned row against its identity map and hands back the instance it already holds, dropping the freshly read values.

open as a page

How does one generic save call in a data-access layer decide whether to insert a new row or update an existing one?

level: middleimportance: must knowfreq 66%

basics

~20 s

It infers newness from a signal: an unset generated key, an absent version value, or a probe read for the row. Each signal can guess wrong or costs a round trip, which is why explicit insert calls exist.

open as a page

How do per-transaction, per-request and conversation-spanning unit-of-work scopes differ, and what does each one cost?

level: middleimportance: must knowfreq 58%

basics

~20 s

The three scopes differ in how long the tracked set and its connection stay alive - one transaction, one request, or a flow spanning requests. Longer scopes reuse more objects, go stale more often, and hold a shared connection longer.

open as a page

Which edits can snapshot comparison at flush miss or report falsely, and what makes a field compare unreliable?

level: middleimportance: should knowfreq 50%

basics

~20 s

Comparison misses an edit when the snapshot shares a mutable reference with the object, so both change together. It reports a false change when a value never round-trips equal: truncated precision, collation folding, or equality that ignores fields.

open as a page

When a partially populated object is copied into a tracked set, how are its nulls, absent collection elements and linked objects treated?

level: middleimportance: should knowfreq 57%

basics

~20 s

A copy-in overwrites the target field by field, so a null in the incoming object stores a null, a collection is made to match the one handed in, and links are followed only where the mapping cascades the copy.

open as a page

Why do two reads of one row inside a tracked set agree even when the isolation level permits non-repeatable reads?

level: middleimportance: should knowfreq 44%

basics

~20 s

Because the second read is served from the identity map: it returns the object already held, so the reads agree by instance reuse, not by an engine guarantee. The held object can still be stale.

open as a page

When a save or delete cascades along an object's links, what happens, and how can a cascade reach too far?

level: middleimportance: should knowfreq 55%

basics

~20 s

A cascade applies one object's state transition to the objects reachable through links declared to cascade, so one save or delete becomes many. It reaches too far when it crosses a link to shared data that other rows still need.

open as a page

Why does it matter whether a data-access layer takes its connection at the boundary or at the first statement?

level: middleimportance: should knowfreq 44%

basics

~20 s

Acquisition timing decides how much of a boundary's duration occupies a pool slot. Taking the connection at the boundary is predictable but pays for pre-work; taking it at the first statement lets a boundary that never queries borrow nothing.

open as a page

A tracking layer loads hundreds of thousands of rows for a report and slows as it runs; what does change detection cost here, and how do you keep those reads out of the tracked set?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Tracking holds each row twice, object plus snapshot, and every flush compares the whole accumulated set, so cost grows as the read proceeds. Read untracked, or project the needed columns into an unmapped shape, so no original values are kept.

open as a page

An object edited outside the tracked set is written back an hour later and another user's change vanishes — why, and how do you fix the write path?

level: seniorimportance: should knowfreq 58%

basics

~20 s

The detached object is a snapshot of the row an hour earlier, and writing it back assigns every mapped field, so a column another writer changed is overwritten with the old value and the update still succeeds.

open as a page

Why does a flush fail on a unique constraint when the same unit of work deleted that row and then inserted a replacement?

level: seniorimportance: should knowfreq 54%

basics

~20 s

The unit of work does not replay statements in call order; it groups them by kind and dependency, sending inserts before deletes. So the replacement meets the original still holding the key. Flushing between them forces the order.

open as a page

After a flush fails part way through sending its statements, what is the unit of work worth afterwards?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Very little. Some statements were sent, one failed, the rest never attempted, and the tracked set no longer describes any real state. Treat it as unusable, and never reuse its objects as if they were saved.

open as a page

Why does comparing mapped objects by reference identify a row reliably inside one tracked set but not across two?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Inside one tracked set the identity map guarantees one instance per row, so reference comparison is a valid row-identity test. Across two sets one row has two instances and the test reports 'different' - compare by identifier at that boundary.

open as a page

Two units of work over the same rows are open at once in one process; what stale-object problems follow and how do you contain them?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Each set has its own identity map, so one row becomes two drifting objects: edits in one are invisible in the other, and the later write can overwrite the earlier. One set per operation, plus a version check, contains it.

open as a page

Why do tracked objects go stale after an insert-or-update statement issued outside the tracked set?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The layer only knows about writes it made itself. A statement sent around it moves the rows while tracked objects keep their old values and old version numbers, so a later write-out can overwrite the newer row.

open as a page

What does keeping the unit of work open through response rendering buy a team, and what does it cost?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Convenience: deferred links still resolve while the view walks the object graph, so nobody declares in advance what a screen needs. It costs a connection held through the slowest phase and statements issued from a layer nobody reviews.

open as a page

When would you prefer explicit change marking over automatic change detection, and what does each choice cost a team?

level: principalimportance: should knowfreq 42%

basics

~20 s

Automatic detection makes the write set a runtime property no reviewer can see, so incidental assignments become real statements. Explicit marking makes writes visible at the cost of discipline: a forgotten mark loses data silently. Choose per path.

open as a page

How do you decide where explicit flush points belong in a codebase, and what does forcing one early cost?

level: principalimportance: should knowfreq 42%

basics

~20 s

Default to flushing at the transaction boundary; treat an explicit flush as a rare exception with a stated reason. Forcing one early takes row locks sooner and holds them to commit, gives up grouping, and moves where errors surface.

open as a page

How would you set unit-of-work and connection-lifetime policy for a service mixing short requests, long batch jobs and multi-step flows?

level: principalimportance: should knowfreq 38%

basics

~20 s

Set policy per workload against one invariant: a unit of work is open only while it is issuing statements. Requests get a thin boundary; batches one unit per chunk with the tracked set discarded between chunks; flows keep detached state.

open as a page

For a large codebase, when is one generic save the right default, and when should writes be explicit inserts and updates?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Default to a generic save where the store assigns keys and writes are ordinary edits; require explicit inserts and updates where the application assigns keys, writes are high volume, or creating an existing thing must fail.

open as a page