When a data-access layer loads a row into a tracked object, what snapshot does it keep, and what is that snapshot for?
answer
- a second copy of the row
- taken at load, read at flush
- compare, do not be told
- no difference, no statement
basics
~20 sA tracking layer copies each loaded row's column values into a private snapshot beside the object. At flush it compares the live object against that snapshot; fields that differ become the UPDATE, and an object with no difference produces none.
solid answer
~50 sWhen a data-access layer that tracks loaded objects reads a row, it materialises the object the caller uses and privately records the values exactly as they arrived. That record is the snapshot, sometimes called the original or loaded values, and it belongs to the unit of work rather than to the object. Application code assigns fields normally; nothing is written at that moment. At flush the layer walks its tracked objects and compares each mapped field against the snapshot — this comparison is what `dirty checking` means. No differing field means no statement at all; one or more differences produce an UPDATE, and after it succeeds the snapshot is replaced with the written values so an unedited object never writes twice. The price is a second copy of every loaded row and a compare over the whole tracked set.
go deeper
Recall the shape: the layer keeps the loaded values, you edit the object, and at flush it compares the two and writes only what differs. No explicit save call is needed for an object that was loaded.
Be able to explain the mechanics: when the snapshot is captured, that the compare runs over the whole tracked set, that it is replaced after a successful write, and that an edit-and-revert produces nothing.
Show that you cost it. Two copies per loaded row and a full-set compare per flush is the price of tracking, and you should know when to read untracked or project instead of loading mapped objects.
Frame the trade: comparison asks nothing of the model but pays memory and flush time, while instrumented notification is cheap at flush and couples the model to the layer. Decide which the codebase can live with.
## What a snapshot is A data-access layer that **tracks** the objects it loads does two things with every row it reads. It materialises the object the application will hold, and it privately records the column values exactly as they arrived — a flat set of **original values** usually called the *snapshot* or *loaded state*. The snapshot belongs to the unit of work, not to the object: application code never receives it, cannot assign to it, and normally has no way to reach it. Its whole purpose is to let the layer answer one question — *what changed since load?* — without the object having to cooperate in any way. The caller assigns fields the way it would on any object in memory. Nothing signals the layer, nothing is queued, and no statement is issued at the moment of assignment. ## What the layer does when the unit of work flushes 1. It iterates the objects it is tracking. 2. For each one it walks the mapped fields and compares the current value against the value held in the snapshot. That comparison is exactly what the term **dirty checking** names. 3. An object with no differing field yields no statement — loading a row and only reading it costs nothing at write time. 4. An object with at least one difference yields an UPDATE. *Which* columns the statement carries is a layer decision: some emit only the differing columns, others emit every mapped column with its current value. 5. Once the statement succeeds the layer replaces the snapshot with the values just written, so the same untouched object flushed again produces nothing. *When* that flush happens — on an explicit call, before a query that would be affected, or at commit — is a separate question from how the change was noticed; this one is only about the noticing. | moment | what the layer does | what it costs | |---|---|---| | load | captures original values beside the object | a second copy of the row's values | | edit | nothing at all | nothing | | flush | compares live values against the snapshot | fields compared x objects tracked | | after write | replaces the snapshot with what was written | nothing beyond the write | | object leaves the set | drops the snapshot | detection stops for that object | ## Noticing is not writing The gap between assignment and statement is deliberate and is usually called **write-behind**: changes accumulate in memory and are turned into statements later, in one go. That is why a method can assign three fields on two objects and produce two statements rather than six, and why an object edited and then edited back to its loaded value produces none. It also means the database sees nothing until the flush, so a query issued from raw SQL in the middle of the method may not see the pending edit. ## Why compare, rather than be told Comparison is the mechanism that asks the least of the model. The object does not have to inherit from a base type, does not have to raise an event on assignment, and does not have to be built by the layer at all — a plain object with plain fields works. That generality is bought with two costs: - **Memory.** Every tracked row is held roughly twice: once as the object, once as the original values. - **Flush time.** The compare visits every tracked object, not just the edited ones, because until it has compared them the layer does not know which ones differ. - **Reads pay both.** A query that loads ten thousand rows nobody intends to modify still snapshots ten thousand rows and still scans them at flush. Other layers make the opposite trade: they have the object report its own writes through generated or intercepted members, keeping a dirty set incrementally and skipping both the copy and the scan — at the cost of requiring the model to be instrumented. Others again keep no tracked set whatsoever and expect the code to name the row and the columns it wants written. ## How to opt out when you do not need it - Read **untracked** (often offered as a read-only read) when a path will never write: the layer materialises objects but keeps no original values, so there is nothing to compare and nothing to spend. - Select just the columns you need into a plain transfer shape rather than a mapped object, so no tracking applies in the first place. - Keep write paths narrow, so the tracked set stays small where flushes actually happen. The one thing to remember about untracked reads is that they are silent: the objects look ordinary, so assigning a field on one compiles, runs, and quietly changes nothing in the database.
- Does keeping a snapshot really mean two copies of every loaded row are in memory?Roughly, yes. The object graph is one copy and the flat array of loaded values is another, so a tracked read of a wide row costs about twice what the same read costs untracked. That doubling is the main reason layers offer a read-only or untracked read mode for large result sets.
- If code assigns a field the same value it was loaded with, is an UPDATE produced?Under snapshot comparison, no: the compare sees equal values on both sides and nothing is written. A layer that instead marks the object dirty when a member is written may treat it as changed and emit a statement, unless it re-verifies the value before flushing. The mechanism, not the assignment, decides.
- Where does the snapshot come from for an object the code created rather than loaded?There is none. A newly created object has no loaded state to compare against, so the layer treats it as an insert rather than diffing it. Its snapshot is established from the values written when the insert goes out, after which ordinary comparison applies to it like any other tracked object.
It is like photocopying a form before handing someone the original to fill in: at the end you lay the two side by side and only the boxes that differ get typed up.
saying these in an interview costs you the question
- Believes an already-loaded tracked object needs a second save call before it will update
- Thinks the snapshot is the object itself, so edits change both sides
- Says the snapshot is written back to the database at commit
- Assumes merely reading a tracked row produces an UPDATE
- Expects change detection to keep working after the object leaves the tracked set
- Thinks tracking is free because the layer is told about every assignment