How do snapshot comparison, change notification, and explicit marking differ as ways a data-access layer notices edits?
answer
- copy and diff, or be told
- generality versus runtime cost
- who must cooperate: model or programmer
- the failure mode is the tell
basics
~20 sSnapshot comparison keeps original values and diffs them at flush: general, but costs a second copy and a full scan. Notification has instrumented members report each write, keeping a dirty set incrementally. Explicit marking makes the code name the write.
solid answer
~50 sThree mechanisms cover almost every layer. **Snapshot comparison** copies the loaded values and diffs them at flush; it asks nothing of the model class, which is why plain objects work, but it holds two copies of every tracked row and scans the whole tracked set on each flush. **Notification** — through generated or intercepted members, or a stand-in wrapper — has the object tell the layer as each member is written, so the layer maintains a dirty set incrementally and flush is nearly free; the price is that the model must be created by the layer and writes that bypass the instrumented members are invisible. **Explicit marking** means the code names the row, or the columns, it wants written; detection costs nothing and the statement is visible at the call site, but a forgotten mark silently loses the write. Many layers mix them: scalar fields compared, collections tracked through wrappers.
go deeper
Know that a layer either keeps the loaded values and compares them, or is told about each write, or expects you to say what to write. The three feel identical from the calling code until something goes wrong.
Explain each mechanism's cost and its blind spot, and be able to say which one a plain object with plain fields can support without any instrumentation.
Talk about the failure modes in production: writes lost because instrumentation was bypassed, writes nobody intended because a compare found a normalised field, and reads paying for tracking they never use.
Weigh the coupling. Notification buys cheap flushes with a model the layer must build; comparison buys plain classes with memory and scan cost. Pick the one whose failure mode your team can detect.
## The question all three answer A layer that writes changed rows has to learn *which* in-memory values differ from what the database holds. There are only three broad ways to learn it: keep a copy and compare, be told as it happens, or be told at the end by the programmer. Each is a different position on the same trade-off — how much the mechanism asks of the model class versus how much it costs at runtime. ## Snapshot comparison At load the layer records the column values in a private structure beside the object. At flush it walks every tracked object, compares each mapped field with the recorded value, and turns the differences into statements. - **Asks nothing of the model.** Plain objects with plain fields work; no base type, no generated members, no instrumentation. - **Catches everything the mapping covers**, including a field assigned deep inside library code that knows nothing about persistence. - **Costs memory**: roughly two copies of every tracked row. - **Costs flush time**: the scan visits all tracked objects, because until it compares them it cannot know which ones differ. - **Has blind spots** where the copy is shallow — a mutable value held by reference and mutated in place can look unchanged on both sides. ## Change notification Here the object reports its own writes. The layer typically hands back an instance whose members have been instrumented — a generated subclass, a rewritten class, or a stand-in that wraps the real object — and each write raises a notification that adds the object, and often the specific field, to a dirty set. - **Flush is close to free**: the dirty set is already known, and unchanged objects are never visited. - **No original values are kept**, so tracked reads cost about half what comparison costs. - **Requires cooperation**: the model must be instantiable by the layer and must route writes through the instrumented members. A direct field write, or an object built with a plain constructor and then attached, may go unseen. - **Marks on write, not on difference**: assigning the same value can still mark the object dirty, producing a statement that changes nothing, unless the layer verifies the value. - **Leaks into the model**: the instance the caller holds is not the class it wrote, which surprises code that reasons about concrete types or serialises the object. ## Explicit marking The layer keeps no tracked set, or keeps one but does not diff it. The code says which row to write and, often, which columns. This is the normal mode for query builders and thin data-access layers. - **Zero detection cost** and zero extra memory. - **The write is visible in the code**, so a reviewer can see exactly which columns a call touches. - **Partial-column writes are natural**, which matters when two writers routinely touch different columns of the same row. - **A forgotten mark is a silently lost write**, and that is the dominant failure mode. - **The model becomes data**: business objects stop being the unit of persistence and the mapping moves into the call site. ## Side by side | | snapshot comparison | change notification | explicit marking | |---|---|---|---| | cost at load | second copy of each row | none beyond the object | none | | cost at flush | scan of the whole tracked set | proportional to edits | none | | model requirements | none | instrumented members | none | | typical blind spot | in-place mutation of a shared reference | writes bypassing instrumentation | the mark nobody wrote | | write of an identical value | no statement | may still write | writes what you named | ## Choosing, and mixing In practice a layer rarely picks one purely. Collections are commonly tracked by wrapper objects that notify on add and remove even where scalar fields are compared, because diffing a collection element by element is far more expensive than diffing a handful of scalars. A comparison-based layer may also offer a per-read opt-out that keeps no original values at all, which is the same thing as choosing explicit marking for that one read path. When you evaluate a layer for a codebase, ask the three questions the mechanisms differ on: what must my classes look like, what does a read that never writes cost, and what happens when someone forgets. The answers, not the vocabulary, are what distinguishes the approaches.
- Why do many comparison-based layers still track collections through notification?Diffing a collection means comparing membership element by element against a recorded copy, which is far costlier than comparing a few scalars and needs a copy of every element. Wrapping the collection so that add and remove report themselves gives the layer the delta directly, so the flush knows which links to write without walking the contents.
- What breaks when notification-based tracking is given an object the layer did not create?Nothing reports its writes. If the instance was built with a plain constructor and then attached, its members are not instrumented, so edits raise no notification and the object never enters the dirty set. Layers that work this way usually require instances to be obtained from them, or fall back to comparison for such objects.
- Does explicit marking mean the layer cannot help with concurrency at all?No. The code still names the row and the layer can still add a predicate on a version column, so a write whose expected version no longer matches affects zero rows and is reported as a conflict. What explicit marking removes is the automatic discovery of what changed, not the ability to guard the statement.
saying these in an interview costs you the question
- Thinks every layer detects changes the same way, by comparing values
- Claims notification-based tracking still keeps a copy of the loaded values
- Says comparison catches in-place mutation of a shared mutable value
- Believes the object the caller holds is always the class it wrote
- Assumes explicit marking cannot use a version check for concurrency
- Treats a forgotten mark as a compile-time error rather than a silent loss