Which edits can snapshot comparison at flush miss or report falsely, and what makes a field compare unreliable?
answer
- two directions: missed and phantom
- shared reference hides the mutation
- a value that never round-trips equal
- precision, collation, custom equality
basics
~20 sComparison misses an edit when the snapshot shares a mutable reference with the object, so both change together. It reports a false change when a value never round-trips equal: truncated precision, collation folding, or equality that ignores fields.
solid answer
~50 sSnapshot comparison is only as good as the copy it took and the equality it uses. If the snapshot holds a **reference** to the same mutable value the object holds, mutating that value in place changes both sides, the compare sees equality, and the write is silently lost — so layers either deep-copy such values or require them to be immutable. In the other direction, a value that cannot round-trip equal produces a **phantom update** on every flush: a timestamp truncated by the column's precision, a number whose scale is adjusted, a string folded by the collation, or a custom equality that ignores a mapped field. A field not yet loaded has no recorded value to compare, and values the database set are unknown until the row is read again. The symptoms are the diagnostic: an edit that never appears, or a row updated on every request.
go deeper
Remember that change detection compares values, so anything that makes two values look equal when they are not, or unequal when they are, breaks it. Prefer assigning a new value over editing one in place.
Explain both directions with a concrete cause each: a shared reference to a mutable value for the missed edit, and precision or collation for the update that repeats every flush.
Diagnose from evidence. Count statements on a read path, read the row back and diff it field by field, and know that phantom updates cause lock contention and spurious optimistic-lock failures.
Set the rules that prevent it: immutable composite values, in-memory types chosen to round-trip their columns, and a test harness that fails when a read path emits a write.
## Why a compare can be wrong Dirty checking answers *did this field change?* by comparing two values. It can therefore be wrong in exactly two directions: it can say **no** when the answer is yes, losing a write, or say **yes** when the answer is no, producing a statement nobody wanted. Both come from the same two decisions the layer made — how deeply it copied the values at load, and what it counts as equal. ## Direction one: the missed edit The classic case is a **shared reference to a mutable value**. Suppose a mapped field holds a composite value — a money amount, a date range, an address, a byte buffer. If the snapshot stores the same reference the object holds rather than a copy of the value, then mutating that value in place changes what both sides point at. The compare puts the value next to itself, finds it equal, and writes nothing. Layers deal with this in one of three ways, and it is worth knowing which one you are relying on: 1. **Deep-copy** the value into the snapshot at load, so the two sides are independent. 2. **Require immutability** for such values, so in-place mutation is impossible and the only way to change the field is to assign a new value. 3. **Do neither**, and document that mutating a composite value in place is not detected. Related misses come from anything outside the compared set. A field the mapping does not cover is not compared. A field not yet loaded — deferred until first access — has no recorded value to compare against and only joins the comparison once it is loaded. Values written by the database itself, through a default or a trigger, are outside the object's knowledge until the row is read again. ## Direction two: the phantom update Here the compare reports a difference on every flush for a row nobody touched. The cause is always the same shape: **the value that comes back is not equal to the value that went out.** | what differs | typical cause | what you see | |---|---|---| | precision | the column stores fewer fractional digits than the in-memory value | a timestamp or amount rewritten every flush | | scale or type | a numeric value normalised on storage, or a widening on read | an equality that is never satisfied | | text form | trailing whitespace or case folded by the column's collation | a name column updated on every request | | null versus empty | an empty string stored as null, or the reverse | an endless alternation between the two | | custom equality | equality defined over a subset of fields, or over identity | changes detected where there are none, or missed | Phantom updates are more than noise. Each one issues a statement, each one takes a row lock for the rest of the transaction, each one moves a version column and so can make a concurrent writer's optimistic check fail, and each one lands in any change-history mechanism that watches the table. A read-only page that quietly writes a hundred rows is a real production problem, and its cause is nearly always a value that does not round-trip. ## How to find either one - **Count the statements a path emits.** A read path that emits UPDATEs is a phantom; assert the count in a test rather than reading logs by eye. - **Read the row back and compare it with what you wrote**, field by field, for the type you suspect. The field where the two differ is the one that will never compare equal. - **Check the column's precision and collation against the in-memory type.** A mismatch here explains most phantoms. - **For a lost write, assign a fresh value instead of mutating in place** and see whether the statement appears. If it does, the snapshot was sharing your reference. - **Compare the mapped field set against what the class actually holds.** An unmapped field is not compared, and a derived field that looks persistent may be nothing of the kind. ## The habit that prevents most of it Make composite values immutable, so the only way to change a field is to assign it, and choose in-memory types whose values survive a round trip to their column unchanged. Those two habits remove the majority of both failure directions before they can be written, and they leave the compare doing what it is good at: noticing an assignment nobody bothered to announce.
- Why is a phantom update worse than a wasted statement?It takes a row lock for the remainder of the transaction, moves any version column, and feeds every change-history mechanism watching the table. A page that was supposed to read now blocks concurrent writers and can make their optimistic checks fail, so a harmless-looking equality bug turns into contention and spurious conflicts.
- How would you prove that an in-place mutation is being missed rather than rolled back?Assign a fresh value to the same field instead of mutating the existing one and flush again. If the statement now appears, the values were fine and the snapshot was sharing your reference. If it still does not, look at whether the field is mapped and whether the unit of work committed at all.
- Does making composite values immutable remove the need for a snapshot?No. Immutability only guarantees that changing a field requires an assignment, so the compare can trust that a differing reference means a differing value. The layer still has to record the loaded values and diff them, because nothing about immutability tells it which fields were assigned.
saying these in an interview costs you the question
- Assumes the compare always deep-copies mutable values into the snapshot
- Thinks a phantom update is harmless because the values are the same
- Blames a lost write on the transaction when the compare never saw a difference
- Ignores column precision and collation when choosing an in-memory type
- Defines equality over a subset of fields and expects full change detection
- Believes a not-yet-loaded field is compared like any other