skip to content

When would you prefer explicit change marking over automatic change detection, and what does each choice cost a team?

level: principalimportance: should knowfreq 42%

answer

  1. who decides the write set
  2. visible in the diff or not
  3. opposite failure modes: too many, too few
  4. decide per path, default reads untracked

basics

~20 s

Automatic detection makes the write set a runtime property no reviewer can see, so incidental assignments become real statements. Explicit marking makes writes visible at the cost of discipline: a forgotten mark loses data silently. Choose per path.

solid answer

~50 s

The trade is between an invisible write set and a manual one. With automatic detection any assignment anywhere in the call graph can become a statement, so the write set is a runtime property that no reviewer can read off the code — which is convenient for rich domain logic and dangerous for wide, hot, or contended rows, where a normalising setter on a read path quietly issues updates. With explicit marking the statement is visible at the call site, partial-column writes are natural, and the blast radius of a change is reviewable; the cost is that a forgotten mark loses the write with no error, and the model drifts toward being plain data. Most mature codebases mix them: detection inside a narrow write boundary, untracked reads everywhere else, and explicit statements for the few paths where write amplification or column-level contention actually matters.

go deeper

for a junior

Understand the two styles: either the layer works out what changed, or your code says what to write. Each one fails in a different way, and neither is simply better.

for a middle

Be able to name the concrete costs — invisible write sets and incidental updates on one side, forgotten writes and more boilerplate on the other — and give one situation that favours each.

for a senior

Argue from production evidence: statement counts per request, contention on shared rows, and the behaviour of history or audit mechanisms when incidental writes reach them.

for a principal

Set an enforceable boundary rather than a preference, choose the failure mode your tooling can see, and be explicit about whether updates carry all columns or only the changed ones.

## What is really being chosen Both mechanisms produce the same statements when everything goes right. They differ in **who decides what gets written, and whether that decision is visible**. Automatic detection derives the write set at runtime from the difference between memory and the loaded values. Explicit marking has the programmer state it. Every argument on either side reduces to a consequence of that one difference. ## The case for automatic detection - **Domain logic stays about the domain.** A method can enforce a rule by assigning fields and never mention persistence at all. - **Nothing is forgotten.** A change made three calls deep still gets written, which removes a whole class of lost-update bugs. - **Statements collapse.** Many edits to one object across a request become one statement, and an edit reverted before the flush becomes none. - **Consistency is easier to reason about at the unit level**: everything the unit touched goes out together. ## The case for explicit marking - **The write is in the diff.** A reviewer can see which columns a change touches without simulating the call graph. - **No incidental writes.** A normalising or defaulting assignment on a read path cannot become an UPDATE, because nothing writes without being asked. - **Column-level precision.** When two workflows routinely touch different columns of the same row, naming the columns lets both succeed where a wide update would clobber. - **Cheaper reads by default**, since a layer that does not diff usually does not keep original values either. - **The failure mode is loud enough to test.** A missing write shows up as data that did not change, which an integration test asserts directly. ## Where automatic detection hurts most | situation | what goes wrong | |---|---| | wide rows with many columns | the compare is expensive and a write may carry columns nobody touched | | rows two workflows edit concurrently | a wide update reverts the other writer's untouched column, or trips its version check | | read paths that normalise on load | phantom updates: locks taken and version columns moved on a page meant to read | | audit or history triggers on the table | every incidental write becomes a permanent history row | | very large tracked sets | scan cost and retention dominate the work the path actually does | ## The column-width question that only shows up under concurrency Even within automatic detection there is a second decision: does the UPDATE carry only the differing columns or every mapped column? Writing all columns is simpler for the layer and lets statements be reused, but it sends the values this unit loaded for columns it never touched — so a concurrent writer's edit to one of those columns is reverted, unless a version predicate rejects the write first. Writing only differing columns avoids that but produces more distinct statements. If your rows are edited by more than one workflow at a time, this is not a detail; it decides whether last-writer-wins applies to fields nobody meant to write. ## How to actually decide 1. **Decide per path, not per codebase.** A rich write path with real invariants is the best case for detection; a bulk maintenance job or a hot single-column counter is the best case for an explicit statement. 2. **Make reads untracked by default.** Most of the pain attributed to detection is really reads paying for a write mechanism they never use. 3. **Ask what your team can detect.** Detection fails silently in the direction of *too many* writes, which observability can catch by counting statements per request. Explicit marking fails silently in the direction of *too few*, which only an assertion on the persisted data catches. Pick the failure your tooling actually sees. 4. **Consider who edits the model.** The more people touch the mapped classes, and the more library code they pass through, the more surprising an invisible write set becomes. 5. **Write down the boundary.** A rule like *mapped objects are loaded writable only inside a write service; every other read is untracked* is enforceable in review and in architecture tests. A vague preference is not. ## What a good answer sounds like Not "detection is convenient" or "explicit is safer", but a stated boundary with the failure mode of each side named, a mechanism for detecting that failure, and an acknowledgement that the same codebase can and usually should run both — because the two failure modes are opposites, and choosing one everywhere means accepting its failure everywhere.

  • Which failure mode is easier to detect in production, and how?
    Too many writes. Counting statements per request or per endpoint makes incidental updates visible immediately, and the count can be asserted in tests. Too few writes — a forgotten explicit mark — produces no signal at all except data that did not change, so it must be caught by assertions on persisted state rather than by observability.
  • How would you enforce that read paths never load writable mapped objects?
    Put the writable loader behind a narrow write-side component and let everything else obtain either untracked results or transfer shapes with no setters. Then assert the dependency direction with an architecture test so that a read-side package cannot reach the writable loader, which turns a review convention into a build failure.
  • Does mixing both mechanisms in one codebase confuse people more than it helps?
    Only if the boundary is implicit. Mixing is fine when the two live in visibly different places — a write service that loads writable objects, and read code that gets transfer shapes — because the type a caller holds tells them which rules apply. It becomes confusing when the same loader returns tracked results on some paths and untracked ones on others.

saying these in an interview costs you the question

  • Picks one mechanism for the whole codebase without naming its failure mode
  • Thinks automatic detection cannot produce writes nobody intended
  • Assumes a forgotten explicit mark surfaces as an error rather than silently
  • Ignores that writing every column can revert a concurrent writer's untouched column
  • Treats read paths as free under detection because they do not edit anything
  • Argues from developer convenience alone with no operational evidence