Moving a service off a tracking mapper onto a query builder, which facilities must the team now provide by hand?
answer
- the work moves, it does not vanish
- no registry of loaded rows
- order of calls is the write order
- one call, one round trip
- row count is the conflict signal
basics
~20 sIdentity, write ordering, batching into fewer round trips, cascade to related rows, the optimistic version check with its row-count test, and any cache invalidation. None of it disappears; it moves from the layer into application code.
solid answer
~40 sEverything the tracked set inferred becomes explicit. There is no identity map, so one row loaded twice is two objects and the later write wins. There is no flush to derive order from, so parents before children is now a property of your call order. There is no grouping, so a per-element loop costs a round trip per element unless you reach for multi-row execution. Cascade, the version predicate in the `WHERE` clause, and the check on the affected-row count are all hand-written. What you get back is that the statements the database runs are the statements in the source — nothing fires that is not on the page.
code
sql · 9 linesUPDATE orders
SET status = ?,
version = version + 1
WHERE id = ?
AND version = ?;
-- zero rows affected means another writer got there first;
-- the untracked layer will not raise for you, so the caller
-- must inspect the count and decide what to do.go deeper
Learn the list itself: identity, ordering, batching, cascade, the concurrency check, invalidation. Knowing what a mapper was doing is the point of the question.
Explain the mechanism behind each item — why no before-image means no inferred statement, why no flush means no natural grouping point.
Bring the operational failures: the half-done version check that reports false success, and the loop that turns one flush into thousands of round trips.
Argue where the trade pays. Predictable statements are worth real money on hot and bulk paths; on a graph-shaped write model the hand-written ordering is a recurring tax.
## What the tracked set was quietly providing A mapper with a tracked set does a surprising amount of work between *you mutate an object* and *the database sees a statement*. It registers each loaded object, keeps a before-image, compares them, decides which statements to emit, decides the order to emit them in, groups similar ones, applies the concurrency check, and follows configured cascades to related rows. Drop the tracked set for a query builder or a hand-rolled row-mapping layer and none of that disappears as a *requirement*; it moves from the layer into your code. The interview question is whether you can list what moved. ## The list you now own 1. **Identity.** With no identity map, one row loaded through two code paths becomes two objects. They can be edited independently and the later write wins wholesale, so a request that loads the same aggregate twice can lose one path's edit entirely. You either thread a single loaded instance through the call chain or you accept that last-write-wins is the semantics. 2. **Write ordering.** A parent must exist before its children reference it, and children must go before a parent is deleted. A mapper derives that order from its mapping metadata. In an untracked layer, the order of your calls *is* the order, so ordering becomes a property of control flow — and control flow is easy to reorganise without noticing you broke it. 3. **Batching.** A mapper can hold many pending statements and send them together at flush. Issuing statements one call at a time gives one round trip each, which is invisible on ten rows and dominant on ten thousand. You have to reach for the layer's multi-row execution deliberately. 4. **Cascade.** Saving or deleting a graph is now several statements, written out, in the right order, at every call site that saves that graph. 5. **The concurrency check.** A mapper can add a version predicate to the update's WHERE clause and raise when no row matched. You write both halves yourself: the predicate, and the check on the affected-row count. Skipping the second half is the common half-done version of this. 6. **Invalidation.** A mapper's tracked set is at least self-consistent within one unit of work. Any cache you add on top of an untracked layer is yours to invalidate, and there is no flush event to hang the invalidation on. ## What the layer gives back | Mechanism | Under a tracked set | Under an untracked layer | |---|---|---| | Change detection | Inferred from a snapshot | Stated as a statement | | Statement order | Derived from mapping metadata | The order of your calls | | Round trips | Grouped at flush | One per call unless grouped | | Concurrency check | Often applied automatically | Version predicate plus row-count check, by hand | | Related rows | Cascade configuration | Explicit statements per call site | | What runs | Sometimes surprising | Exactly what is written | The last row is the reason teams take the trade. Nothing fires that is not on the page: no fetch behind a property read, no flush before a query, no statement whose source is a configuration file. On a hot read path or a bulk write path, that predictability is worth more than the conveniences given up. ## Where teams actually get burned - **The half-done optimistic check.** The version column is in the WHERE clause but nobody looks at the affected-row count, so a conflicting write silently does nothing and the user is told it succeeded. - **Round-trip amplification.** A loop that calls the layer per element replaces one flush with ten thousand round trips. This is the single most common performance regression after the move. - **Order-of-call fragility.** Extracting a helper reorders two writes, and a foreign key fails in production under data that the test fixtures never produced. - **Two copies of one row.** Two services in one request each load, edit and save the same row; the second overwrites the first's column with the value it read before the edit. ## How to keep it sane Concentrate every statement for one table behind a small number of functions rather than spreading them across services, so ordering, the version predicate and the row-count check exist in one place each. Make multi-row writes the default shape for anything that runs in a loop. And write tests that re-read the row after the call instead of asserting on the object in hand — under an untracked layer, the in-memory object proves nothing about what was written.
- Why is the affected-row count the critical half of a hand-written optimistic check?The version predicate only stops the wrong row from being written; it reports nothing. If no row matched, the statement succeeds and updates zero rows. Without inspecting the count the code cannot distinguish a conflict from a normal write, so it reports success while the user's edit was dropped.
- What makes write ordering more fragile than it looks?The order is your control flow, not declared metadata, so any refactor that extracts, reorders or parallelises calls can break a foreign-key dependency. Tests with small fixtures often insert in an order that happens to work, and the failure appears only under data where the parent is genuinely new.
- Does losing the identity map matter for read-only paths?Much less. Duplicate objects only conflict when both are written, so a report or export path loses little beyond some memory to duplicate rows. It bites on write paths where two parts of one request load, edit and save the same row independently.
saying these in an interview costs you the question
- Says the tracked set's work simply disappears with it
- Adds a version predicate but never checks the affected-row count
- Writes per-element statements in a loop and calls it equivalent
- Assumes the database preserves an ordering the code did not send
- Thinks two loads of one row stay consistent with each other