How does one generic save call in a data-access layer decide whether to insert a new row or update an existing one?
answer
- the caller left a question unanswered
- unset key, absent version, probe read
- free guess versus a round trip
- application-assigned keys break the guess
- explicit insert fails loudly instead
basics
~20 sIt infers newness from a signal: an unset generated key, an absent version value, or a probe read for the row. Each signal can guess wrong or costs a round trip, which is why explicit insert calls exist.
solid answer
~50 sA generic save has to answer a question the caller did not: does this object already have a row? Layers infer it three ways. **Unset key** - if the key field is null or zero and keys are store-generated, treat it as new; free, but wrong whenever the application assigns keys itself. **Absent version** - a mapped version column that is null on a fresh object marks it new; also free, but only where a version is mapped. **Probe read** - select the row by key and branch on the result; always correct at the moment it runs, but costs a round trip per save and can still race two concurrent inserts. Explicit insert and update calls skip the inference entirely: the caller states intent, a wrong assumption fails loudly, and there is nothing to probe.
go deeper
Remember that one save call hides a decision. Learn the three signals - unset key, absent version, probe read - and that a probe costs an extra trip to the store while the other two are free guesses.
Explain where each signal breaks and what the wrong branch does. The application-assigned key turning every save into an update that matches nothing is the case interviewers reach for most.
Talk about failure shape. An inferred insert on a store-generated key duplicates a row and raises nothing, so argue for the unique constraint and for explicit intent on write-heavy or import paths.
Own the policy rather than the call. Decide who assigns keys, whether inference is permitted at all in this codebase, and what the store must enforce regardless - constraints are the only guarantee that survives a wrong guess.
One save call, two possible statements. A generic save is a convenience that asks the layer to answer a question the caller did not answer: **does a row for this object already exist?** Every mechanism for answering it is either a guess from local evidence or a read. ## The three signals layers use | Signal | How it decides | Cost | Where it fails | |---|---|---|---| | Unset generated key | key is null or zero, so treat the object as new | free, no round trip | the application assigns keys, so a brand-new object already carries one and is read as existing | | Absent version value | a mapped version column with no value marks a fresh object | free, no round trip | no version is mapped, or the field cannot hold "absent" and its default collides with a real version | | Probe read | select the row by key, then branch on hit or miss | one extra round trip per save | two concurrent writers can both miss and both insert | Layers combine these - a version check where a version exists, an unset-key check otherwise, a probe as the fallback - and some let the caller override the verdict with a flag on the object or a hint on the call. ## Why the unset-key test breaks The unset-key test assumes the **store** hands out keys. Plenty of designs do not: the application assigns a generated identifier or uses a natural key, precisely so that the object is complete before it is ever saved. Now the key is set from the first instant, the inference says "existing", and the save emits an update. Two outcomes follow, both bad: - the update matches **zero rows**. Some layers detect the empty result and raise; others simply return, and the write is silently lost. - the layer falls back to a probe read, and the "free" inference quietly became a round trip on every save in the system. The remedies are all about restoring evidence the layer can trust: map a version column that is genuinely absent on new objects, carry an explicit is-new flag, or - simplest - stop using the generic call on that path and say insert. ## Why an explicit insert is not just an update in disguise Stating intent buys three things a guess cannot: 1. **Loud failure.** An insert against an existing key violates the key constraint and raises. An inferred update against a missing row often reports nothing at all. Failing is better than losing a write. 2. **No probe.** The layer does not need to read before writing, which matters most exactly where saves are frequent. 3. **Better batching.** A known statement shape can be grouped with others of the same shape; a per-object guess that may resolve either way is harder to group. The symmetrical point holds for the update side: an explicit update expresses "this row exists and I am changing it", so a zero-row result is an anomaly worth surfacing rather than a normal outcome. ## What a wrong guess actually costs - **Inferred update on a new object**: zero rows affected. The best case is an exception; the common case is a silent no-op that surfaces later as missing data. - **Inferred insert on an existing object**: a key-constraint violation, which is loud and therefore fine - *unless* the key is store-generated, in which case the insert succeeds and you now have **two rows** describing one thing. This is the more dangerous outcome precisely because nothing fails. - **A probe that was correct and then stopped being correct**: between the probe and the write, another writer can insert the same row. The probe narrows the window; it does not close it. Only a unique constraint in the store closes it, which is why the constraint should exist even when the layer probes. ## Where this decision does not exist A layer with no tracked set never faces the question: the code says insert or says update, and the statement it emits is the statement it was given. The tradeoff is symmetrical - more typing, no inference to be wrong about. Some codebases deliberately mix, using the generic save for straightforward object-per-row work and dropping to explicit statements on the paths where an insert must be an insert.
- Your objects carry application-assigned identifiers. What breaks in a generic save, and what would you change?The unset-key test always reports the object as existing, so the save emits an update that matches zero rows - either an error or a silently lost write. Give the layer better evidence or take the decision away from it: map a version column that is genuinely absent on new objects, carry an explicit new flag, or call insert directly on the creation path and keep the generic save for edits.
- A save infers insert on an object that already has a row. Is that a safe failure?Only when the key is the one being inserted, because the key constraint then raises and nothing is corrupted. If the key is store-generated, the insert succeeds with a fresh key and you have two rows for one thing, with no error anywhere. That silent duplication is the outcome to design against, usually with a unique constraint on the natural key.
- How much does the probe-read strategy really cost in a write-heavy path?One extra round trip per saved object, which is invisible on a single save and dominant on a batch: the probes cannot be batched with the writes and they serialise against network latency. It also does not remove the race, so the unique constraint is still required. Where writes are frequent, stating insert or update is both cheaper and stricter.
saying these in an interview costs you the question
- Says the layer always reads the row before deciding
- Thinks a null-key test works with application-assigned keys
- Believes a generic save is always cheaper than an explicit insert
- Says an inferred update on a missing row always raises
- Assumes a probe read removes the concurrent-insert race