For a large codebase, when is one generic save the right default, and when should writes be explicit inserts and updates?
answer
- who answers does this row exist
- key ownership decides most cases
- inference costs a probe or a guess
- creation with meaning wants loud failure
- constraints outlive conventions
basics
~20 sDefault to a generic save where the store assigns keys and writes are ordinary edits; require explicit inserts and updates where the application assigns keys, writes are high volume, or creating an existing thing must fail.
solid answer
~50 sThe question is who answers "does this row exist" - the layer, by inference, or the caller, by saying so. A generic save is the right default when the store generates keys, the objects are straightforward, and the cost of a wrong guess is low: it removes noise and keeps write code short. Explicit calls earn their extra typing where the inference is unreliable or the failure is expensive - **application-assigned keys**, which defeat the unset-key test; **write-heavy paths**, where a probe read per object is real latency; and **paths with intent**, where creating something that already exists is a business error that must raise rather than quietly become an update. Most large codebases end up mixed, and the thing worth standardising is not the call but the rule for which paths use which - plus the store constraints that hold either way.
go deeper
The point to take away is that one save call and separate insert and update calls are a real choice, not two spellings of the same thing. The difference is whether the caller or the layer decides that the row already exists.
Argue the tradeoff concretely: what the inference costs, where it breaks, and what an explicit call buys in failure behaviour. Naming application-assigned keys as the case that settles it shows you have met the problem.
Answer with a mixed policy and its boundary - generic save for ordinary aggregate writes, explicit statements on ingest and on creation paths that carry meaning - and say how you would verify the boundary holds.
Reason about what survives turnover. Conventions decay, so put the guarantee in the store's constraints and the boundary in the code structure, and judge the policy by which failure mode your organisation absorbs better.
This is a policy question, not a correctness one: both styles work. What differs is who carries the risk of the insert-or-update decision, and how a mistake shows up. ## The two positions, honestly stated **The generic save** says: the layer already tracks state, so let it infer whether this object is new. Write code shrinks to one verb, the same repository method serves creation and edit, and nobody forgets which one to call. **Explicit insert and update** say: intent belongs to the caller. The statement emitted is the statement written, an insert against an existing key fails loudly, and there is no inference to be wrong about or to become wrong when the key strategy changes. Neither is a maturity level, and neither is the "advanced" answer in an interview. A codebase can hold both, and the interesting judgment is where the line between them falls rather than which side wins outright. ## What actually decides it | Factor | Points to generic save | Points to explicit writes | |---|---|---| | Who assigns keys | the store generates them, so unset means new | the application assigns them, so the free inference is broken | | Write volume | occasional writes, latency irrelevant | hot paths where a probe read per object is measurable | | Semantics of creation | create-or-update is genuinely what the domain wants | creating an existing thing is an error someone must see | | Team and codebase size | small, one convention easily held in heads | large, where a silent no-op crosses many hands before anyone notices | | Idempotency needs | replaying a request should converge on one row | replay must be rejected, not absorbed | The key-assignment row is the one that most often settles the argument by itself. Application-assigned identifiers are chosen deliberately - so an object is complete before it is stored, or so an identifier can be generated without a round trip - and they make the cheapest inference unusable. A codebase that made that choice and then leaned on a generic save is asking the layer a question it cannot answer. ## A workable mixed policy 1. **Generic save as the default for aggregate writes** through repository methods, where the store owns key generation. 2. **Explicit insert on creation paths that carry meaning** - registration, provisioning, anything where "it already exists" is a distinct outcome the caller must handle. 3. **Explicit statements for volume work** - imports and ingest paths - where inference costs a probe and batching is the point. 4. **Store constraints regardless.** Unique constraints on natural keys are the guarantee that does not depend on any of the above being followed correctly. Point four is the one that survives turnover. Conventions decay; a constraint does not. ## What the choice costs Standardising on explicit writes costs *fluency*: more call sites, a caller who must know which state the object is in, and some duplication between create and edit paths. Standardising on the generic save costs *legibility of failure*: a wrong guess is a silent no-op or a duplicate row rather than an exception, and the person who eventually debugs it has no line of code to look at, only an inference. There is a second cost that is easy to miss on both sides: the choice leaks into testing. Inference-based writes need tests that cover the new-object and existing-object branches separately, because a single happy-path test passes under either verdict. Explicit writes need tests that the right call was made on the right path, which is easier to see but easier to duplicate. Ask which of those your organisation is better at absorbing. Teams whose main risk is a rarely-exercised path failing quietly should push toward explicit intent. Teams whose main risk is inconsistent write code across many services benefit more from one convention held everywhere. ## How to make the policy real A written rule that nobody can check is not a policy. Make the boundary structural: keep the ingest path in its own component with its own write API, so a reviewer can see which side a change is on. Where the codebase forbids the generic call on a given path, prefer a repository surface that does not expose it there at all - a narrow interface enforces more than a convention does. And accept a mixed outcome as the goal, not a failure: the useful artefact is a short list of paths where inference is banned and why, not a global ban that the next urgent change quietly breaks.
- Which single fact about a codebase most often settles this argument?Who assigns identifiers. When the store generates keys, the unset-key inference is free and reliable, and a generic save is a reasonable default. When the application assigns them, the object is complete before it is saved, the inference reads every new object as existing, and the layer is being asked a question it has no evidence to answer - so intent has to come from the caller.
- If you standardise on explicit writes, what do you lose?Fluency and uniformity. Callers must know which state an object is in, create and edit paths diverge, and shared repository methods split in two. In a large codebase that also means more surface for inconsistency between services. The gain is that every wrong assumption becomes an exception with a line number instead of a missing row discovered weeks later.
- Why is a unique constraint part of this decision rather than a separate concern?Because it is the only part that holds when the policy is not followed. Whichever style a path uses, a duplicate created by a wrong inference, a race between two probes, or a replayed request is caught by the constraint and by nothing else. It converts the worst outcome - a silently duplicated row - into a rejected statement, independently of any convention.
- How would you keep such a policy alive as the team changes?Make it structural rather than documentary. Put the paths that must not infer behind a narrow write interface that does not offer the generic call, keep ingest code in its own component so a reviewer can see which rule applies, and keep the written rule to a short list of banned paths with reasons. A rule nobody can check is one urgent change away from gone.
saying these in an interview costs you the question
- Claims explicit inserts and updates are always the safer default
- Says the generic save is fine because the layer figures it out
- Picks a policy without asking who assigns the keys
- Treats it as style rather than a question about failure shape
- Relies on convention alone with no constraint behind it