How do you decide which rows in a model carry a version, and what does the granularity of that choice cost?
answer
- granularity is the whole decision
- coarse guards rules, fine avoids collisions
- no concurrent editing, no guard needed
- versionless compares old column values
- conflict rate is the feedback signal
basics
~20 sPut versions where concurrent editing is real and where an invariant needs protecting. A version on a parent guards rules spanning its children but makes independent child edits collide; versions only on children collide less and guard nothing across rows.
solid answer
~50 sGranularity is the whole decision. A version on a **root row** means any change anywhere under it conflicts with any other - which is exactly right when the rules being protected span several rows, and needlessly painful when two people edit unrelated children. Versions on **individual rows** give the opposite: near-zero false conflicts, and no protection for anything that spans them. Rows nobody edits concurrently need no version at all; the column is not free in review attention even if it is cheap in bytes. Some layers also offer a **versionless** variant that compares the previously loaded values of all columns, or only the changed ones, which needs no schema change but is fragile against imprecise types and, in the changed-columns form, lets two disjoint edits both succeed. Treat the conflict rate per entity as the feedback signal.
go deeper
A version is worth adding where two people can really edit the same record at once. Rows written once and only read afterwards do not need one.
Contrast a version on a parent with versions on each child: the first collides on any change under the parent, the second protects nothing that spans them.
Diagnose from evidence - false conflicts mean too coarse, invariants breaking only under load mean too fine - and know why comparing old column values is a fallback rather than a default.
Own it as a contract question: placement decides what tokens clients echo, the every-writer-advances rule has to hold across teams, and the right granularity follows the workload, so make it measurable.
## The question is not whether, but where Adding a version column is cheap; adding it everywhere by reflex is still a decision worth making consciously, because each one changes how two concurrent writers interact. Three questions decide it per row type: 1. **Is this row edited concurrently in practice?** Reference data written once by an import and read forever does not need a guard. 2. **Is the state that must stay consistent contained in this row?** If yes, a version here is sufficient. 3. **Does a rule span several rows?** Then no per-row version protects it, and the granularity question begins. ## Root versus child When an object is really a small cluster - an order with its lines, a document with its sections - the version can sit on the root, on each row, or on both. | Placement | What it protects | What it costs | |---|---|---| | Root only | Rules that span the cluster; any concurrent change under the root collides | False conflicts between genuinely independent edits | | Each row | Exactly that row's contents | Nothing spanning rows; two valid child writes can be jointly invalid | | Root plus rows | Both, if the root is advanced deliberately on child writes | Most contention, and the discipline to remember the root every time | The honest framing is that a coarse version is a **deliberate serialisation**: it makes concurrent editors of one cluster take turns. That is a feature when the cluster has invariants, and pure friction when it does not. The failure to avoid is picking coarse by default and then discovering the conflict rate on the busiest cluster in the system. ## The versionless variant Some layers can guard a write without a dedicated column by putting previously loaded values into the `WHERE` clause - either all mapped columns, or only the ones being changed. It is the fallback for tables whose schema you cannot alter. - **No schema change**, which is sometimes the entire argument for it. - **Wide predicates.** Every column in the comparison, on every write, and the statement no longer reads as a simple keyed update. - **Type fragility.** Approximate numbers, large text, and values the engine normalises on storage may not compare equal to what was read back, producing conflicts nobody can explain. - **Requires the full prior state.** A partially loaded row cannot supply what the comparison needs. - **Changed-columns-only is weaker still.** Two writers editing disjoint fields both succeed - fine when the fields are independent, wrong when a rule ties them together, which is precisely the case a guard is for. A dedicated column avoids all of this for the price of one small field, which is why it is the default whenever the schema is yours. ## Reading the feedback Granularity is a hypothesis, and production answers it: - **False conflicts dominate** - users rejected while editing genuinely unrelated things - the version is too coarse. Move it down, or split the cluster. - **Invariant violations appear under load** and cannot be reproduced by hand - the version is too fine, or the rule needs the parent brought into the write deliberately. - **Conflicts cluster on one entity or one screen** - the workflow is the problem, not the guard. A form that holds a record open for twenty minutes will collide no matter how the columns are arranged; shortening the window, or splitting the record so people edit different rows, beats any tuning of the mechanism. Measuring this needs the conflict to be recorded with the row type involved, which is worth building before the argument arrives rather than during it. ## Organisational edges Two things make this a lead's problem rather than a developer's. First, the rule that every writer advances the version has to hold across teams and jobs, and it only holds if it is written down somewhere people read. Second, granularity leaks into the client contract: what the boundary hands out as a token, and therefore what a submission must return, follows the placement. Changing a version from a child to its root later means changing what every client echoes back. ## A defensible default Start with a dedicated version column on rows that are edited by people, place it on the root of any cluster that has rules spanning its rows, leave it off write-once reference data, and use the deliberate parent advance rather than a coarse version where only some child writes touch the shared rule. Then watch the conflict rate and move, because the correct granularity is a property of the workload, not of the model diagram.
- When is comparing previously loaded column values instead of a version column defensible?When the schema is not yours to change. It needs the full prior state of the row, produces wide predicates, and misbehaves with approximate or normalised types; the changed-columns-only form additionally lets two disjoint edits both land. With a schema you control, a dedicated column is smaller and exact.
- How would you tell that a version's granularity is too coarse?Users are rejected while editing unrelated parts of the same cluster, and re-submitting succeeds unchanged. That signature - conflicts whose two edits do not actually interact - means the guard is serialising work that never needed serialising, so the version belongs lower down or the cluster should be split.
- Does a high conflict rate always mean the versioning is wrong?No. It often means the workflow holds records open too long, or that one screen concentrates edits on a single row. Shortening the editing window, splitting the record, or making the edit a delta the server can re-apply usually helps more than rearranging columns.
saying these in an interview costs you the question
- Adds a version to every table without asking who edits concurrently
- Assumes per-row versions protect a rule spanning several rows
- Treats a coarse version as free rather than as deliberate serialisation
- Prefers comparing all old column values on a schema they control
- Uses changed-columns-only comparison where fields are interdependent
- Never measures how often conflicts actually occur