skip to content

You are deciding where to put JPA @Version fields across a domain model. What does versioning every entity cost, and how would you decide which ones actually need it?

level: principalimportance: nice to knowfreq 28%

answer

  1. Version = row is the unit of conflict
  2. Wide entity + disjoint writers = false conflicts
  3. Hot row: retries amplify, throughput inverts
  4. Counters -> atomic delta or sharded rows
  5. Bulk/native writes bypass the version

basics

~20 s

A version column makes conflict detection row-shaped: any two writers to the same row conflict, even on unrelated fields. Version entities with genuinely concurrent writers and lost-update risk. For hot counters and aggregates, change the model — atomic SQL updates or split rows — rather than adding a version and retrying.

solid answer

~60 s

Versioning is cheap per row and expensive per **contention point**, so I decide by write pattern, not by habit. Costs of blanket versioning: - **False conflicts.** The unit of conflict is the row. Two users editing unrelated columns of a wide entity collide, and users see failures with no semantic cause. - **Retry amplification.** On hot rows, every loser retries, increasing contention — throughput can fall as load rises. - **Contract weight.** Once a version is exposed to clients (form field, ETag) it becomes API surface. - **Blind spots.** Bulk JPQL and native updates do not maintain it, so a batch job can silently invalidate the guarantee people believe they have. Where I do put it: entities with multiple concurrent writers where a lost update is a real business error — money, stock, status transitions, anything edited through long-lived forms. Where I do not: single-writer or import-owned tables, append-only rows, and hot aggregates — those get `update ... set n = n + ?`, per-partition rows, or event-and-project designs. And if a wide entity conflicts falsely, I split the hot fields into their own row.

go deeper

for a junior

Know that a version column protects a whole row and that not every table needs one — tables with a single writer gain nothing.

for a middle

Explain false conflicts on wide entities and name a non-conflicting alternative for counters, plus the bulk-update blind spot.

for a senior

Reason about contention quantitatively — retry amplification, latency tails, which rows are hot — and propose splitting entities or changing the write operation rather than tuning retries.

for a principal

Own the policy: versioning declares the row as the unit of conflict, so decide per aggregate, design hot paths to be conflict-free, document which writers bypass the ORM, and instrument conflict rates as a first-class signal.

## The unit of conflict is the row A `@Version` column says: *any* change to this row invalidates *any* other in-flight change to it. That is a design decision about granularity, and it is the whole story of when versioning helps or hurts. If a row represents one cohesive thing that one actor changes at a time, row-level granularity matches reality. If a row is a wide record whose columns belong to different workflows — a `Customer` holding marketing preferences, credit limit, address and login timestamps — then two unrelated jobs touching different columns conflict for no semantic reason. Users experience random failures; engineers add retries; the retries hide the mismatch. ## The costs of versioning everything **False conflicts.** As above: contention that does not correspond to any real business conflict. It grows with entity width and with the number of independent writers per row. **Retry amplification.** Under optimistic control the loser wastes a whole transaction. On a contended row, more concurrency produces more losers, each retrying, each consuming a connection and queries. Beyond a threshold throughput *decreases* as load increases — a congestive collapse that looks like a database problem but is an application-design problem. **Latency tails.** Even with modest contention, retried requests sit at the tail of the distribution. If a hot row sits on the critical path of a common request, p99 tracks contention rather than work. **Contract weight.** The moment a version travels to clients — hidden field, ETag, mobile cache — it is part of your external contract. Changing its type or removing it later breaks clients. **Silent blind spots.** Bulk JPQL, native SQL, imports, and other services writing the same tables do not maintain the version unless explicitly written (Hibernate's HQL `update versioned ...` form does). A team can believe a table is protected while the nightly job walks straight past the mechanism. **Schema and migration cost.** Adding a version to an existing table needs a default and a backfill; multiply that across every entity and it stops being free. ## Where versioning earns its place - Rows where a **lost update is a business error**: balances, inventory, invoice totals, entitlement state. - Entities edited through **long conversations** — a form open for minutes — where the version is the only thing spanning the gap between requests. - **State machines**, where two concurrent transitions must not both apply; the version prevents `PAID` and `CANCELLED` both landing on top of `PENDING`. - Aggregates whose invariants span child rows, where you deliberately want any child change to invalidate a concurrent parent-level decision (this needs the increment to be forced explicitly, since inverse-side child changes do not move the parent's version by themselves). ## Where something else is better - **Counters and totals.** A single row incremented by everyone is the worst case for optimistic control. `update stats set n = n + ? where id = ?` is atomic, never conflicts, and needs no retry. If you must read-modify-write, shard the counter across N rows and sum on read. - **Append-only facts.** Insert events and project the aggregate; inserts do not conflict. - **Single-writer tables.** Import- or job-owned tables where concurrency does not exist gain nothing but a column and a migration. - **Immutable rows.** Nothing to lose. - **Wide, multi-workflow entities.** Split the frequently-mutated fields into their own entity so conflict granularity matches the workflows. This is usually the highest-value change: it converts contention into non-contention rather than managing it. ## The decision procedure I use 1. **Who writes this row, and how often concurrently?** No concurrency, no version. 2. **If two writers overlap, is the outcome wrong?** If last-write-wins is genuinely acceptable — a preference, a cached score — skip it. 3. **Do writers touch disjoint fields?** If yes, and contention is real, split the entity before adding a version. 4. **Is the row hot?** Then choose an operation that cannot conflict (atomic delta, partitioned rows) rather than a version plus retries. 5. **Does the edit span requests?** Then you need the version *and* a transport for it, plus a conflict UX. ## Operating it afterwards Whatever the choice, instrument it: count conflicts per entity and per operation. A conflict rate that is always zero suggests untested protection or a table with one writer; a rising rate is an early signal that a design assumption expired. And write down which paths bypass the ORM, because those are where the guarantee quietly stops holding. The summary I would give a team: versioning is not a safety blanket to apply uniformly. It is a statement that a row is a unit of conflict. Make that statement where it is true, and change the model where it is not.

  • Why does adding retries to a hot versioned row sometimes make throughput worse?
    Every loser wastes a full transaction and then competes again, so higher concurrency produces more losers and more repeat attempts. Past a threshold the system spends most of its capacity on work that will be discarded, and throughput falls as load rises. The fix is to remove the conflict — an atomic increment, sharded rows, or an append-and-project model — not to tune the retry.
  • How would you handle a wide entity where two teams' jobs conflict on unrelated columns?
    Split the hot or independently-written fields into their own entity with its own row, so conflict granularity matches the workflows. That converts a false conflict into no conflict at all. If splitting is not feasible, versionless optimistic strategies that compare only the changed columns are an alternative, at the cost of longer UPDATE predicates.
  • What paths can silently bypass a version column, and how do you keep track of them?
    Bulk JPQL and HQL updates (unless written as `update versioned`), native SQL, database jobs, and any other service writing the same table. Keep an explicit list of writers per table, prefer the versioned bulk form, and add tests that assert a conflicting write is detected so a bypass shows up as a failing test rather than as a support ticket.

saying these in an interview costs you the question

  • Adding @Version to every entity by reflex and calling it a best practice.
  • Treating a chronically high conflict rate as normal load rather than as a modelling problem.
  • Using a versioned entity for a global counter instead of an atomic increment.
  • Forgetting that bulk and native writes do not maintain the version.
  • Ignoring that exposing the version to clients turns it into part of the API contract.

context