skip to content

Most relational database engines ship READ COMMITTED as the default isolation level. Make the case for that default, and describe how you decide that a particular workload needs something stronger.

level: principalimportance: should knowfreq 40%

answer

  1. Isolation is a spend, not a virtue
  2. Kills dirty reads, no retry contract
  3. Per-statement view = no version bloat
  4. Escalate the transaction, never the system
  5. Write skew and tie-out reports are the real triggers

basics

~20 s

It is the cheapest level that removes the clearly unacceptable anomaly. Reads never block and never abort, version retention stays short, and most OLTP work is single-statement anyway. Escalate only where a real invariant spans statements and constraints or row locks cannot express it.

solid answer

~60 s

The case for the default rests on three properties. **It removes the anomaly nobody can tolerate** — dirty reads — and stops there. **It has no retry contract**: reads take no lasting locks and produce no serialization failures, so ordinary application code needs no retry loop, which matters enormously for a default that unaware developers inherit. **It is operationally cheap**: per-statement snapshots mean no long-lived view pinning old row versions, so cleanup keeps up and a long report does not bloat storage. It also fits the workload. Most OLTP writes are single statements, and a single statement is internally consistent at this level; most remaining invariants are expressible as declarative constraints, which hold regardless of level. I escalate when an invariant genuinely spans statements and cannot be pushed down: multi-row consistency checks, write skew on a resource-counting rule, or reports that must tie out across queries. And I escalate the *transaction*, not the system — narrowly, with a bounded retry loop and idempotent work, and I measure the abort rate afterwards.

go deeper

for a junior

Say that READ COMMITTED is the default because it is cheap and blocks little, and that stronger levels cost concurrency.

for a middle

Add that a single statement is consistent at this level and that most OLTP work fits that shape, with constraints covering much of the rest.

for a senior

Bring in the retry contract of stronger levels, version-retention and cleanup effects, and a concrete escalation checklist for the paths that need it.

for a principal

Argue defaults as a bet about median workload and median developer, quantify the cost of each escalation, and lay out an ordered decision procedure ending in per-transaction escalation with retries, idempotency and abort-rate observability.

## Framing the decision Isolation is a spend, not a virtue. Every step up buys the elimination of some anomaly and charges in one or more of: blocking, aborted transactions requiring retries, version retention, and developer obligation. A default level is a bet about what the median application needs and what the median developer will handle without being told. ## Why READ COMMITTED wins that bet **It removes the one anomaly with no defensible use.** Reading data that may be rolled back can turn a transient state into a permanent business decision — an email sent about an order that never existed, a downstream record keyed to a row that vanished. There is no ordinary application that wants this. Everything READ COMMITTED still permits, by contrast, has a shape that competent code can handle locally. **It imposes no retry contract.** This is the underrated argument. Stronger snapshot-based levels reject conflicting transactions with serialization failures, which is correct behaviour but *only* if the application retries — and retry loops must be written, must be bounded, and must wrap work that is safe to re-run. A default that silently requires every caller to implement retry semantics will produce user-visible errors in every codebase that did not read the manual. READ COMMITTED's reads never fail this way; the only waiting is on write locks, which is intuitive and observable. **Its operational profile is benign.** Because the read view is per statement, no long-running transaction pins an old view of the database. In MVCC engines this is the difference between routine version cleanup and unbounded bloat: a long-lived transaction-scoped snapshot forces the engine to retain every row version that snapshot might still need. Analytics-style queries and idle-in-transaction sessions are common in the wild; a default that punishes them with storage growth and cleanup stalls is a bad default. **It matches the workload shape.** OLTP writes are overwhelmingly single statements against few rows, and a single statement is internally consistent even here. Most of the remaining correctness surface is covered by mechanisms independent of isolation: unique indexes, foreign keys, check constraints, and the write-lock-plus-recheck behaviour that makes `SET x = x - 1 WHERE x >= 1` safe. The residual gap is smaller than it first appears. **Concurrency is high and predictable.** Readers do not block writers, writers do not block readers, and the only contention is on the rows genuinely being written. Latency is stable and tail behaviour is easy to reason about. ## What the default actually costs The bill is paid in developer obligation, and it is a real bill: - Read-modify-write across round trips must be locked or version-checked, or updates are lost silently. - Multi-query reports do not tie out; the numbers straddle commits. - Write skew — two transactions each validating an invariant over rows the other is about to change — is invisible and unreported. The failure mode of these is *silent wrong data*, not an error, which is exactly the failure mode that survives testing and surfaces in production months later. That is the honest counter-argument to the default and should be stated when making the case. ## When I escalate My decision procedure, in order: 1. **Can the invariant be expressed as a declarative constraint?** Unique, foreign key, check, exclusion. If yes, do that — it holds at every level, needs no discipline from callers, and cannot be forgotten by the next person. 2. **Can the write be a single statement whose guard lives in the `WHERE` clause?** If yes, do that, and branch on the affected-row count. This covers most counters, inventory decrements, and state-machine transitions. 3. **Is it a read-modify-write across round trips?** Use an explicit row lock inside a short transaction, or an optimistic version column with a bounded retry when the gap spans requests. 4. **Does the invariant span rows the transaction only *reads*?** "At most N active X", "at least one on-call Y", "the sum across these rows stays non-negative". This is write skew and none of the above catches it. Now escalate — or materialise the invariant into a single lockable row so step 2 or 3 applies again. 5. **Do multiple queries have to describe the same instant?** Financial reports, exports, reconciliations. Escalate that transaction to a transaction-scoped view. ## How I escalate **Per transaction, not per system.** Raising the global default changes behaviour for every code path, including ones nobody has reviewed, and moves failures from silent-and-rare to loud-and-everywhere without anyone having written a retry. The checklist for an escalated path: - a bounded retry loop around the whole transaction, with jitter; - work that is safe to re-run — no side effects emitted before commit, no non-idempotent external calls inside the transaction; - the transaction kept short, because a longer transaction-scoped snapshot means both a bigger conflict window and more retained versions; - observability on abort rate, retry counts, and lock waits, since escalation converts silent corruption into visible failures and you need to see them; - a documented reason on the code path, so it is not "simplified" back later. ## The position to argue READ COMMITTED is the right default because defaults should be safe against the anomaly with no legitimate use, cheap to run, and forgiving to code written by people who never thought about isolation. It is *not* a claim that it is sufficient everywhere. The engineering work is knowing which handful of paths in a system carry a genuine cross-statement invariant, handling those explicitly, and refusing to pay the escalation cost on the other ninety-odd percent.

  • What is the strongest argument against READ COMMITTED as a default?
    Its failure mode is silent. Lost updates, read skew and write skew produce wrong data with no error, so they pass tests, survive code review, and surface in production as unexplained inconsistencies long after the fact. A stronger default converts those into loud serialization failures, which are noisier but far easier to find and fix.
  • Why raise the level per transaction rather than globally?
    A global change affects every code path, including ones nobody audited, and stronger levels come with a retry contract that existing code does not implement — so you convert silent rarity into visible, widespread errors. Escalating a specific transaction confines both the cost and the required retry discipline to the path that actually needs it.
  • How would you handle an invariant such as 'at most three active sessions per user' without raising the isolation level?
    Materialise the invariant so a single statement or a single lock can enforce it: keep a counter column on the user row and update it with a guarded single statement, or take an explicit lock on the user row before counting and inserting. Alternatively express it declaratively where the engine supports it. Any of these converts a multi-row read-then-write check into something the level can already protect.

saying these in an interview costs you the question

  • Recommending SERIALIZABLE everywhere as the safe default without acknowledging retry and throughput costs
  • Ignoring that stronger levels require an application retry loop to be correct
  • Claiming READ COMMITTED is 'good enough' without naming the invariants it fails to protect
  • Overlooking that declarative constraints hold at every level and often remove the need to escalate
  • Treating the isolation level as the only concurrency-control tool available

context