skip to content

You are deciding whether a high-throughput OLTP service should run all of its transactions at SERIALIZABLE. What throughput and abort-rate effects do you expect, and what alternatives would you weigh for the invariants that actually need protection?

level: principalimportance: should knowfreq 38%

answer

  1. cost scales with conflict, not traffic
  2. hot key + long transaction = abort storm or queue
  3. goodput can fall as load rises (wasted retries)
  4. constraints > single statement > explicit lock > targeted level
  5. benchmark with realistic skew; measure abort rate and p99

basics

~20 s

Expect throughput to fall with contention, not uniformly: uncontended paths change little, hot rows and wide scans degrade sharply through waits or aborts, and abort rates rise superlinearly with transaction length. Weigh declarative constraints, single-statement updates, explicit locking and targeted use before adopting it globally.

solid answer

~60 s

Globally enabling SERIALIZABLE is a defensible default, but it must be a measured decision. **What to expect.** Cost scales with conflict, not with traffic. Non-overlapping transactions are barely affected. Overlapping ones pay either blocking (locking implementations) or aborts (optimistic ones), and the abort rate grows superlinearly with contention and with how long each transaction stays open — so a p99 dominated by a few long transactions is the thing that breaks first. Wide read sets from unindexed predicates are the second amplifier, since both mechanisms track what you read. Every application path also needs retry plumbing and idempotent side effects, which is real engineering cost. **Alternatives to weigh.** Push invariants into declarative constraints — a unique index or a check enforces a rule at any isolation level with no schedule analysis. Use explicit locking on a handful of hot paths so conflicts become short waits instead of redone work. Restructure read-modify-write into single statements. Apply SERIALIZABLE selectively to the transactions that carry a multi-row invariant. **How to decide:** benchmark at realistic contention, watch abort rate and p99, and keep the level where the invariant lives.

go deeper

for a junior

Know that stronger isolation costs throughput and produces retries, and that short transactions help.

for a middle

Explain that cost tracks conflict rather than volume, and name transaction length and read-set width as the main levers.

for a senior

Give a measurement plan — realistic skew, abort rate, attempts per commit, p99 and goodput — and the cheaper alternatives per invariant.

for a principal

Own the policy: default level, which invariants are enforced by constraints versus schedule analysis, transaction-length budgets, hot-key strategy, and the limits of a per-database guarantee in a multi-service system.

## Framing the decision Run everything at SERIALIZABLE is attractive because it removes a whole class of reasoning from application developers: no anomaly analysis, no lock-order discipline, no defensive re-reads. That is a genuine engineering saving. The question is what it costs in this workload, and whether cheaper mechanisms already cover the invariants that matter. ## Where the cost actually appears **It scales with conflict, not with volume.** Transactions whose read and write sets do not overlap are effectively unaffected: a locking implementation finds no conflicting lock; an optimistic one records dependencies but never sees a dangerous pattern. A partitioned workload — per-tenant, per-user, per-account rows — can run at SERIALIZABLE almost for free. **Hot keys are where it hurts.** A single row every transaction touches (a global counter, a sequence table, one shared inventory row) becomes the serialization point. Locking makes everyone queue behind it, so throughput on that path approaches one transaction per lock-hold time. Optimistic detection instead lets everyone run and aborts most of them; each abort wastes the work already done, so goodput can fall as offered load rises — a congestion-collapse shape rather than a plateau. **Transaction duration is the multiplier.** Conflict probability rises roughly with the window during which a transaction is exposed. A transaction holding a database transaction open across a remote call, an HTTP retry, or user think-time is orders of magnitude more likely to conflict than a 2 ms one. In most incidents, transaction length — not the isolation level — is the root cause. **Read-set width is the second multiplier.** Both mechanisms key off what a transaction read. An unindexed predicate scans and therefore locks or marks far more than it needed; add the index and both waits and aborts drop. Missing indexes often show up first as abort-rate regressions. **Second-order costs.** Retry logic on every write path; idempotency or an outbox for external side effects; tracking memory and possible granularity escalation in optimistic engines; more lock-manager memory and escalation risk in locking ones; harder capacity modelling, because throughput is no longer a simple function of CPU. ## The alternatives, roughly in order of preference **1. Declarative constraints.** If the invariant is expressible as a unique index, foreign key, check constraint or exclusion constraint, use it. The engine enforces it at every isolation level, with no dependency analysis, no retries and a clear error. Many invariants people reach for SERIALIZABLE to protect — only one active membership, no overlapping bookings — are constraint-shaped. **2. Statement-level atomicity.** Rewrite read-decide-write as one statement that reads and writes the current row (`SET stock = stock - 1 WHERE id = ? AND stock >= 1`, checking the affected row count). This eliminates the window rather than defending it, and it is the single highest-leverage fix for counter and inventory contention. **3. Explicit locking on the few hot paths.** Take a row lock on read, in a consistent order. Conflicts become short, bounded waits instead of repeated wasted work, and no retry loop is needed. Cost: deadlock risk and reduced parallelism on those rows — acceptable for a handful of paths, unmanageable as a global convention. **4. Selective SERIALIZABLE.** Apply the level to the specific transactions that enforce a multi-row invariant no constraint can express — balance checks across rows, capacity checks over a range, approval workflows. Remember the caveat: in optimistic implementations the guarantee only holds among transactions that *all* run at SERIALIZABLE and touch the same data, so selective means per-invariant, not per-statement. **5. Global SERIALIZABLE.** Correct and simple to reason about. Reasonable when the workload is well partitioned, transactions are short, and the team benefits from a uniform rule; a strong default for smaller-scale or correctness-critical systems. ## How to decide, concretely - Benchmark at *realistic contention*, not with uniform random keys — the skew is the whole story. Model the real hot keys. - Measure abort rate, attempts per successful commit, p99 latency and goodput, not just average throughput. - Inventory the invariants: which are constraint-expressible, which are single-statement-expressible, which genuinely need serial-equivalence. - Set and enforce transaction-length budgets and forbid remote calls inside transactions before touching the isolation level at all. - Instrument first: abort rate and retry counts per transaction type must be visible before you can operate the level. - Verify what the engine actually implements under that name, and remember the guarantee is per-database: replicas, caches and cross-service writes sit outside it. ## The judgement There is no universal answer, which is the point of the question. A good answer says: default to the strongest level the workload can afford, but pay for correctness with the cheapest mechanism that expresses the invariant — constraints first, single statements next, targeted locking or targeted SERIALIZABLE after that — and treat transaction length and read-set width as the levers that decide whether global SERIALIZABLE is affordable at all.

  • Your abort rate under SERIALIZABLE jumps from 0.2% to 8% after a release, with no traffic change. Where do you look first?
    At what the transactions read and how long they stay open. The usual causes are a lost or unused index turning a narrow lookup into a scan, which widens the read set every mechanism tracks, and a newly added remote call or extra work inside the transaction that lengthens the conflict window. Also check for a new hot key — a shared counter or status row the release introduced.
  • Is running everything at SERIALIZABLE enough to guarantee correctness across a system of several services?
    No. The guarantee applies within one database engine over one set of transactions. Reads served from replicas, values cached in the application, and workflows that span two services with separate transactions all fall outside it, so an invariant spanning services still needs its own mechanism — ownership of the invariant by one service, idempotency keys, or a saga with compensation.

saying these in an interview costs you the question

  • Assuming SERIALIZABLE imposes a flat percentage overhead independent of contention.
  • Ignoring transaction length and read-set width, which usually dominate the cost more than the level itself.
  • Proposing unlimited retries as the answer to a high abort rate, which lowers goodput further.
  • Reaching for the isolation level when a unique index, check constraint or a single self-referencing UPDATE would enforce the invariant more cheaply.
  • Believing the guarantee extends across replicas, caches or multiple services.

context