skip to content

Some databases offer a mode where COMMIT returns before the transaction is guaranteed on durable storage. What does the system gain, exactly what can be lost, and when is that trade defensible?

level: seniorimportance: should knowfreq 42%

answer

  1. Ack from memory, flush shortly after
  2. Loses a bounded window of recent commits
  3. Process crash: usually nothing; power loss: the window
  4. No corruption — lost transactions, not torn data
  5. Per-transaction knob, not a global switch

basics

~20 s

You gain much lower commit latency and far higher small-transaction throughput, because commits no longer wait on the storage flush. You risk losing the most recent commits — a bounded window, typically well under a second — if the machine loses power or the OS panics. The database still comes back structurally intact.

solid answer

~60 s

Relaxed or asynchronous commit acknowledges the client as soon as the transaction's record is in memory, letting the flush happen shortly afterwards. Because the per-commit flush disappears from the critical path, latency drops and small-transaction throughput can rise by an order of magnitude. The precise loss profile matters, and candidates usually get it wrong: on a crash you lose **the most recent committed transactions**, bounded by the flush interval — typically sub-second. You do **not** get a corrupted database, and you do not lose atomicity: whatever did make it is applied whole, so the surviving state is consistent, just slightly stale. It is a lost-transactions failure, not a torn-data failure. Defensible when the data is regenerable or low-value per row and volume is high: telemetry, clickstream, audit-ish logs, caches, derived tables, bulk loads you can rerun. Indefensible when the client acted on the acknowledgement irreversibly — payments, order placement, identity, anything a user or another system was told about and cannot be re-derived. The good pattern is per-transaction, not global: run the system durable by default and relax specifically for the high-volume, low-value paths.

go deeper

for a junior

Know that some systems answer COMMIT before the data is safely on disk, and that a crash can then lose the newest transactions.

for a middle

State the gain and the bounded loss window, and give an example of data where that is acceptable.

for a senior

Give the exact failure profile per crash type, insist on per-path configuration, and separate it from disabling write barriers.

for a principal

Turn it into policy: classify data by replayability and external consequence, set defaults accordingly, and tie the chosen window to a stated recovery-point target with monitoring.

## What the mode actually changes In strict durability the sequence is: prepare the record, force it to storage, wait for the device, then answer the client. In relaxed/asynchronous commit the force-and-wait moves off the critical path: the record is in memory (and usually already in the OS page cache), the client is answered, and a background process flushes within some interval. The entire difference is *when the client is told*. Nothing about atomicity, isolation, or constraint checking changes. ## The exact failure profile This precision is what separates a senior answer: - **Process crash only** (the database is killed, machine keeps running): usually **nothing is lost** if the record already reached the OS page cache, because the kernel still holds it and will write it out. Relaxed commit is nearly free against this failure class. - **OS panic or power loss**: transactions committed within the flush window are gone. The window is a configurable interval, typically fractions of a second. - **What you get back**: a database that starts cleanly and is internally consistent as of some point slightly in the past. Not corruption, not half-applied transactions, not broken indexes. Recovery still guarantees all-or-nothing per transaction. So the honest sentence is: *'I trade a bounded window of recently acknowledged transactions for latency; I do not trade structural integrity.'* ## Why the acknowledgement is the dangerous part The damage is not the missing rows, it is the **lie already told**. If the API returned 201 Created, the email went out, the message was published to a queue, or another service recorded a foreign reference, the outside world now disagrees with the database and no restart fixes that. The rule follows: relaxed durability is acceptable exactly when nobody outside the database acted irreversibly on the acknowledgement, or when the writer can replay. This is why the same physical table can justify different settings by path: the ingestion writer for events may relax; the admin API that a human uses to change one of those events should not. ## Where it is a good trade - **High-volume telemetry, metrics, clickstream, IoT samples** — individually worthless, replayable from the producer, and the flush cost dominates ingestion. - **Bulk loads and ETL** where the input file is still available and the job is restartable. - **Caches, materialized projections, search-index side tables** — rebuildable from the source of truth. - **Non-critical bookkeeping** such as last-seen timestamps or view counters. ## Where it is not - **Money movement, ledgers, payments.** - **Order or booking creation confirmed to a user.** - **Identity, credentials, permissions, consent records** — losing a revocation is a security event. - **Anything emitting an external side effect on commit**, unless the outbox and the state share the same relaxed fate and the consumer is replay-safe. - **Audit and compliance data** where a regulator's question is 'what did you have at time T'. ## Neighbouring knobs, and one that is different in kind Relaxed commit is one point on a scale that also includes 'flush locally but do not wait for a replica', 'wait for a replica to receive but not flush', and 'wait for a replica to make it durable'. These all preserve structural integrity and vary the loss window and its failure domain. A genuinely different and more dangerous class is disabling the device's flush honouring — write barriers off, a lying consumer SSD, a virtualization layer discarding flushes. That does not just widen the loss window; it can allow the storage to reorder writes so that on-disk state after a power cut is not any valid past state. That is how real corruption happens, and it is why 'we set fsync off' and 'our SAN ignores flushes' are not the same conversation as 'we chose async commit'. ## How to answer Name the gain (latency and small-transaction throughput), state the loss precisely (bounded window of acknowledged transactions on power loss or OS panic; nothing on a mere process crash; no corruption), give the decision rule (did anyone act irreversibly on the acknowledgement, and can the writer replay?), and prefer per-transaction control over a global switch.

  • Does asynchronous commit risk leaving the database corrupted or half-applied?
    No. Each transaction is still atomic and the engine still recovers to an internally consistent state; you simply come back missing the most recent commits within the flush window. Corruption is a different failure mode, caused by storage that ignores or reorders flushes, not by choosing to acknowledge before flushing.
  • Which data would you keep on strict durability even in a system that is mostly telemetry?
    Anything with an external irreversible consequence or no replay path: billing and usage records that generate invoices, account and permission changes, consent and audit records, and the outbox rows that trigger side effects in other systems. The volume of these is usually a tiny fraction of the traffic, so keeping them strict costs almost nothing overall.
  • How would you validate that relaxed commit actually helps before rolling it out?
    Confirm the workload is flush-bound rather than CPU- or lock-bound — many small transactions, commit latency tracking device flush latency — and measure with the setting applied to a representative load. If the same throughput is achievable by batching transactions or raising concurrency, do that instead, because it costs no durability at all.

It is the difference between handing your letter to a courier who has already left, and waiting for the recipient's signature. Almost always fine, occasionally the van is lost — and the problem is that you already told everyone it arrived.

saying these in an interview costs you the question

  • Claiming asynchronous commit can corrupt the database or leave half-applied transactions
  • Not knowing the loss is bounded by a short flush window
  • Treating it as a global switch instead of a per-transaction or per-path decision
  • Applying it to payment, identity or audit data because 'the window is small'
  • Confusing it with disabling storage write barriers, which is a genuinely different and more dangerous change

context