skip to content

questions

5

When a database returns success for a COMMIT, what exactly has it promised the client, and what has it not promised?

level: juniorimportance: must knowfreq 58%

answer

  1. Committed = survives crash / power loss
  2. Stable storage, not memory
  3. Not media loss, not node loss, not a bad DELETE
  4. Lost commit response = unknown, not failed
  5. 'Durable against what?' — crash / node / site

basics

~20 s

It promises the committed changes survive a crash or power loss of that node: the data reached stable storage, not just memory. It does not promise the data survives losing that machine or its disk, nor that any replica already has it.

solid answer

~60 s

A successful COMMIT means: **this transaction's effects will still be there after the process dies, the OS panics, or the power is cut** — the changes reached durable storage rather than sitting only in volatile memory. It also means the effects are all-or-nothing and visible to subsequent readers. What it does *not* promise, by default, is broader than people expect: - It does not promise the disk survives. Durability is against crashes, not against media failure, deletion, or an accidental `DROP TABLE`. Backups cover that, not ACID. - It does not promise any **other machine** has the data. With asynchronous replication, a commit acknowledged on the primary can be lost entirely if the primary dies and a replica is promoted. - It does not promise anything about a data centre or region failing. So durability is a scoped promise. The right interview move is to ask 'durable against what?' — process crash, node loss, disk loss, site loss — because each level needs a different mechanism and costs a different amount of commit latency.

go deeper

for a junior

Define it plainly — after COMMIT returns, the change survives a crash or power loss because it reached stable storage — and know it is not a backup.

for a middle

Add the scope: single node by default, replicas may not have it, and a lost commit response means unknown rather than failed.

for a senior

Discuss the levels (buffered, node-durable, replica-acknowledged, cross-site) and the failure class each survives.

for a principal

Frame it as a recovery-point decision per workload, including the latency each level imposes on every commit and where the application must still be idempotent.

## The core guarantee Durability is the D in ACID: once the engine tells the client 'committed', the transaction's effects survive a subsequent failure. The failure the classical definition has in mind is a **crash** — the process is killed, the kernel panics, the power is pulled — anything that wipes volatile memory but leaves the storage intact. After restart, the database comes back with that transaction applied. The practical implication is a rule about ordering: the engine must not acknowledge the commit until enough information to reconstruct the transaction is on stable storage. Everything else — pages still dirty in the buffer pool, the data file not yet updated — is fine, because the recorded information is sufficient to restore the state on restart. ## Why this matters to application code Durability is the boundary at which an application is allowed to tell the outside world that something happened. Before the commit returns, you must not send the confirmation email, charge the card, hand the message to another system, or return 201 Created. After it returns, you may. Systems that acknowledge to users before the commit returns will, sooner or later, tell someone their order exists when it does not. Equally important: if the commit call **fails or the connection drops**, the outcome is *unknown*, not 'failed'. The transaction may have committed on the server just as the network broke. Correct clients treat a lost commit response as indeterminate and resolve it by re-reading state or by making the operation idempotent with a client-supplied key, rather than blindly retrying and creating a duplicate. ## What the promise excludes **Media failure.** If the drive holding the data is destroyed, ACID durability says nothing. Recovery from media loss is the job of backups, redundant storage, and replicas. **Node loss.** Single-node durability means the data is on *that node's* stable storage. If the node is gone — hardware death, a terminated cloud instance, a failover to a replica — anything the replicas had not yet received is gone with it. **Logical damage.** A committed `DELETE` is durable. Durability makes mistakes permanent as reliably as it makes correct work permanent. Point-in-time recovery, not ACID, is the answer. **Anything the storage layer lies about.** If a device acknowledges a flush while data is still in a volatile write cache, the engine's promise is broken by the hardware beneath it. Enterprise-grade devices either honour flush semantics or protect their cache with a battery or capacitor. ## Durability is a spectrum, not a boolean Production systems configure a point on a scale: 1. **Buffered only** — the commit is acknowledged from memory. Fastest, loses recent commits on a crash. This is not durability; it is durability deliberately traded away. 2. **Flushed to the node's stable storage** — survives process and OS crash and power loss on that node. The classical ACID D. 3. **Acknowledged by one or more replicas** — survives losing the node entirely; the promoted replica has the transaction. 4. **Acknowledged in another failure domain** (rack, zone, region) — survives site loss, at the cost of cross-site round-trip latency on every commit. Each step up adds latency to every commit and buys a wider class of survivable failure. 'Is your system durable?' is therefore an incomplete question; 'what failures does a returned commit survive, and what is your acceptable data-loss window?' is the real one. ## Answering well Give the crisp definition first — committed effects survive crash and power loss because the record reached stable storage. Then immediately scope it: single-node by default, not media loss, not node loss unless replication is synchronous, and note that a lost commit response means unknown rather than failed. That combination separates a candidate who memorized ACID from one who has operated a database.

  • Your client sends COMMIT and the connection drops before any response arrives. Did the transaction commit?
    Unknown — it may have committed on the server just as the link failed. The client must resolve the ambiguity rather than assume failure: re-read the state to see whether the effect is present, or design the operation to be idempotent with a client-supplied key so a retry cannot duplicate it. Blind retry is how duplicate orders and double charges are created.
  • Does durability protect against someone running DELETE FROM orders with no WHERE clause?
    No — it guarantees the opposite: once that delete commits, it is permanently applied. Protection against logical damage comes from backups and point-in-time recovery, plus access controls and review on destructive statements. Confusing the two is a common gap.

COMMIT returning is like getting a stamped receipt rather than a verbal promise. The stamp survives the clerk fainting — it does not survive the building burning down unless someone filed a copy elsewhere.

saying these in an interview costs you the question

  • Saying durability means the data is written to the data file immediately
  • Believing a committed transaction survives losing the machine even with asynchronous replication
  • Treating durability as protection against accidental deletes or bad migrations
  • Treating a lost commit response as 'it failed' and retrying blindly
  • Acknowledging work to users or downstream systems before the commit returns

context

open as a page

Why does making a database commit genuinely durable cost latency, and what physically has to happen before the engine can acknowledge the commit?

level: middleimportance: must knowfreq 50%

basics

~20 s

Before acknowledging, the engine must get the transaction's record onto storage that survives power loss — an fsync that the device honours, not just a write into the OS page cache. That is a physical round trip to the device, and adds a network round trip too if a replica must confirm.

open as a page

Your primary database node dies and a replica is promoted, yet several transactions the application had already been told were committed are missing. How is that possible, and how would you configure the system so it cannot happen?

level: seniorimportance: must knowfreq 48%

basics

~20 s

With asynchronous replication, the primary acknowledges a commit once it is durable locally; replicas receive it slightly later. If the primary dies before shipping those transactions and you promote a replica, they are lost. Preventing it means requiring at least one replica to acknowledge before the client is told committed.

open as a page

Some databases offer a mode where COMMIT returns before the transaction is guaranteed on durable storage. What does the system gain, exactly what can be lost, and when is that trade defensible?

level: seniorimportance: should knowfreq 42%

basics

~20 s

You gain much lower commit latency and far higher small-transaction throughput, because commits no longer wait on the storage flush. You risk losing the most recent commits — a bounded window, typically well under a second — if the machine loses power or the OS panics. The database still comes back structurally intact.

open as a page

You are setting the durability configuration for a new platform whose data ranges from payment ledgers to high-volume device telemetry. How do you decide what each dataset's commit should wait for, and what evidence backs the decision?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Classify data by two questions: can the writer replay it, and did anyone outside act irreversibly on the acknowledgement? That yields an acceptable loss window per dataset, which maps to a commit level — relaxed local, node-durable, or replica/quorum-acknowledged — and the latency budget each costs. Then verify by testing real failover.

open as a page