Why does making a database commit genuinely durable cost latency, and what physically has to happen before the engine can acknowledge the commit?
answer
- Buffers → page cache → device cache → media
- write() is not durable; fsync is
- Cost is per commit, not per byte
- Batch + concurrency are the levers
- Replica ack = extra network round trip
basics
~20 sBefore acknowledging, the engine must get the transaction's record onto storage that survives power loss — an fsync that the device honours, not just a write into the OS page cache. That is a physical round trip to the device, and adds a network round trip too if a replica must confirm.
solid answer
~60 sA commit is a **synchronization point**. Writing into the operating system's page cache is not durable — a power cut loses it — so the engine issues a flush (fsync or an equivalent barrier) and waits until the storage device reports the data is on non-volatile media. That wait is real physics: on a spinning disk it is a rotation and possibly a seek (milliseconds); on an SSD it is a program operation plus cache-flush handling (tens to hundreds of microseconds); on a device with a battery-backed or capacitor-protected write cache, the acknowledgement can come from that cache safely and is much faster. This matters because the cost is **per commit**, not per row. A thousand single-row transactions pay a thousand flushes; the same thousand rows in one transaction pay one. That is why chatty auto-commit loops are so much slower than batched work, and why commit throughput often tracks the storage device's flush rate rather than its bandwidth. If durability is defined to include a replica, add at least one network round trip to a peer that must itself make the record durable before answering. Cross-zone or cross-region replicas add that distance to every commit.
code
text · 8 lineslevel survives typical wait
---------------------------- ----------------------------- -------------
engine buffer (memory) nothing beyond a clean exit ~0
OS page cache (write()) process crash ~0
device cache, unprotected process + OS crash ~0
device cache, capacitor/BBU power loss ~10-50 us
non-volatile media (fsync) power loss 0.1-10 ms
remote replica acknowledged loss of the whole node + network RTTgo deeper
Know that COMMIT waits for data to reach storage that survives power loss, which is why committing in a tight loop is slow.
Explain the buffer/page-cache/device-cache/media layering, that only a flush crosses to durable, and that cost is per commit so batching helps.
Quantify by device class, explain why concurrency raises aggregate throughput, and refuse the disable-barriers shortcut with the reason.
Turn it into a budget: which workloads pay a synchronous cross-zone commit, which run node-durable, and what hardware or topology change moves the whole curve.
## The layers a write passes through 1. **Database buffers** — in process memory. Lost on process crash. 2. **OS page cache** — a normal `write()` lands here. Survives the process dying, lost on kernel panic or power cut. 3. **Device write cache** — volatile DRAM inside the drive or controller, unless it is battery/capacitor backed. 4. **Non-volatile media** — platters or NAND. Survives power loss. Durability requires the transaction's record to reach layer 4 (or a layer-3 cache that is power-protected). Only the database can force that; a plain `write()` never does. The engine issues `fsync`/`fdatasync` (or opens with O_DSYNC, or uses the platform equivalent) and **blocks until the device answers**. That block is the latency. ## Why the wait cannot be optimized away for a single commit The device has to actually place the bytes somewhere that survives losing power, and report that it did. The magnitudes, roughly: - **HDD**: ~5–10 ms for a flush that requires rotation, sometimes a seek. - **Consumer SSD**: tens to hundreds of microseconds, highly variable, sometimes much worse under garbage collection. - **Enterprise SSD / NVMe with power-loss protection**: the on-device cache is capacitor-backed, so a flush can be acknowledged from cache — often tens of microseconds. - **Battery-backed RAID controller cache**: similar effect; the controller answers from protected DRAM. - **Cloud network-attached storage**: adds network latency to each flush; typically the dominant term. This is why 'durability is cheap on modern hardware' is only true when the hardware has power-loss protection. On a laptop SSD lying about flushes, durability is fast *and fake*. ## Per-commit, not per-byte The crucial mental model: durable commit cost is dominated by **flush count**, not data volume. Consequences: - 10,000 inserts in auto-commit mode = 10,000 flushes. The same inserts in one transaction = one. Batching is often a 10–100x throughput change with no other code difference. - Small transactions on a high-latency device are bound by the device's flush rate, so throughput per connection is roughly 1/flush-latency. More concurrent connections raise total throughput because independent commits' flushes can overlap and be serviced together. - Adding CPU or bandwidth does not help a flush-bound workload; changing the storage's flush characteristics does. ## Replication adds a second cost If 'durable' means 'a second machine has it', commit latency includes a round trip to that machine, plus that machine's own durability step if the configuration requires it. Same-rack is sub-millisecond; cross-availability-zone is single-digit milliseconds; cross-region is tens to over a hundred milliseconds — which is why cross-region synchronous commit is a deliberate, expensive design choice rather than a default. Some systems soften this by requiring only that the replica *received* the record into memory, not that it flushed it, trading a narrow failure window for much lower latency. ## What you can do about it - **Batch.** Fewer, larger transactions. The single biggest lever. - **Raise concurrency.** Independent transactions' flushes overlap, so total commit throughput scales beyond one connection's flush-latency limit. - **Buy power-loss-protected storage.** Turns a milliseconds-scale flush into microseconds legitimately. - **Segregate by value.** Keep strict durability for money and identity; consider relaxed modes for data whose loss window is genuinely acceptable. - **Never** solve it by disabling the device's flush honouring or mounting with barriers off. That converts a latency problem into silent data loss on power failure. ## Answering well Name the layers, say that only a device-honoured flush makes a commit durable, quantify roughly, and stress that the cost is per commit so batching and concurrency are the levers — then note that including a replica adds a network round trip to every commit.
- A batch job doing 50,000 single-row inserts in auto-commit takes minutes. What is the first thing you change and why?Group the inserts into transactions of a few hundred to a few thousand rows so the job pays one durable flush per batch instead of one per row. Commit cost is dominated by flush count, not data volume, so this alone is commonly a 10–100x improvement. Batch size is then tuned against lock-hold time, rollback size, and retry granularity.
- Why can total commit throughput exceed one divided by the flush latency?Because independent transactions commit concurrently, and their durability work overlaps — flushes issued at the same time are serviced together rather than strictly one after another. So a single connection is bounded by flush latency, while a busy server with many connections is bounded by the device's aggregate flush capacity, which is much higher.
- Someone proposes disabling write barriers on the filesystem to speed up commits. What do you say?That is not a performance tuning knob, it is silently turning durability off: the device may acknowledge writes that are still in volatile cache, so a power cut can lose acknowledged commits and, worse, leave the on-disk state inconsistent. If lower latency is genuinely needed, use storage with power-loss protection or an explicitly configured relaxed-commit mode, so the loss window is a documented decision instead of a hidden one.
saying these in an interview costs you the question
- Thinking a normal write() to the file is durable
- Believing commit cost scales with the number of rows rather than the number of commits
- Claiming fsync is free on SSDs without mentioning power-loss protection
- Proposing to disable write barriers or flush honouring as a tuning step
- Assuming a synchronous replica in another region costs nothing on the commit path