Committing a transaction usually requires flushing the write-ahead log to stable storage with an fsync-style call. How does that shape commit latency and throughput, and what does group commit do about it?
answer
- commit = fsync of log through commit record
- latency floor = device flush latency
- group commit: one flush, many waiters
- deliberate delay trades latency for batch size
- relaxed sync = bounded loss window, still consistent
basics
~20 sCommit latency is floored by the storage device's flush latency, since the commit record must be durable before acknowledgement. Group commit batches many concurrent commits into one flush, so throughput scales with concurrency even though single-commit latency does not improve.
solid answer
~60 sA commit is durable only once its log record is on stable storage, which means a real flush — an fsync or a write with a durability barrier — not just a write into the OS page cache. That flush is the dominant cost of a small transaction: microseconds on NVMe with a power-protected cache, but potentially several milliseconds on spinning or network-attached storage. **Group commit** exploits the fact that the log is one ordered stream. Sessions committing concurrently each wait for the flushed log position to pass their own commit record; one flush that advances that position satisfies all of them. So under concurrency the engine amortises one flush across many commits, and throughput rises even though each individual commit still waits roughly one device flush. Some engines deliberately add a tiny delay before flushing to collect a bigger batch — trading a little latency for far fewer flushes. Diagnostically, if commit latency tracks device flush latency and throughput plateaus at low concurrency, group commit is either not engaging or the log device is the bottleneck.
code
text · 7 linesT1 appends COMMIT at LSN 900 -> waits for flushedLSN >= 900
T2 appends COMMIT at LSN 940 -> waits for flushedLSN >= 940
T3 appends COMMIT at LSN 980 -> waits for flushedLSN >= 980
log writer: single fsync -> flushedLSN = 980
T1, T2, T3 all released by that one flushgo deeper
Know that commit forces the log to disk and that many tiny transactions are therefore slow compared with batching.
Explain the flush cost, why the log is sequential, and how group commit lets concurrent commits share a single flush.
Diagnose it from wait events and latency distributions, and reason about remedies in order: batching, concurrency, log device, tuning, then relaxing durability with an explicit loss window.
Own the durability policy: decide per workload what recovery-point objective is acceptable, whether per-transaction settings are the right split, and how storage choices change the commit-latency floor for the whole platform.
## Why a flush is required at all Writing to a file usually just copies bytes into the operating system's page cache. If the machine loses power, those bytes are gone. Durability requires telling the storage stack \"do not return until this is really persisted\" — fsync, fdatasync, an O_DSYNC write, or the platform equivalent. Only then may the engine tell the client the transaction is committed. That is one synchronous round trip through filesystem, block layer, device queue and, on some hardware, the physical medium. Its latency does not shrink with faster CPUs. ## The latency floor So a trivially small transaction — update one row, commit — costs roughly: some microseconds of work, plus one flush. Typical orders of magnitude: - NVMe SSD with power-loss protection: tens of microseconds. - Consumer SSD honouring flushes properly: a few hundred microseconds. - Spinning disk: several milliseconds (a rotation). - Network-attached or replicated storage: a network round trip on top. This is why chatty application code that commits per row is so much slower than batching work into fewer, larger transactions: you pay one flush per commit regardless of how little the transaction did. It is also why a workload can be \"slow\" with an idle CPU and an idle data disk — everything is queued behind the log device. ## Group commit The redeeming property is that the log is a single sequential stream and durability is expressed as a position: a session's commit is durable once the flushed position passes its commit record. If twenty sessions append commit records and one flush pushes the durable position past all twenty, all twenty are durable. One device flush, twenty commits. Implementations differ in the details — a leader session performs the flush while followers wait, or a dedicated writer thread drains a queue — but the effect is the same: **flush cost per commit falls as concurrency rises**. Single-commit latency stays about one flush; aggregate throughput climbs until the log device saturates on bandwidth rather than on flush count. Some engines expose a deliberate small delay: wait N microseconds (or until M transactions are waiting) before issuing the flush, so more commits join the batch. That raises the latency of the earliest committer slightly in exchange for a much better flushes-per-second ratio at high concurrency. It only pays when there is enough concurrency to fill the window; on a lightly loaded system it is pure added latency. ## The knobs that trade durability for speed The other lever is to stop flushing synchronously at all. Engines offer settings that acknowledge a commit once the log record is in memory or in the OS cache, with flushes happening periodically instead. The trade is precise and worth stating carefully: - **What you lose**: the most recently committed transactions, up to the flush interval, if the *machine* crashes or loses power. This is a bounded data-loss window — a recovery-point objective measured in milliseconds or seconds. - **What you keep**: the write-ahead ordering itself, so the database still recovers to a *consistent* state — just an older one. A process crash where the OS survives generally loses nothing, since the bytes are already in the page cache. - **What you must not confuse it with**: disabling flushing entirely or lying about it (a device or virtualization layer that ignores flush commands). That breaks the ordering guarantee and can leave a genuinely corrupt database, not merely a stale one. These settings are often per-transaction, which is the mature answer: keep synchronous durability for the payment transaction and relax it for the click-tracking insert, in the same database. ## Diagnosing it Signs that log flushing is the bottleneck: commit latency clustering near the device's flush latency; throughput flat as concurrency rises (batching not engaging); high write IOPS with small write sizes on the log device; sessions waiting on a commit/log-flush wait event. Remedies in rough order of preference: batch application work into fewer transactions; raise concurrency so group commit can amortise; put the log on a device with power-loss-protected cache; tune the group-commit delay; and only then consider relaxing synchronous durability, with the loss window written down and accepted by whoever owns the data. ## Interview framing \"Commit is one durable sequential flush, so latency is device-bound and throughput depends on how many commits share a flush. Group commit amortises it; relaxing synchronous commit removes it but buys you a bounded window of lost commits.\"
- An application inserts 10,000 rows one row per transaction and it is slow, even though CPU and the data disk are idle. What is happening and what is the first fix?Each commit pays its own log flush, so the run costs about 10,000 device flushes regardless of how little work each statement does. The first fix is to batch the inserts into far fewer transactions, which collapses 10,000 flushes into a handful. If the rows genuinely arrive independently, raising concurrency lets group commit amortise flushes instead.
- If a team turns off synchronous commit, what exactly can they lose, and what can they not lose?They can lose the most recently acknowledged commits — everything not yet flushed, bounded by the flush interval — if the machine loses power or the OS crashes. They do not lose consistency: write-ahead ordering is preserved, so the database recovers to an earlier but valid state, and a crash of just the database process typically loses nothing since the records are already in the OS cache. What would break this is hardware or a hypervisor that ignores flush commands, which risks genuine corruption.
A courier who will only depart when the mailbag is sealed. Each letter waits for one departure; but if twenty letters arrive before the van leaves, they all ride on the same trip.
saying these in an interview costs you the question
- Believing a plain write to the log file is enough for durability
- Saying group commit reduces the latency of an individual commit
- Claiming relaxed synchronous commit can corrupt the database (it costs recent commits, not consistency)
- Not distinguishing an OS/process crash from power loss when reasoning about the loss window
- Assuming faster CPUs or more memory fix a commit-flush bottleneck