Why does a log-structured engine need a commit log when it flushes its memory buffer to sorted files anyway, and when can old log segments be discarded?
answer
- memory is volatile
- replay rebuilds the buffer
- sync every write or periodically
- discard after the flush
basics
~20 sThe memory buffer is lost on a crash, so the commit log is what makes acknowledged writes recoverable; on restart it is replayed into a new buffer. A log segment can be discarded once every buffer holding its writes has been flushed to sorted files.
solid answer
~50 sBetween a write's acknowledgement and the next flush, the data exists only in the **in-memory buffer**, which a crash wipes out. The **commit log** closes that gap: each write is appended to it first, and on restart the engine **replays** the log from the last flushed point to rebuild the buffer. How durable an acknowledged write is depends on how the log is **synced**: syncing before acknowledging (often in small groups) protects every acknowledged write at a latency cost; syncing periodically is faster but can lose the last interval on a crash. Log segments are **reclaimed** once every buffer containing their writes has been flushed, since the sorted files now hold that data. If the log grows past its space limit, the engine forces flushes to free segments, and a long, unflushed log means a long replay on restart.
go deeper
Know that the commit log protects writes that are still only in memory and is replayed after a crash.
Explain sync policies and what each means for an acknowledged write, and when a log segment can be freed.
Diagnose slow restarts and log-space pressure, including tables that pin old segments, and choose a sync policy for a workload.
Be ready to set durability expectations for acknowledged writes across servers and replicas, and justify the latency cost.
## The gap the log closes In a log-structured engine a write lands in a **sorted memory buffer** and is acknowledged; it reaches a durable sorted file only when that buffer is flushed, perhaps minutes later. Without further protection, a crash in that window would lose every acknowledged write in the buffer. The **commit log** (write-ahead log) fixes this. It is a plain, append-only file of mutations in arrival order: 1. a write is **appended** to the log; 2. it is inserted into the memory buffer; 3. it is acknowledged. After a crash, the engine reads the log from the point that is known to be flushed and **replays** each mutation into a new buffer, restoring the state it had. ## How durable is "acknowledged"? Appending to a file does not by itself put bytes on disk; the operating system may hold them in memory. Engines offer a choice of **sync policy**: | policy | behaviour | trade-off | |---|---|---| | sync before acknowledging (often grouped) | writes wait for the log to be forced to disk, several writes sharing one sync | every acknowledged write survives a crash; higher latency | | periodic sync | writes are acknowledged at once, the log is synced on a timer | fast; a crash can lose writes from the last interval | In replicated stores the choice interacts with replication: a periodic-sync loss on one server may be recoverable from other copies, while a store whose log sits on a replicated storage layer gets redundancy from that layer. ## When a segment can go The log is split into **segments**. A segment is needed only while some of its mutations exist solely in memory. Once every buffer that contains writes from the segment has been **flushed** to sorted files, the segment is redundant and can be deleted, recycled or archived. Consequences: - **A quiet table can pin the log**: if one rarely written table's buffer holds a write from an old segment and never fills, that segment cannot be freed. Engines therefore force flushes of the tables holding the oldest segments when log space runs low. - **Replay time follows unflushed data**: the more data sits unflushed, the longer a restart takes. Flushing before a planned shutdown makes restarts fast. - **The log is shared**: one log usually serves every table on the server, so its space limit is a server-wide budget. ## What the log is not - It is **not** read by normal queries; reads use the memory buffer and sorted files. - It is **not** the long-term storage; sorted files are. - It is **not** the same as a change stream offered to clients, even if some stores build one from it. ## Interview angle Show that you know why the log exists, how sync policy decides what "acknowledged" means, and when segments are freed — including the quiet-table effect and restart replay time.
- Why put the commit log on a different device from the data files?The log is written sequentially and synced often, while flushes and compaction write large files. Sharing a device makes log syncs wait behind compaction I/O, raising write latency; a separate device keeps log appends fast.
- What happens at restart if the log is very large?The engine must replay every unflushed mutation into memory before the server is fully ready, so restart takes longer. Keeping the log bounded, and flushing before planned shutdowns, keeps replay short.
saying these in an interview costs you the question
- Saying reads are served from the commit log
- Believing an acknowledged write is always on disk regardless of sync policy
- Thinking log segments can be deleted as soon as a write is acknowledged
- Treating the commit log as the permanent copy of the data