In a store that appends every write to a write log, which flush granularities decide when those appends reach disk, and what does each cost?
answer
- appending is not flushing
- three granularities, three exposures
- acknowledged to the caller, not yet on disk
- a dead process differs from a dead machine
basics
~20 sA flush policy decides when appended writes reach disk: flush on every write, flush on a timer of about a second, or leave flushing to the operating system. Each trades write-path cost against writes at risk.
solid answer
~50 sAppending an entry and flushing it to disk are two acts: the append lands in the operating system's buffers, and a flush pushes it to the device. The flush policy picks the granularity. **Flush on every write** holds the exposure to roughly the append in progress, but every write waits on the device, so write throughput becomes bound by how fast the device accepts flushes. **Flush on a timer of about a second** puts almost nothing on the write path and leaves roughly an interval's worth of appends unflushed. **Leaving flushing to the operating system** costs least and exposes most, with one asymmetry worth knowing: if only the process dies, the operating system still writes those buffers out, whereas losing the machine loses them. Under the last two policies a write is routinely acknowledged to the caller before it is flushed to disk.
go deeper
Remember that writing to the log and getting it onto the device are two different steps, and that a setting decides how often the second one happens. Being told a write succeeded is not the same as it being on disk.
Name the three granularities and price each one in both currencies: what is at risk, and what the write path pays. Explain why flushing on every write bounds the tier's write throughput to the device's flush rate.
Separate a killed process from a lost machine, since the operating-system policy behaves very differently in each. Mention back pressure, failed flushes and the device's own cache as the reasons a nominal interval is not a promise.
Treat the policy as a lever with a measurable price on every write, and insist the exposure is measured on the real storage rather than quoted from a document. Then ask whether this tier should be holding the data in question at all.
## Appending and flushing are two acts When the process appends an entry to the write log it hands bytes to the operating system, which holds them in its own buffers and writes them to the device on its own schedule. A **flush** is the explicit act of pushing those buffers to the device and waiting for it to say so. The **flush policy** is the setting that decides how often the store performs that act. Everything at risk in this posture lives in the gap between the two: the appends that exist in memory somewhere but not yet on the device. ## The three granularities Stores that offer a write log expose granularities along these lines. The names and the exact number vary between stores, and some offer only one: - **Flush on every write.** The store flushes after each appended entry. What is at risk is at most the append in progress. The price is that every write waits for the device, so the write path inherits the device's latency and the tier's write throughput becomes bound by how many flushes the device accepts per second. Some stores soften this by letting callers arriving together share a single flush, so the per-write cost is amortised across a batch; others do not. - **Flush on a timer (about a second).** A background activity flushes at a fixed interval. The write path pays almost nothing, because the caller never waits for the device. What is at risk is roughly an interval's worth of appends. This is the usual middle setting because it converts an open-ended exposure into a bounded one for nearly no latency. - **Leave flushing to the operating system.** The store initiates no flush; the operating system writes its buffers out whenever it decides to. The write path is as cheap as it gets, and what is at risk is whatever the operating system is holding, which can be many seconds' worth. | Flush policy | What is at risk if the machine is lost | Cost on the write path | |---|---|---| | Flush on every write | At most the append in progress | A device round trip per write, sometimes shared between concurrent callers | | Flush on a timer (about a second) | Roughly the interval's worth of appends | Almost none; the flush happens off the caller's path | | Leave flushing to the operating system | Whatever the operating system is still buffering | Lowest; the store initiates no flush at all | ## The asymmetry between a dead process and a dead machine These two failures are not the same, and the difference is sharpest under the last policy. If the **store process** dies while the machine keeps running, the operating system still owns the buffered appends and writes them out; nothing is lost that the store had already appended. If the **machine or its power** is lost, everything not yet on the device goes with it. A candidate who collapses the two will describe the operating-system policy as far worse than it is for the common case of a process being killed, and far better than it is for a power cut. ## The acknowledged write that is not yet flushed The sharpest object in this whole category is **a write acknowledged to the caller but not yet flushed to disk**. Where the acknowledgement sits relative to the flush is itself a design decision: - A store that flushes before answering pays the device's latency inside every write call, and the caller's success reply means the bytes are on the device. - A store that answers first has a window, however short, in which the caller believes a write happened that the device has never seen. Under the timer policy and the operating-system policy, this is the normal state of affairs rather than an edge case. Naming the raw exposure - the appends since the last flush - is the mechanism; deciding whether a given workload can live with it is a separate conversation. ## What varies, and what to check before quoting a number - **Back pressure.** When the device cannot keep up, some stores slow or block writers so the exposure stays near the nominal interval; others keep buffering, and the amount unflushed grows past it. - **Failed flushes.** Some stores stop accepting writes when a flush fails; others record the failure and carry on, which quietly widens the exposure. - **The device's own cache.** A device with a volatile write cache can report a flush complete before the bytes are durable unless that cache is protected or the flush is honoured end to end. The honest answer to "is it on disk" does not stop at the store. - **The storage underneath.** Per-write flush costs behave nothing alike on a local device and on network-attached storage; the policy that is affordable on one can halve throughput on the other. - **What the policy does not change.** No flush policy makes the log shorter. Log size is a function of how much has been appended since it was last compacted, and it is a separate problem with a separate remedy.
- Why does flushing on every write cap write throughput rather than just adding latency?Because each write has to wait for the device to confirm the flush, the tier can complete no more writes per second than the device accepts flushes. Latency and throughput are the same constraint here. Stores that let concurrent callers share one flush raise that ceiling, because one device round trip then covers a batch rather than a single write.
- If the flush policy is a timer of about a second, is the exposure exactly one second?Treat it as a nominal figure, not a guarantee. It holds while flushes complete promptly. If the device slows down, some stores apply back pressure to writers and keep the exposure near the interval, while others keep buffering and it grows. A failed flush that the store only records can widen it further.
- Does choosing a stricter flush policy reduce how long the process takes to replay at start?No. Flushing decides when appended bytes reach the device, not how many entries the log holds. Replay applies every entry in the log regardless of when each was flushed, so a stricter policy costs write-path time without shortening a start. Shortening the log is what compaction is for.
saying these in an interview costs you the question
- Says a write is on disk the moment the server acknowledged it.
- Says a process crash loses appends the operating system already holds.
- Treats flushing on every write as free because appends are sequential.
- Assumes the timer policy's exposure holds however the device behaves.
- Thinks a stricter flush policy also keeps the write log smaller.
- Believes a successful flush proves the device itself has the bytes.