Why do most brokers force writes to persistent media in batches or on a timer rather than per record?
answer
- the cost does not divide
- one round trip covers many records
- latency floor on every write
- policy bounds the window instead
- copies answer a different failure
basics
~20 sForcing costs a round trip to the device that cannot be shared between records, so doing it per record collapses throughput to the device's operation rate. Batching amortises one force over many records and leaves a bounded loss window instead.
solid answer
~50 sA broker's write path is cheap because it appends to a file and lets the operating system's file cache absorb it; thousands of records cost almost nothing. A forced write is a different kind of operation: it waits on the storage device, and that wait cannot be amortised if every record demands its own. Per-record forcing therefore pins the broker's throughput to how many forced operations the device can complete per second, which is orders of magnitude below what the same hardware sustains for buffered appends, and it adds that wait to every writer's latency. So platforms in this class typically force on a policy — after a number of records, or every so many milliseconds — which turns an unbounded exposure into a bounded one, and buy their protection against single-machine failures from extra copies instead. Tightening the policy moves throughput down and shrinks the loss window; the operator chooses where on that curve a stream sits, if the platform exposes the choice at all.
go deeper
Know that pushing bytes to a real device is much more expensive than handing them to memory, and that brokers therefore do it periodically rather than for every single record.
Explain that one forced write covers a whole batch, so the cost does not divide, and name the two policy shapes: after so many records, or every so many milliseconds.
Argue the trade with numbers: what a tighter policy costs in sustained rate and write latency on your hardware, and what window the current policy actually leaves at peak ingest.
Own it as a small set of durability tiers rather than a per-stream argument, and say which tier is the default when nobody chooses, since that is the one most streams will end up on.
## What a force actually costs An ordinary append is a memory copy. The broker hands bytes to the operating system, they land in its **file cache**, and the call returns. Because the cost is memory-speed, a broker can absorb an enormous number of these without the storage device being involved at all yet. A **forced write** is a different animal. It instructs the operating system to get the cached bytes onto persistent media and not to return until the device confirms them. That is a round trip to hardware, and crucially it is a *serialising* one from the writer's point of view: the answer to that record cannot be sent until it completes. The important property is not that a force is slow in absolute terms — on modern hardware it can be fast. It is that **the cost does not divide**. One force covering ten thousand records costs about the same as one force covering one record. Per-record forcing throws that away and pins the broker's ceiling to the device's forced-operation rate. ## Two costs, not one 1. **Throughput.** The broker can no longer let the cache absorb a burst. Its sustained rate becomes a hardware property rather than a software one. 2. **Latency, on every write.** The device round trip is now inside the writer's wait, and it is added to whatever the acknowledgement rule already makes the writer wait for. That second point is what makes per-record forcing unattractive even on fast storage: it does not merely cost throughput at peak, it raises the floor of every single write's latency, including the quiet ones. ## Bounding the window instead Rather than forcing everything, platforms in this class typically express a **forcing policy**, and an operator should know which form theirs takes: - **After N records** — the loss window is bounded in records, so its size in time shrinks as traffic rises and grows when traffic is thin. - **Every T milliseconds** — the window is bounded in time, so its size in bytes grows with the ingest rate. This is usually the easier one to reason about, because a durability statement in seconds survives a traffic change. - **Left to the operating system** — the broker never asks, and the window is whatever the operating system's own write-out behaviour produces. Perfectly common, and the honest description is "we do not control this, we control copies". | Forcing policy | Loss window on power loss | Throughput and write latency | |---|---|---| | Every record | Effectively none on that machine | Pinned to the device's forced-operation rate; every write pays a round trip | | Every N records or T milliseconds | Bounded, and quotable in bytes at peak | One round trip amortised over a batch; writers rarely wait on it | | Left to the operating system | Unbounded from the broker's point of view | Highest; the device is never in the writer's path | ## Why this is an acceptable trade at all Because the alternative protection is already there. Keeping several copies of a stream on separate machines covers the failure that happens most often by far — one machine going away on its own — without any device round trip in the write path. Forcing is bought for a different event entirely: the machines going down together. Most operators decide that a bounded window of a second or so, combined with copies, is the right posture for most streams, and reserve tighter forcing for the few where it is not. ## Where the trade lands differently - On a design where durability comes from an **underlying shared replicated store**, the broker is not choosing a forcing policy at all; it inherits the store's durability contract, and the question becomes what that store promises and how fast it can promise it. - On **hardware whose device cache is protected against power loss**, a force is confirmed almost immediately, so the price of a tight policy drops sharply. The semantics do not change — an unforced record is still unforced — but the trade-off curve does. - On a **rented cluster**, the policy is often not exposed. The right move is to find out what the provider publishes and treat it as a fixed input to the posture rather than a knob. - On broker designs that keep exactly one **mirrored copy** rather than a configurable count, the same question applies to a pair of machines instead of a set, and the correlated-failure exposure is correspondingly narrower to reason about. ## What an interviewer is listening for That you know the cost is an un-amortisable device round trip rather than "disks are slow"; that you can name the two policy shapes and what each bounds; and that you present forcing and copies as answers to two different failure questions rather than as two strengths of the same dial.
- Which forcing policy shape is easier to write into a durability contract, and why?The time-based one. A window expressed in milliseconds stays meaningful as traffic changes, and can be multiplied by the current ingest rate whenever the bytes figure is wanted. A record-count bound silently produces a much longer exposure in time when a stream goes quiet, which is exactly when nobody is watching it.
- Does fast storage make per-record forcing reasonable again?It narrows the gap rather than closing it, and a device cache that is protected against power loss helps most. But the round trip is still in every writer's path and still does not amortise, so the ceiling is still a hardware property. It becomes a live option for a small number of streams, not a sensible default.
saying these in an interview costs you the question
- Says forcing is slow simply because disks are slow
- Thinks batching forces makes each individual write faster
- Believes a forcing policy removes the need for copies
- Cannot name what the policy bounds — time or records
- Assumes every platform lets an operator set this