skip to content

When a reader asks for a record written seconds ago, what decides whether the volume is touched at all?

level: middleimportance: should knowfreq 46%

answer

  1. written once, read again immediately
  2. memory outside the process or inside it
  3. distance from the newest record decides
  4. a reader far behind loses the shortcut

basics

~20 s

Whether that record is still resident in memory somewhere: the operating system's file cache, the broker process's own memory, or neither. On designs that cache little, and for any reader that has fallen far behind, every such read goes to the volume.

solid answer

~50 s

A record that was appended seconds ago may be served from one of three places, and platforms differ over which. It can sit in **the operating system's file cache**, because the host keeps recently written file data in memory and hands it straight back. It can sit in **in-process memory**, because the broker itself retains recent records rather than delegating to the host. Or it can sit in neither, in which case the read is a real transfer from the volume and competes with the appends still arriving. Which one applies decides what a caught-up reader costs and what a fan-out of many readers costs: when residency holds, one write can be served to several readers with no extra media work. A reader that has fallen far behind falls out of all of this — its records left memory long ago.

go deeper

for a junior

The point to hold on to is that a record written a moment ago is often still in memory, so reading it back may not touch the storage at all. That is why keeping up is cheap and falling behind is not.

for a middle

Explain the three residencies and who controls eviction in each. Be explicit that a restart empties all of them, and that memory residency speeds reads without doing anything for the write path or for durability.

for a senior

Demonstrate that you reason from the distance between a reader and the newest record. Be able to trace why a large re-read on one record-serving node raises write service times for producers writing to the same volume.

for a principal

The tradeoff to own is where the memory budget lives — with the host or inside the process — and what each choice implies for node shape, restart behaviour and how a fan-out heavy estate is provisioned before anyone measures it.

## The question behind the question A broker is often described as "writing everything to disk", which makes people assume every read is a media transfer. In practice, the read that dominates a healthy cluster — a reader asking for what was written a moment ago — frequently never reaches the volume. Understanding *why* is what separates someone who can reason about broker storage from someone repeating a slogan, and it is the reason two clusters with identical media can behave completely differently under fan-out. ## Three places a recently written record can be - **The operating system's file cache.** When the broker writes file data, the host keeps a copy of it in memory and serves later reads of the same bytes from there. The broker itself does nothing special; it reads the file back and the host answers from memory. The memory is managed by the host, shared with everything else running on it, and reclaimed under pressure without asking. - **In-process memory.** The platform keeps recent records inside its own process, in structures it controls, and answers reads from there. It decides what to keep and what to evict, and the memory is accounted to the broker rather than to the host. - **Neither.** Some designs deliberately cache very little, or the record has already aged out of whatever held it. The read is then a genuine transfer, and it shares the volume's throughput with the appends still arriving. ## What each residency costs | residency | who controls eviction | what a restart does | the main cost | |---|---|---|---| | the operating system's file cache | the host, under memory pressure | memory is empty; early reads hit the volume | shared with every other process on the host | | in-process memory | the broker | memory is empty; early reads hit the volume | the memory is spent and unavailable for anything else | | neither | not applicable | no change; behaviour is already steady | every read competes with the write path | Two points are worth stating flatly because they are routinely confused. First, none of these three is a durability mechanism: a record resident in memory is not thereby safe, and whether an acknowledged write has reached durable media is a separate question with its own answer. Second, memory residency does not change the write path at all — appends still have to be written out at the incoming rate, so no amount of memory rescues a volume that cannot sustain it. ## Why reader behaviour decides the outcome Residency is not a property of the cluster; it is a property of the *distance* between a reader and the newest record. Three cases: 1. **Caught up.** The reader asks for bytes written seconds ago, they are still resident, and the read is nearly free. Several readers of the same stream in this state all hit the same resident bytes, so fan-out is cheap. 2. **Slightly behind.** The reader is asking for records that are near the edge of whatever memory holds them. Behaviour is unstable — sometimes served from memory, sometimes from the volume — and shows up as inconsistent read service times. 3. **Far behind.** The records left memory long ago. Every batch is a media transfer, and those transfers compete with the appends. This is the mechanism by which a large re-read raises write latency for everyone else on the same node. ## What varies between platforms This is one of the areas where a claim that is true of the platform you know best is simply false elsewhere: - Leaning on **the operating system's file cache** is one platform family's explicit doctrine, and it presumes a deployment where the host's spare memory is the broker's cache. Another family keeps records inside the process instead, which gives it control and costs it memory. - A broker that **removes a record on acknowledgement** and is keeping up may serve almost everything from memory and use the volume mainly so that a restart does not lose the queue. Its storage character is only exposed when it falls behind. - A design whose primary store is a **remote service** has no local residency question in this form at all; its equivalent question is what it keeps near the compute and for how long. ## What residency is not It is not durability, it is not a substitute for volume throughput, and it is not a free way to serve readers who are behind. It is the explanation for a specific, very common observation: that a healthy broker with many readers can be doing far less media work than its ingress rate suggests, and that the moment readers fall behind, that pleasant arrangement disappears and the volume's real character becomes visible.

  • Why does a reader that has fallen far behind cost more than one that is caught up?
    The records it wants are no longer in whatever memory held the recent ones, so each batch is a real transfer from the volume. Those transfers share throughput with the appends that are still arriving, which is why a large re-read can raise the write service time seen by producers on the same record-serving node.
  • Does holding recent records in memory protect them if the record-serving node dies?
    No. Residency decides the cost of serving a read, not survival. Whether an acknowledged write has actually reached durable media, and what is lost if the node dies before it does, is a separate durability question — and memory that has not been written out is precisely what is at risk.
  • Why do several readers of the same stream often cost little more than one?
    Because they are all asking for the same recently written bytes, and if those bytes are resident, each extra reader adds network transfer but no extra media work. Fan-out stays cheap exactly as long as the readers stay close to the newest record; it stops being cheap the moment one of them drifts.

saying these in an interview costs you the question

  • Says every read is a media transfer because the records are on disk.
  • Assumes the memory that speeds reads also makes the write durable.
  • Thinks caching works identically on every platform because they all append.
  • Believes a reader far behind is served as cheaply as one that is caught up.
  • Treats memory held inside the broker process and host file-cache memory as interchangeable.