skip to content

A broker acknowledges a record as stored, yet a power cut on that machine loses it — why?

level: juniorimportance: must knowfreq 62%

answer

  1. answered is not the same as stored
  2. bytes sit in volatile memory first
  3. the operating system's file cache
  4. forcing is a separate instruction
  5. accepted but not forced equals the loss window

basics

~20 s

An acknowledgement usually means the bytes were accepted into the operating system's file cache, which is volatile memory, not that they were forced onto persistent media. Everything accepted but not yet forced is the loss window a power cut takes.

solid answer

~40 s

Two different events get the same word. **Acceptance** is the broker handing bytes to the operating system, which copies them into its file cache and returns immediately; the broker can answer the writer at that point. **Forcing** is a separate, explicit instruction that pushes those cached bytes onto persistent media and waits for the device. On most platforms in this class the writer is answered after acceptance, and forcing happens later — after a number of records, on a timer, or whenever the operating system decides to write the cache out. So an acknowledged record survives the broker process dying, because the operating system still holds the bytes and will write them out, but a machine that loses power drops everything accepted since the last force. That span is the loss window.

go deeper

for a junior

Remember the one sentence: a record that has been acknowledged may still be sitting in memory rather than on a disk. Answered and stored-for-good are not the same event.

for a middle

Explain the two steps — acceptance into the operating system's file cache, then a separate forced write to persistent media — and say which one the writer's answer usually waits for. Be able to define the loss window.

for a senior

Quote the window in seconds and in bytes at peak for the cluster you run, say whether the platform lets you change it per stream, and be honest that a process crash and a power cut are not the same test.

for a principal

Frame it as a posture the estate commits to: which streams are allowed an unforced window at all, what number goes in the contract with their owners, and how that choice is checked on streams nobody reviews again.

## Two events that share one word When a writer sends a record and the broker answers "stored", two quite different things may or may not have happened to those bytes. - **Acceptance.** The broker hands the bytes to the operating system with an ordinary write. The operating system copies them into its **file cache** — ordinary volatile memory that it manages — and returns at once. Nothing has reached a physical device yet. The broker is now free to answer the writer. - **Forcing.** A separate, explicit instruction tells the operating system to push those cached bytes onto persistent media and not to return until the device reports them stored. Nothing about an acknowledgement implies the second event happened. On most brokers and streaming platforms in this class, the answer to the writer is sent after acceptance, and forcing runs on its own schedule behind it. ## Committed is not durable It is worth keeping two words apart deliberately, because the whole subject lives in the gap between them. - **Committed** means the acknowledgement rule was satisfied — whatever the operator asked the server to wait for before answering: no wait at all, the leader alone, a majority of copies, or every caught-up copy. - **Durable** means the bytes are on persistent media and will still be there after the machine loses power. A record can be fully committed under the strictest acknowledgement rule available and still be durable nowhere. | State of a record | Survives the broker process dying | Survives the machine losing power | |---|---|---| | Accepted into the file cache on one node | Yes — the operating system still holds it | No | | Accepted into the file cache on several nodes within one cluster | Yes | Only if those machines lose power separately | | Forced onto persistent media | Yes | Yes | The middle row is the one that surprises people. Extra copies genuinely protect the record, because a machine that crashes on its own leaves the others holding the bytes and able to write them out. They protect it much less when the machines go down together, because each of them was holding an unforced span in memory at that moment. ## Sizing the loss window The **loss window** is the span of records that were accepted but not yet forced. It has two useful units, and an operator should be able to quote both: 1. **Time** — how long it has been since the last force. If the platform forces roughly once a second, the window is roughly a second of ingest. 2. **Bytes or records** — that time multiplied by the write rate on that node. A node taking 50 MB per second with a one-second forcing policy is carrying up to about 50 MB of exposed data. Both numbers move with load: the same forcing policy produces a far larger window at peak than at night. That is why a durability posture written in seconds is easier to reason about than one written in records. ## What varies between platforms This is a place where platforms in the same class genuinely differ, and an interview answer that ignores that reads as one product's manual: - Some platforms expose a forcing policy as a per-stream override on top of a cluster default; others expose only a cluster-wide one; a rented cluster may expose none at all and simply publish what it does. - On designs where durability comes from an underlying shared replicated store rather than from per-node copies, the broker is not the thing making this decision. The store's own durability contract is what the acknowledgement inherits, and the question becomes what that contract says. - Broker designs that keep exactly one mirrored copy answer once the pair holds the record — which is still memory on both machines unless something forced it. - Hardware changes the price without changing the meaning. A device whose own cache is protected against power loss can confirm a force very quickly, so forcing costs less there; it does not make an unforced write safe. ## What an operator does with this 1. Find out what the platform actually does before it answers a writer, and whether that is configurable. 2. State the loss window in seconds and in bytes at peak ingest, per node. 3. Decide whether an event that takes several of those machines at once is in the threat model for this stream. 4. Record the answer as part of the stream's durability posture, next to its copy count and its acknowledgement rule, rather than leaving it as an accident of the defaults.

  • If the broker process is killed but the machine stays up, are the unforced records lost?
    No. They are in the operating system's file cache, not inside the broker's own memory, so the operating system still holds them and writes them out normally. That is exactly why a process crash and a power cut are different events, and why people who have only ever seen process restarts conclude the acknowledgement was durable.
  • How would you express this leaf's risk in a number an owner of the stream can act on?
    Quote the loss window twice: in time since the last force, and in bytes at peak ingest per node. "Up to about a second, roughly 50 MB per node at peak" is actionable; "writes are buffered" is not. Then say whether an event taking those machines together is considered plausible for this stream.
  • Does a stricter acknowledgement rule close the loss window?
    No. Waiting for more copies changes how many machines hold the record before the writer is answered, not whether any of them has put it on persistent media. Both settings exist because they answer different questions, and quoting a strict acknowledgement rule as proof of durability is the common error.

An acknowledgement is a receipt from the sorting room, not from the vault: the parcel is inside the building and logged, but it has not gone behind the steel door yet, and a fire in the sorting room takes everything still on the bench.

saying these in an interview costs you the question

  • Assumes an acknowledgement means the bytes reached the device
  • Thinks unforced records live inside the broker process and die with it
  • Treats committed and durable as the same word
  • Believes extra copies make forcing unnecessary in every failure
  • Cannot say how large the loss window is in time or bytes
  • Assumes every platform in this class behaves identically here