skip to content

The data volume under a broker node is full — what happens to writes landing on that node, and which streams feel it?

level: juniorimportance: must knowfreq 62%

answer

  1. no free space, no partial answer
  2. refusal, not congestion
  3. the blast radius follows the volume
  4. neighbours stop although they behaved
  5. reads serve, progress writes may not

basics

~20 s

A full data volume is a refusal, not a slowdown: the node stops accepting writes, and it stops accepting them for every stream whose data lives on that volume, not only the one that filled it. Reads usually keep serving.

solid answer

~40 s

A store appends records to files on a data volume, so an accepted write needs free space. When the volume is exhausted the node cannot answer partially — it protects itself by refusing writes. What that looks like differs by design: some platforms mark the node, or every stream on it, unwritable and return an explicit error; some fail each write attempt individually; some hold the caller until its send timeout expires, which the application experiences as a hang. The important part is the blast radius: the refusal follows the **volume**, so streams that were sized correctly stop accepting writes because a neighbour grew. Reads usually keep serving bytes that are already there, and nothing already stored is discarded by the exhaustion itself.

go deeper

for a junior

Recall the shape: no free space means the node refuses writes rather than slowing down, and the refusal covers every stream on that volume. Records already stored are not thrown away.

for a middle

Explain why a partial answer is impossible for an append-only store, and why the refusal reaches clients as an error on some platforms and as a hang on others. Say which side of the traffic survives.

for a senior

Show that you treat it as a shared-capacity incident: name the blast radius across co-located streams, check whether reader progress is itself a refused write, and reach for relief that does not destroy history first.

for a principal

Frame it as governance. A volume several teams share with no bound on any stream is an unpriced shared resource whose failure is collective; the design question is who is allowed to consume it and who pays when they do.

## What a full data volume means to an append-only store A broker or streaming platform holds records by **appending** them to files on a **data volume** — the block device the node writes to. The store rolls those files closed at some size or age; each closed unit is a **segment**, and the one still being appended to is the **active segment**. Every accepted write needs room on that volume, and so does the bookkeeping the store keeps beside the records. When free space reaches zero — or reaches whatever floor the platform reserves for itself so it can still function — there is no partial answer available. The store cannot write half a record and remain an append-only store, and it cannot make room on the write path, because removing history is governed by a policy and a background pass, not by pressure. ## Refusal, not degradation This is the distinction interviewers listen for. A full volume is not congestion that clears when a burst passes. Nothing a writer does will make room. The condition persists until an operator changes something. Designs differ in how the refusal reaches the client, and saying so is part of a good answer: - some platforms flip the node, or every stream held on the exhausted volume, into a state where writes are rejected with an explicit error while reads continue; - some fail each individual write attempt and leave the writing application's retry policy to decide what happens next; - some hold the caller until its send timeout expires, so the application sees a hang rather than an error and only discovers the cause from the node's own signals. In all three shapes the answer to "can this write be accepted?" is no, and retrying harder makes it worse: retries multiply request load on a node that can accept none of it. ## The blast radius is the volume, not the stream | Question an operator asks | What is true | |---|---| | Which streams stop accepting writes? | Every stream whose data lives on the exhausted volume, whichever one grew | | Do reads stop? | Usually not — serving bytes already on the volume needs no new space | | Is anything already stored lost? | No; exhaustion refuses new writes, it does not discard old records | | Does adding writers or retries help? | No — it adds load to a node that can accept nothing | | Whose incident is it? | Everyone's on that node, which is why it escalates faster than teams expect | This is the operational point of the whole subject. One ungoverned stream, one backfill nobody warned about, one team's experiment — and the neighbours that were sized correctly go down with it. The mitigation is therefore a capacity story about the volume, not a tuning story about the stream that misbehaved. ## The read path is not entirely safe "Reads keep working" needs one qualification. On platforms where a reader's progress is recorded back onto the cluster — sometimes into a stream of its own that lives on the same volume — a node refusing writes also refuses to record that progress. Readers then appear healthy, consume records, and re-consume the same records after a restart, because nothing durable recorded how far they got. On platforms where progress is held client-side or in an external store, that symptom does not appear. Expect the difference rather than assuming either. ## Why the volume filled Exhaustion is almost never a surprise in the arithmetic; it is a surprise in **eligibility**. The bytes on the volume are the bytes the store was not allowed to remove: history still inside its configured bounds, plus anything a rule or an un-advanced reader pinned there, plus whatever else shares the device. A volume that several streams share with no bound on any of them has no defence at all — the sum of what may be stored is simply unbounded, and the first stream to grow decides everyone's fate. ## What it costs to fix in a hurry The fastest way to restore writes is to shorten **the retention window** — the span of history still readable on the store right now. That window *is* the replay budget: everything a reader could still be rewound into, everything an investigation could still read, everything a rebuild could still consume. Shortening it converts history into free space immediately and permanently. That makes it powerful and makes it the last rung, not the first: relief that does not destroy history — more space, moving the pressure elsewhere, releasing bytes something is pinning — is tried before the one action that cannot be taken back. A junior answer that gets "it refuses writes, for everything on that volume, and reads generally survive" is already a good answer. The rest is what seniority adds.

  • Why does the store not simply overwrite its oldest records to make room?
    It is an append-only store, not a ring buffer. Which records may go is decided by a configured bound and carried out by a background pass over whole segments; overwriting on demand would silently move the earliest readable position under every reader, destroying replay guarantees exactly when something is already wrong.
  • Can readers still make progress on a node that is refusing writes?
    They can usually still fetch records, because serving existing bytes needs no space. Where the platform records reader progress back onto the cluster, that record is itself a write and can be refused, so readers consume normally and then re-consume after a restart. Where progress is held elsewhere, this does not happen.
  • Does the node recover on its own once the burst passes?
    No. A burst that passes leaves the bytes behind; nothing on the write path frees space. Recovery needs either the background removal pass to become able to remove something, or an operator to add space, move data, or release what is pinning it.

saying these in an interview costs you the question

  • Says the node just gets slower until space frees up by itself
  • Assumes only the stream that grew stops accepting writes
  • Expects the store to overwrite the oldest records to make room
  • Thinks retrying more aggressively will get the write accepted
  • Declares a total outage without checking that reads still serve
  • Believes records already stored are discarded when the volume fills