skip to content

A node's data volume keeps filling although the configured age bound has not changed and write rates are flat — where do you look, and what is the structural fix?

level: seniorimportance: should knowfreq 55%

answer

  1. ask what may not leave
  2. ineligible bytes, not extra traffic
  3. flat record rate hides byte growth
  4. who pins, who was never bounded
  5. the sum of bounds must fit

basics

~20 s

Look for bytes that are not eligible for removal rather than for extra traffic: records something pins, streams nobody bounded, and payloads that grew while the record rate stayed flat. The fix is a bound per stream whose sum, plus headroom, fits the volume.

solid answer

~40 s

Stop asking what is arriving and start asking what may not leave. With a steady record rate and an unchanged age bound, growth comes from **ineligible or unaccounted bytes**: records a hold protects; on platforms that bound a per-subscriber backlog, bytes a subscriber that has not advanced still holds; a stream created on the same volume with a longer bound or no bound at all; non-store files sharing the device; and payloads that got bigger while the record count did not. The structural repair is not a bigger device. It is that **every stream on a shared volume carries its own bound, and the sum of those bounds plus headroom provably fits the volume** — plus treating an abandoned subscriber as a capacity object with an owner, not as a harmless leftover.

go deeper

for a junior

Recall that a record occupies space until the store is allowed to remove it, and that something can make old records unremovable. Growth without new traffic points at permission, not volume.

for a middle

Explain the eligibility categories: inside the bound, pinned by a rule, pinned by a subscriber where the bound is per subscriber, or never bounded at all. Also that byte growth hides behind a flat record rate.

for a senior

Demonstrate the diagnosis order — what may not leave, what changed on the date growth started, what shares the device — and land on the fix being a bounded, checked sum rather than more space.

for a principal

Own the oversubscription policy: how much the estate may promise on one device, who reviews it, and who is accountable for a stream with no bound or a subscriber with no owner.

## Start from eligibility, not from traffic When a volume grows without the write rate growing, the arithmetic has not changed — the *permission* has. The right first question is: **which bytes on this volume are not eligible for removal, and why?** Everything else follows from the answer, and a diagnosis that starts by hunting for a traffic spike usually finds nothing and wastes the hour you had. ## Where ineligible or unaccounted bytes come from 1. **A subscriber that has not advanced.** This is the classic, and it is also the claim that most needs qualifying, because platforms split here. Where the bound applies to a **per-subscriber backlog**, everything that subscriber has not read is ineligible at any age; a consumer switched off before a holiday, or a test subscriber nobody deleted, holds bytes forever and no retention setting on the stream will move them. Where the bound applies to **the stream itself**, the opposite failure occurs: the pass removes on schedule regardless of who read what, so the volume does not grow but the reader silently loses records. Establish which shape you run before you diagnose, because the two produce opposite symptoms from the same neglect. 2. **A hold that suspends removal.** A rule that some records must be kept regardless of the configured bounds does exactly that, quietly, and often applies to a subset nobody is monitoring separately. 3. **Bytes nobody bounded.** A stream created last month with no bound set, a stream whose bound is far longer than the rest, an inbound mirrored copy landing on the same device, or non-store files — archived diagnostics, crash dumps, old artefacts — sharing the volume. A shared device with no per-stream bound has no defence: the total that may be stored is simply unbounded. 4. **The same records got bigger.** A flat rate is usually measured in records, not bytes. A new field, a larger payload, or compression switched off changes byte volume without changing anything a rate dashboard shows. ## What to check against what changed | Observation | What to establish | |---|---| | One stream's bytes grow, all others flat | Whether that stream has a bound at all, and what scope it applies to | | Bytes grow but the record count is flat | Average record size over the same period | | Growth started on a known date | What was created, held, mirrored or switched off that day | | Oldest data is far older than the bound | Whether a hold or an un-advanced subscriber pins it | | Store accounting disagrees with free space | Non-store files on the same device, or unlinked files still held open | ## The structural fix: bound the sum, not the intent The durable repair is capacity governance on the device: - **Every stream carries a bound.** Not "the volume is big" and not "the teams are careful" — a declared bound per stream, including the ones created by automation. - **The sum of the bounds, plus headroom, fits the volume.** If the arithmetic does not close on paper, the volume is oversubscribed and the only question is which week it fills. - **Headroom is not decoration.** It absorbs removal working a whole segment at a time, any data movement in progress, and the growth between one capacity review and the next. - **An abandoned subscriber is a capacity object.** Where the bound is per subscriber, subscriber lifecycle *is* a storage control: a subscriber with no owner should be retired deliberately, and retiring it is what releases the bytes. - **Isolate what you cannot predict.** A stream whose growth is driven by something outside your control belongs on its own device or node, so that when it misbehaves it takes down only itself. ## What is not the fix Running the removal pass more often does nothing to bytes that are ineligible — the pass is not the constraint, permission is. Buying a larger device without bounding anything buys time proportional to the extra space and changes no outcome. And shortening **the retention window** — the span of history still readable, which is the replay budget — is relief, not a fix: it trades permanent history for space and leaves the same unbounded sum in place to fill the volume again, slightly later, with less budget to investigate it. The answer an interviewer is listening for is the reframing: on a shared device, capacity is not a property of any one stream's settings but of the sum of everything allowed to be stored there, and until that sum is bounded and checked, every correctly-sized stream on the device is one neighbour away from refusing writes.

  • How does an abandoned subscriber become a storage problem rather than a reader problem?
    On platforms where the bound applies to a per-subscriber backlog, everything that subscriber has not read is ineligible for removal at any age, so a subscriber nobody owns pins bytes indefinitely. Where the bound applies to the stream instead, the same neglect produces silently lost records rather than a growing volume.
  • Bytes are growing while the record rate is flat — what explains that?
    Record size. Rate dashboards usually count records, so a new field, a larger payload, or compression switched off raises bytes without moving the rate at all. Compare average record size across the period before concluding that traffic is unchanged.
  • Why is a larger device not the structural answer?
    Because nothing about the failure was the device's size; it was that the total permitted to be stored is unbounded. A larger device postpones the same incident in proportion to the extra space, and leaves every correctly-sized stream still exposed to its neighbours.

saying these in an interview costs you the question

  • Hunts for a traffic spike when the rate is genuinely flat
  • Assumes old records are always eligible for removal
  • Treats an abandoned subscriber as harmless leftover state
  • Proposes a bigger device while nothing is bounded
  • Suggests running the removal pass more often against pinned bytes
  • Reads a flat record rate as proof that bytes are flat