skip to content

Why does a broker cluster's health list track free space as projected time to exhaustion rather than percent used?

level: middleimportance: should knowfreq 55%

answer

  1. how long, not how full
  2. fill rate makes the deadline
  3. retention is the subtracted term
  4. handles run out on bytes-free clock
  5. worst node, not the average

basics

~20 s

Because a percentage says nothing about how long you have. The same 80% is months away on a quiet stream and an hour away on a busy one; a projection from the measured fill rate turns the reading into a deadline you can act on.

solid answer

~50 s

A broker node's volume drains at a rate you can actually measure: the bytes written per second, multiplied by the copies the node holds, minus whatever retention removes. Dividing the free space by that net rate gives a time to exhaustion, which is the form an operator can act on — it says whether there are hours or weeks, and it stays meaningful when nodes have different volume sizes. A percentage does not, because the same figure means wildly different things at different write rates. Its twin on the health list is handle headroom: the remaining open-file and connection handles on a node, which is consumed by connection and partition counts rather than by bytes and therefore runs out on a completely different clock. Both are read per node, because the fullest node is what matters, not the cluster average.

go deeper

for a junior

Recall that a node stores records on a volume that can fill up, and that knowing how fast it is filling matters more than knowing how full it is right now.

for a middle

Explain the arithmetic: bytes written, times the copies this node holds, plus overhead, minus what retention removes — and why a balanced cluster has no meaningful deadline at all.

for a senior

Show why the reading is per node and why handle headroom belongs beside it: a node out of handles presents as a network or storage fault and is diagnosed for far too long.

for a principal

Argue what the platform team publishes: a fleet-minimum projection as a scheduling input, so capacity work is planned rather than triggered by a page at four in the morning.

## Why a percentage is the wrong form A node's data volume is consumed by an ongoing process, not by a one-off event. So the useful question is never "how full is it" but "how long until it is full". Those two questions have very different answers from the same percentage: - A node at 80% whose free space is shrinking by a percentage point a week has about five months. - A node at 80% receiving a backfill at full line rate may have under an hour. - A node at 40% whose oldest records are about to age out may never fill at all, because retention is about to return space. A percentage collapses all three into one number. A projection separates them, and a projection is what a human can act on, because it converts the reading into a deadline. ## The arithmetic behind the projection The rate is not the raw write rate of the stream. What a single node stores per second is roughly: - the bytes written per second to the partitions it serves, **plus** - the bytes it receives as a follower copy of partitions led elsewhere, **plus** - any index and bookkeeping overhead the platform keeps alongside the records, **minus** - the bytes retention removes per second as old records age out or are compacted away. That last term is what makes the projection honest and also what makes it fragile: a cluster in steady state is roughly balanced, with removal matching arrival, and the projection is effectively infinite. What produces a finite, shrinking projection is an imbalance — a raised retention period, a new stream, a doubled write rate, a copy being rebuilt onto this node, or removal that has stopped. So the projection is best read as a **change detector**: any transition from "no meaningful deadline" to "eleven days" is the event, whatever the percentage says. ## Handle headroom: the twin on a different clock The second headroom reading is the number of open-file and connection handles a node may still take, against the cap the operating system enforces. It matters because it is exhausted by things that have nothing to do with bytes: | Reading | Consumed by | Grows with | Typical symptom when exhausted | |---|---|---|---| | Free space | Bytes stored, times copies held, minus retention | Write rate and retention period | The node can no longer store what it accepts | | Handle headroom | One or more handles per connection and per open file of stored data | Client count and partition count | The node refuses new connections or cannot open new storage files | The consequence of missing the second one is that the failure looks like something else entirely. A node that cannot accept a connection reads to a client team as a network fault or an authentication problem; a node that cannot open a new file for the next segment of stored data reads as a storage fault even though the volume has ample space. Both are a counted resource quietly hitting a ceiling, and both are cheap to watch. ## Read per node, never per cluster Averaging headroom across the cluster is the classic way to lose it. Partitions are rarely spread perfectly evenly, and a cluster that has been grown by adding nodes typically has an older set holding far more data than the new ones. The reading that matters is the **worst node**, because that is the one that will hit the wall first, and the consequences of one node running out are not confined to that node — its partitions are the ones that stop being served properly. So the health list carries the minimum projection and the minimum handle headroom across the fleet, not the mean of either. ## What the projection cannot tell you - It cannot see a step change that has not happened yet — a scheduled backfill, a retention change approved for next week, a new stream about to be created. - It assumes the current rate continues, so it swings wildly during a traffic spike; smoothing the rate over a sensible interval is part of making it usable. - It does not say what the node will actually do when the space runs out, which depends entirely on the platform's design and is a separate subject. - It does not distinguish space that retention will return soon from space that will not be returned at all, unless the removal term is measured rather than assumed. ## The habit to carry Headroom readings are the only ones on a cluster health list that are genuinely predictive: copies behind and unserved partitions tell you about something that has already happened, while a projection tells you about something that has not. That is exactly why they belong on the list despite never being an emergency at the moment you read them — they are the readings that let a team schedule work instead of being woken by it.

  • Why can a node's projected time to exhaustion be effectively infinite even at high write rates?
    Because retention removes bytes at roughly the rate they arrive. In steady state the volume is a moving window over recent records, so occupancy plateaus. A finite projection appears only when something breaks that balance — a longer retention period, a new stream, a rebuilt copy landing on the node, or removal that has stalled.
  • A node is refusing new client connections while its volume is half empty. Which headroom reading explains that?
    Handle headroom. Connections and open storage files each consume handles against an operating-system cap that has nothing to do with free space, so a node with plenty of volume can still be unable to accept another connection or open the next file. It reads as a network or storage fault until you check the count.

A fuel gauge reading a quarter tank tells you almost nothing on its own; the range-to-empty figure, computed from how fast you are actually burning fuel, is what decides whether you pass this service station or the next one. Free space is the gauge; the projection is the range.

saying these in an interview costs you the question

  • Alerts on percent used and calls the reading done
  • Averages headroom across nodes and misses the fullest one
  • Forgets follower copies consume space on the node too
  • Ignores handle headroom because the volume looks fine
  • Treats a connection refusal as always a network problem
  • Projects from an unsmoothed rate during a traffic spike