skip to content

Every reader group's record lag is climbing and the delivery interval has doubled. Why read the cluster's own health signals before the reader numbers?

level: seniorimportance: should knowfreq 62%

answer

  1. consequence against cause
  2. everyone at once is a tell
  3. four readings, no baselines
  4. detectors, not discriminators
  5. check the change window first

basics

~20 s

Because when a cluster signal is bad, every downstream number is a consequence rather than a cause. Reading the reader-side numbers first sends investigation to the wrong teams for what is really one unserved partition, one lagging copy set, or a frozen coordination plane.

solid answer

~50 s

The cluster's own readings — unserved partitions, copies behind, coordination-plane health, and headroom — are read first because each one, when bad, produces exactly the symptom being reported. A partition nobody serves stops every reader assigned to it, so a group's record lag climbs without a single reader being slow. A shortfall of current copies can throttle or reject the write path where a floor is enforced. A coordination plane that cannot decide leaves failures unrepaired and clients holding stale routing. Headroom running out degrades a node in ways that look like anything else. The fact that every group degraded at once is itself the tell: independent readers rarely slow together. Reading the cluster list first does not rule out a reader-side cause; it re-orders the search so the cheapest, highest-yield explanations are eliminated in a minute.

go deeper

for a junior

Recall the order: check whether the cluster itself is healthy before blaming whoever is reading from it, because a cluster fault makes every reader look slow.

for a middle

Explain the mechanism for at least two of the four readings — for instance that an unserved partition stops its readers entirely, so a group's record lag climbs with no reader being slow.

for a senior

Demonstrate the sweep under pressure: four readings, the change-window question, and the explicit hand-off to the reading and processing side when the list comes back clean.

for a principal

Set the contract: which four readings the platform team guarantees to publish for every cluster, so that every incident in the estate starts from the same minute-long sweep rather than from local habit.

## Why order matters at all During an incident every dashboard is bad at once, and the expensive mistake is not missing a number — it is reading a true number as a cause when it is a consequence. "All readers are behind" is true and nearly useless on its own, because it is produced by a reader-side problem, a processing-side problem, and at least four distinct cluster-side problems. Fixing the order of reading is the cheapest possible improvement to time-to-diagnosis, because the cluster list is short, takes seconds, and has no ambiguity in it. There is also a signal in the coincidence itself. Independent reader groups, owned by different teams, running different code, rarely start falling behind within the same minute by chance. **Simultaneity across unrelated consumers is evidence of a shared cause**, and the only things they share are the cluster and the network between them and it. ## The short list, in order 1. **Unserved partitions.** Any non-zero value outside a change window. This is the only reading on the list that means something is already failing rather than degrading. 2. **Copies behind.** Whether the caught-up copy sets are short, and for how long — transient during a change, significant when flat. 3. **Coordination-plane health.** Whether there is exactly one agreed decision-maker, whether the role set is at full membership, and whether its own state is still advancing. 4. **Headroom, per node.** The worst node's projected time to volume exhaustion and its remaining open-file and connection handles. Four readings, and none of them needs a baseline or a judgement call to interpret. ## What each one does to the numbers being reported | Cluster reading | Effect on a reader group's record lag | Effect on the delivery interval | |---|---|---| | A partition nobody serves | Climbs without bound for whichever members hold it; the group's total looks like general slowness | Unmeasurable for that partition — nothing arrives to be timed | | Caught-up copies short | Usually none, unless the write-path floor is crossed and writes are rejected | Rises if writes are being slowed or retried before acceptance | | Coordination plane frozen | None at first; then unbounded on partitions whose failure was never repaired | Rises for anything relying on a routing answer that has gone stale | | Headroom exhausted on a node | Climbs for every group holding partitions on that node | Rises for every stream with a partition there | The left column is four different incidents. The right two columns are the same two symptoms. That is the whole argument for reading left to right. ## Why the reader-side numbers are still worth having None of this says the reader numbers are unreliable — they are the reason anyone noticed. The point is what they can and cannot distinguish: - A group's record lag says something is not keeping up. It does not say whether records are not arriving, not being handed over, or not being processed. - The delivery interval says the whole path got slower. It does not say which hop. - Both are **aggregates over a path that includes the cluster**, so they inherit every cluster fault as their own symptom. So they are excellent detectors and poor discriminators, which is precisely the profile of a signal you alert on and then stop consulting. ## The change-window caveat An announced change legitimately moves several of these at once: a rolling restart pushes the copies-behind reading up repeatedly, briefly moves partitions between nodes, and will make readers stall for as long as reassignment takes. An operator who reads the four cluster signals during a change window and concludes "incident" has read them correctly and interpreted them wrongly. Knowing whether a change is in progress is therefore part of reading the list, not context you fetch afterwards, and it is the first question to ask when several readings move together in a tidy pattern. ## When the cluster list comes back clean If all four readings are normal and have been normal through the period in question, the cause is on the reading or processing side rather than in the cluster's own organs, and the investigation moves there: what the readers are doing with the records, whether the work downstream of them is completing, and whether anything is actually progressing at the far end. That is a different discipline with its own traps — most notably that a cluster can read perfectly healthy while no useful work is being done anywhere — and the value of the four-reading sweep is that it takes about a minute to reach that conclusion honestly rather than assuming it. ## The habit Read the organs, then the flow, then the work. The organs are few, unambiguous and fast; the flow is sensitive but non-specific; the work is where the business impact actually lives. Inverting that order is how an incident spends forty minutes in a reader team's code for a partition that nobody has been serving since the deploy.

  • What does it mean when only one reader group is behind and the four cluster readings are normal?
    The shared explanations are eliminated, so the cause is specific to that group: its own capacity, its processing time per record, or a downstream dependency it waits on. That is a genuinely different investigation from a cluster incident, and reaching it in a minute rather than an hour is the whole point of the sweep.
  • Why is the coincidence of several cluster readings moving together sometimes reassuring rather than alarming?
    Because a tidy, repeating pattern across copies behind, brief reassignments and stalled readers is the signature of a planned change rather than a fault. Checking whether a change window is open is part of reading the list; the same readings mean something quite different inside one.

saying these in an interview costs you the question

  • Starts in the reader team's code before checking the cluster
  • Reads a group's climbing record lag as proof the reader is slow
  • Treats simultaneous degradation across teams as coincidence
  • Ignores whether a planned change window is in progress
  • Assumes a normal cluster reading proves nothing is wrong anywhere
  • Checks headroom as a cluster average instead of per node