skip to content

Operating Signals & Alerting

The handful of numbers that say whether a broker cluster itself is healthy, and which of them deserve to wake a human. Asked because dashboards stay green while nothing is being delivered.

part ofBroker & streaming operationsoverview, primer and where to startread it →
on this pageshow

questions

21

A cluster's health list shows one partition with no node currently serving it. What happens to writes and reads for that partition?

level: juniorimportance: must knowfreq 74%

answer

  1. nobody owns that slice
  2. failure, not slowdown
  3. one partition, whole cluster average
  4. availability now, not margin
  5. steady-state value is zero

basics

~20 s

Writes and reads for that partition fail: with no node holding the job of serving it, there is nothing to accept or answer requests. The stream's other partitions keep working, so the outage is a slice, not the whole cluster.

solid answer

~40 s

An unserved partition means no node currently owns that partition, so requests routed to it cannot be accepted or answered — clients see errors or time out, depending on how the client handles an unroutable request. It is availability loss that is already happening, which is what separates it from a shortfall of caught-up copies, where a node is still serving and only the survivable failure count has shrunk. Because a partition is one slice of a stream, aggregate throughput dips rather than collapsing, so an averaged dashboard can hide it — which is exactly why the count sits on the health list as its own number. In steady state the only acceptable value is zero.

go deeper

for a junior

Recall the plain meaning: no node is serving that partition, so writes and reads for it fail, while the rest of the stream keeps working. Say that the healthy value is zero.

for a middle

Explain why the cluster-wide throughput panel barely moves while one slice is fully down, and why this reading is availability loss whereas a shortfall of caught-up copies is only reduced margin.

for a senior

Show the first-minute sequence: which partitions, which node vanished, whether the coordination role set can reassign at all, and only then the reader side. Name the blast radius as a keyspace slice.

for a principal

Weigh what the estate publishes: whether every cluster exposes this count in the same shape, and what a rented cluster that hides it forces stream owners to infer from request errors instead.

## What the reading is A **stream** is a named, durable channel of records. Most platforms in this class cut a stream into **partitions** so several nodes can carry one stream between them, and each partition is assigned to a node that serves it: that node accepts writes for it and answers reads against it. An **unserved partition** is a partition for which no node currently holds that job. The health list carries it as a **count across the whole cluster** — how many partitions are in that state right now. Outside a deliberate change, the only acceptable steady-state value is zero, and that is what makes it such a cheap reading: there is no threshold to argue about and no baseline to learn. ## Why this is availability, not durability The two headline cluster readings are often quoted together, and they mean different things: | Reading | What it says | What is happening to clients | |---|---|---| | Copies behind | Fewer current copies of a partition exist than configured | Usually nothing yet — the partition is still served, but it survives fewer failures | | Unserved partition | No node is serving a partition at all | Requests for that partition fail now | A shortfall of caught-up copies is **erosion of margin**. An unserved partition is **loss already suffered**. That is why it is read first during an incident and why it is never left to a long confirmation period. ## How a partition ends up unserved - The node serving it is gone — crashed, fenced off, or stopped — and no replacement has taken over yet. - A replacement exists but is not considered eligible to take over, so the cluster refuses to hand the partition to it. (How a replacement is chosen and promoted is a separate subject; from the health list you only see that it has not happened.) - The coordination role set that assigns partitions cannot make a decision, so nothing is reassigned. - A change was applied that left a partition assigned to a node that no longer exists. In every case the reading is the same and the first diagnostic move is the same: find out **which** partitions and **which** node last served them, because the count alone does not say. ## Why the aggregate dashboard hides it If a stream has fifty partitions and one is unserved, total accepted throughput falls by roughly a fiftieth, minus whatever clients retry. On a cluster-wide throughput panel that is noise. But for every client whose records route to that partition — on platforms that route by key, that is a fixed, repeatable slice of the keyspace — the service is fully down, and it stays down until the partition is served again. The same is true on the reading side: a reader assigned that partition stops making progress entirely while its peers look fine, so a per-group record-lag average is also misleading. Counting unserved partitions directly is the only way this shows up honestly. ## How different platforms express it The reading exists in some form wherever a partition or a queue has an **owning node**, but its shape varies: - Where a partition has a leader and a set of follower copies, an unserved partition is a partition whose leadership is vacant. - Where a queue lives on one node with mirrors elsewhere, the equivalent is a queue whose host is down and whose mirror has not been accepted as the new host. - Where storage is shared rather than replicated per record, a partition can be unserved even though none of its data is at risk — the data is intact and merely unreachable. - On a rented cluster the count may not be exposed at all; what you see instead is the request error rate for the affected stream, which is the same fact seen from outside. It is worth knowing which of these your cluster is, because the recovery time differs sharply: promoting an existing current copy is fast, and rebuilding a partition from a surviving copy is not. ## Reading it in the first minute 1. Confirm the count is non-zero and get the list of affected partitions and streams — the blast radius is the set of keys that route there, not the cluster. 2. Check whether a node disappeared at the same moment; a simultaneous drop in the node count makes the story obvious. 3. Check whether the coordination role set is healthy, because if it cannot make decisions the partition will not be reassigned no matter how many candidate copies survive. 4. Only then look at the reader side: readers stuck behind an unserved partition are a symptom, not a second incident. The point to carry out of this: an unserved partition is not a degraded partition. There is no slower path and no queue holding requests for later. For that slice of the stream, the service is off.

  • A client keeps sending records to a stream while one of its partitions is unserved. What happens to those records?
    Records routed to the unserved partition are never accepted, so nothing durable exists for them. Depending on the client they surface as immediate errors or as timeouts after a buffer fills; records routed to the stream's other partitions are accepted normally. The cluster does not hold them aside and deliver them later, so a client that discards failures loses them.
  • Why does the count of unserved partitions belong on the health list even when it is nearly always zero?
    Because it is a reading with no baseline to learn and no threshold to argue about: any non-zero value outside a change window means part of a stream is down. Signals that need tuning get tuned away; this one is checked in a second and is wrong only when something really is wrong.

saying these in an interview costs you the question

  • Thinks the cluster buffers writes for a partition nobody serves
  • Believes records are automatically rerouted to a healthy partition
  • Reads a small dip in total throughput as proof nothing is down
  • Confuses an unserved partition with a partition short of copies
  • Assumes an unserved partition always means its data is lost
  • Goes to the reader teams first when readers on that partition stall
open as a page

A stream's delivery interval is four minutes while its writes are acknowledged in six milliseconds — what is the interval measuring?

level: juniorimportance: must knowfreq 62%

basics

~20 s

The delivery interval covers the whole path: from a write being accepted to a reader finishing the work on that record. Acknowledgement time covers only the first hop, so the wait in the stream and the reader's own processing are what make up the four minutes.

open as a page

A node's mean publish latency is 9 ms while its p99 is 700 ms — which number does a writing client feel, and why?

level: juniorimportance: must knowfreq 70%

basics

~20 s

The p99. Most requests a node serves are cheap, so the mean tracks them and stays flat, while the tail is where a writer actually waits — and one writer in a hundred waiting 700 ms is what a customer reports.

open as a page

When a messaging cluster reports every signal it publishes as normal, what does that actually guarantee about the pipeline's work?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A green cluster dashboard guarantees only that the broker accepted writes, stored them, served reads and recorded that readers moved forward. It says nothing about whether any record was understood, transformed, or turned into useful work downstream.

open as a page

An alert on a reader group's record lag fires on a single sample and pages nightly on write bursts; what does a usable condition need instead?

level: middleimportance: must knowfreq 62%

basics

~20 s

A usable condition names three things: which signal it reads, the threshold that counts as a breach, and a sustained window the breach must hold across. Size that window longer than one full write-and-drain cycle of the flow it watches.

open as a page

Forty partitions report a caught-up copy set smaller than their configured copy count for an hour. What does that reading mean?

level: middleimportance: must knowfreq 70%

basics

~20 s

Those partitions are still being served but now survive fewer failures: copies that should be current are not keeping up. A brief spike during a change is normal; an hour with no change in progress means something is not catching up.

open as a page

A node reports wait time and service time per request — what does each measure and what does a rise in each mean?

level: middleimportance: must knowfreq 62%

basics

~20 s

Wait time is how long a request sat in the node's request queue before a handler took it; service time is the handling itself. Rising wait means too little handler capacity for the arrival rate; rising service means the work got slower.

open as a page

A reader group's record lag sits near zero all day while no business work completes — how can both be true at once?

level: middleimportance: must knowfreq 55%

basics

~20 s

Near-zero reader record lag says records are being read and progress is being recorded — not that the work succeeded. A reader can move past every record while the processing throws, writes nowhere, or is rejected downstream.

open as a page

Why does a broker cluster's health list track free space as projected time to exhaustion rather than percent used?

level: middleimportance: should knowfreq 55%

basics

~20 s

Because a percentage says nothing about how long you have. The same 80% is months away on a quiet stream and an hour away on a busy one; a projection from the measured fill rate turns the reading into a deadline you can act on.

open as a page

How does a record carrying its writer's timestamp rather than the node's arrival timestamp change a measured delivery interval?

level: middleimportance: should knowfreq 48%

basics

~20 s

A write-side timestamp puts the writer's own accumulation, batching and retries inside the measured interval and trusts a clock you do not control. An arrival timestamp is set by the cluster on acceptance, so it is comparable across records but hides everything that happened before the record was accepted.

open as a page

Why is one aggregate request-latency percentile per broker node a weaker signal than separate publish and fetch series?

level: middleimportance: should knowfreq 46%

basics

~20 s

Because writes, reads and metadata calls do unrelated work with unrelated costs. One percentile over all of them moves when the traffic mix changes, not only when something slows, and it never says which side of the node is hurting.

open as a page

Every planned node-by-node cluster change pushes the copies-behind signal past its threshold and pages the on-call, so the team keeps raising the threshold; what should they change instead?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Raising the threshold trades away the detection the alert existed for at every other hour. Split the signals a planned change is expected to move from those it never should, size the sustained window past the normal catch-up time, and bound any suppression to the announced change window.

open as a page

A nightly reader's unread backlog legitimately grows to millions of records and drains by morning, so no fixed record threshold works; how would you phrase the alert condition?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Alert on the shape of the curve rather than its height: fire when the unread backlog fails to fall across a sustained window, when it is still above a level at a deadline the owner states, or when the reader's stored read position stops advancing.

open as a page

A cluster's clients are served normally, but the role set holding its metadata has no agreed decision-maker. Why does that still top the health list?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Because the cluster has lost the ability to change, not to serve. Existing assignments keep working, so request dashboards stay green, while the next node failure cannot be repaired — leaving the cluster one ordinary fault away from unserved partitions.

open as a page

Every reader group's record lag is climbing and the delivery interval has doubled. Why read the cluster's own health signals before the reader numbers?

level: seniorimportance: should knowfreq 62%

basics

~20 s

Because when a cluster signal is bad, every downstream number is a consequence rather than a cause. Reading the reader-side numbers first sends investigation to the wrong teams for what is really one unserved partition, one lagging copy set, or a frozen coordination plane.

open as a page

Delivery intervals computed as the reading host's clock minus a record's write-side timestamp sometimes come out negative — what is actually being measured?

level: seniorimportance: should knowfreq 52%

basics

~20 s

The difference between two hosts' clocks, plus the real interval. A negative result proves the offset between the writing host and the reading host is larger than the journey itself, so every sample from that pair is wrong by that offset — the small ones ruinously.

open as a page

What must a synthetic probe record share with real traffic before it can settle a dispute between cluster-side and reader-side delivery-time numbers?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The same client path, the same cluster, the same acceptance rules and coverage of every part of the stream — written and observed by one host so the timing needs no cross-host subtraction. Anything the probe does not share is a segment it silently excludes.

open as a page

Publish latency is normal but a node's handler-pool saturation climbed from 40% to 85% over an hour — why act on that?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Because occupancy fills before waiting appears. Duration stays flat while a pool has spare handlers and then rises steeply once it does not, so saturation and request-queue depth usually turn first and buy lead time that latency does not.

open as a page

Which pipeline failures can no signal a broker cluster publishes ever reveal, and what makes the cluster blind to them?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Anything that happens after hand-off is invisible: a reader recording progress without doing the work, a transform writing its output nowhere, a destination rejecting every write, a filter dropping everything. The cluster observes transport, not consequences.

open as a page

Across an estate of many streams, which signal honestly reports that a pipeline is working, who must publish it, and what does requiring it everywhere cost?

level: principalimportance: should knowfreq 40%

basics

~20 s

The honest signal is a count of completed business outcomes, published by the team that owns the pipeline, because only that team can define completion. Requiring it estate-wide costs per-pipeline instrumentation, extra stored series, and a definition argument per pipeline.

open as a page

Across an estate of hundreds of streams, which broker alert conditions should a platform team own centrally, which must each stream owner set, and what happens to a stream whose owner sets none?

level: principalimportance: nice to knowfreq 40%

basics

~20 s

Conditions on the cluster's own health are properties of the cluster, so one number serves the estate and the platform team owns it. Conditions on flow are properties of a stream's commitment, so the owner sets the number; an unset stream gets a non-paging default.

open as a page