A cluster's health list shows one partition with no node currently serving it. What happens to writes and reads for that partition?
answer
- nobody owns that slice
- failure, not slowdown
- one partition, whole cluster average
- availability now, not margin
- steady-state value is zero
basics
~20 sWrites and reads for that partition fail: with no node holding the job of serving it, there is nothing to accept or answer requests. The stream's other partitions keep working, so the outage is a slice, not the whole cluster.
solid answer
~40 sAn unserved partition means no node currently owns that partition, so requests routed to it cannot be accepted or answered — clients see errors or time out, depending on how the client handles an unroutable request. It is availability loss that is already happening, which is what separates it from a shortfall of caught-up copies, where a node is still serving and only the survivable failure count has shrunk. Because a partition is one slice of a stream, aggregate throughput dips rather than collapsing, so an averaged dashboard can hide it — which is exactly why the count sits on the health list as its own number. In steady state the only acceptable value is zero.
go deeper
Recall the plain meaning: no node is serving that partition, so writes and reads for it fail, while the rest of the stream keeps working. Say that the healthy value is zero.
Explain why the cluster-wide throughput panel barely moves while one slice is fully down, and why this reading is availability loss whereas a shortfall of caught-up copies is only reduced margin.
Show the first-minute sequence: which partitions, which node vanished, whether the coordination role set can reassign at all, and only then the reader side. Name the blast radius as a keyspace slice.
Weigh what the estate publishes: whether every cluster exposes this count in the same shape, and what a rented cluster that hides it forces stream owners to infer from request errors instead.
## What the reading is A **stream** is a named, durable channel of records. Most platforms in this class cut a stream into **partitions** so several nodes can carry one stream between them, and each partition is assigned to a node that serves it: that node accepts writes for it and answers reads against it. An **unserved partition** is a partition for which no node currently holds that job. The health list carries it as a **count across the whole cluster** — how many partitions are in that state right now. Outside a deliberate change, the only acceptable steady-state value is zero, and that is what makes it such a cheap reading: there is no threshold to argue about and no baseline to learn. ## Why this is availability, not durability The two headline cluster readings are often quoted together, and they mean different things: | Reading | What it says | What is happening to clients | |---|---|---| | Copies behind | Fewer current copies of a partition exist than configured | Usually nothing yet — the partition is still served, but it survives fewer failures | | Unserved partition | No node is serving a partition at all | Requests for that partition fail now | A shortfall of caught-up copies is **erosion of margin**. An unserved partition is **loss already suffered**. That is why it is read first during an incident and why it is never left to a long confirmation period. ## How a partition ends up unserved - The node serving it is gone — crashed, fenced off, or stopped — and no replacement has taken over yet. - A replacement exists but is not considered eligible to take over, so the cluster refuses to hand the partition to it. (How a replacement is chosen and promoted is a separate subject; from the health list you only see that it has not happened.) - The coordination role set that assigns partitions cannot make a decision, so nothing is reassigned. - A change was applied that left a partition assigned to a node that no longer exists. In every case the reading is the same and the first diagnostic move is the same: find out **which** partitions and **which** node last served them, because the count alone does not say. ## Why the aggregate dashboard hides it If a stream has fifty partitions and one is unserved, total accepted throughput falls by roughly a fiftieth, minus whatever clients retry. On a cluster-wide throughput panel that is noise. But for every client whose records route to that partition — on platforms that route by key, that is a fixed, repeatable slice of the keyspace — the service is fully down, and it stays down until the partition is served again. The same is true on the reading side: a reader assigned that partition stops making progress entirely while its peers look fine, so a per-group record-lag average is also misleading. Counting unserved partitions directly is the only way this shows up honestly. ## How different platforms express it The reading exists in some form wherever a partition or a queue has an **owning node**, but its shape varies: - Where a partition has a leader and a set of follower copies, an unserved partition is a partition whose leadership is vacant. - Where a queue lives on one node with mirrors elsewhere, the equivalent is a queue whose host is down and whose mirror has not been accepted as the new host. - Where storage is shared rather than replicated per record, a partition can be unserved even though none of its data is at risk — the data is intact and merely unreachable. - On a rented cluster the count may not be exposed at all; what you see instead is the request error rate for the affected stream, which is the same fact seen from outside. It is worth knowing which of these your cluster is, because the recovery time differs sharply: promoting an existing current copy is fast, and rebuilding a partition from a surviving copy is not. ## Reading it in the first minute 1. Confirm the count is non-zero and get the list of affected partitions and streams — the blast radius is the set of keys that route there, not the cluster. 2. Check whether a node disappeared at the same moment; a simultaneous drop in the node count makes the story obvious. 3. Check whether the coordination role set is healthy, because if it cannot make decisions the partition will not be reassigned no matter how many candidate copies survive. 4. Only then look at the reader side: readers stuck behind an unserved partition are a symptom, not a second incident. The point to carry out of this: an unserved partition is not a degraded partition. There is no slower path and no queue holding requests for later. For that slice of the stream, the service is off.
- A client keeps sending records to a stream while one of its partitions is unserved. What happens to those records?Records routed to the unserved partition are never accepted, so nothing durable exists for them. Depending on the client they surface as immediate errors or as timeouts after a buffer fills; records routed to the stream's other partitions are accepted normally. The cluster does not hold them aside and deliver them later, so a client that discards failures loses them.
- Why does the count of unserved partitions belong on the health list even when it is nearly always zero?Because it is a reading with no baseline to learn and no threshold to argue about: any non-zero value outside a change window means part of a stream is down. Signals that need tuning get tuned away; this one is checked in a second and is wrong only when something really is wrong.
saying these in an interview costs you the question
- Thinks the cluster buffers writes for a partition nobody serves
- Believes records are automatically rerouted to a healthy partition
- Reads a small dip in total throughput as proof nothing is down
- Confuses an unserved partition with a partition short of copies
- Assumes an unserved partition always means its data is lost
- Goes to the reader teams first when readers on that partition stall