skip to content

A cluster's clients are served normally, but the role set holding its metadata has no agreed decision-maker. Why does that still top the health list?

level: seniorimportance: should knowfreq 52%

answer

  1. serving continues, changing stops
  2. green dashboards, no repair
  3. exactly one decision-maker
  4. two is worse than none
  5. next node loss is unrecoverable

basics

~20 s

Because the cluster has lost the ability to change, not to serve. Existing assignments keep working, so request dashboards stay green, while the next node failure cannot be repaired — leaving the cluster one ordinary fault away from unserved partitions.

solid answer

~40 s

The coordination plane holds cluster metadata and decides which node serves which partition. While it is unable to decide, nodes keep serving whatever they were already assigned, so throughput, latency and error rate all look normal. What stops is every change: reassigning a partition after a node is lost, admitting a new node, creating or altering a stream, and on some designs recording cluster-level state at all. The cluster is therefore running without its repair mechanism, and the failure that would normally be absorbed in seconds instead becomes an unserved partition. That invisible fragility is why the reading sits at the top of a health list even though nothing has failed yet, and why it is read before any data-path number.

go deeper

for a junior

Recall that something in the cluster has to remember which node serves which partition, and that if it stops working the cluster keeps serving but cannot fix itself.

for a middle

Explain why data-path metrics stay normal: nodes serve from assignments they already hold, so only change operations depend on the plane being able to decide.

for a senior

Demonstrate the operational response — read the three facts, then freeze every planned change, because a restart run against a plane that cannot reassign creates unserved partitions on purpose.

for a principal

Frame the estate question: which clusters expose coordination health at all, and what a managed offering's silence here obliges stream owners to infer from failing change operations.

## What the coordination plane is, for the purpose of this reading Every cluster in this class has some role set that holds authoritative cluster metadata — which streams exist, how many partitions each has, which node serves each partition — and that decides what changes when the membership changes. How that role set is composed, whether it runs inside the serving nodes or beside them, and how many members it needs are questions about cluster shape. For a health list, none of that matters: what you read are a few facts about whether it is currently able to decide. ## The three facts worth reading 1. **Is there exactly one agreed decision-maker?** Not zero, not two. Zero means no change can be made. Two means contradictory changes can be made, which is worse. 2. **Is the role set at full membership?** A role set that needs a majority to act and is one member short is still working and has no margin — the same relationship the caught-up copy set has to a partition. 3. **Is its own state still advancing?** The metadata the plane maintains is itself written and committed. If that commit stops advancing, the plane is present but frozen, and it will accept requests it never completes. Any of the three going wrong presents the same way to clients: nothing. ## Why the data path stays green A node that is already serving a partition does not consult the coordination plane for each request. It serves from an assignment it already holds, and clients that already know where to connect keep connecting there. So for as long as nothing changes, the cluster is entirely functional: - Writes are accepted and acknowledged as usual. - Reads are served at normal latency. - Copies keep replicating from the nodes leading their partitions. - Readers keep making progress and their record lag stays flat. What has been lost is not service but **repair**. The cluster in this state is running without the mechanism that would normally hide a node failure from anybody. ## Why the next fault is the expensive one | Event | With the coordination plane healthy | With it unable to decide | |---|---|---| | A serving node is lost | Its partitions are reassigned in seconds | Those partitions become unserved and stay that way | | A new node is added | Work is placed on it | It sits idle; nothing is assigned | | A stream is created | It exists and accepts writes | The request fails or hangs | | A client asks where to connect | It is told the current owner | It may be told a stale owner | The last row is the subtle one: a cluster that cannot update metadata can keep handing out an answer that used to be true. Clients then fail against a node that no longer serves what they want, and the symptom looks like a client problem. ## Why two decision-makers are worse than none With none, the cluster is frozen but consistent: it makes no decisions, so it makes no wrong ones. With two, each may issue assignments the other does not know about, and serving nodes can act on contradictory instructions — two nodes believing they lead the same partition, or a partition being handed away while it is still being written to. Platforms defend against this with membership rules and with fencing so that a stale decision is rejected, but the health reading exists because the defence is not free and not universal. A count of decision-makers that is anything other than one is an immediate signal. ## What you can see, by platform - Where coordination runs as a **separate service beside the cluster**, its health is a second system to monitor, with its own membership and its own storage, and an operator has to remember it exists at all. - Where coordination is **built into the cluster** as a coordination quorum among some of its own nodes, the readings come from the same place as everything else, and losing too many nodes takes the coordination plane and the data with it. - On a **rented cluster** you usually see none of this. The provider operates it, and the only visible trace is that change operations become slow or start failing while ordinary traffic is fine — which is the same fact, observed from outside. ## How to read it in practice Treat it as a fragility signal rather than an incident signal. The right response is not to reassure yourself from the request dashboard — it will be green — but to stop anything that would provoke a change: pause a rolling restart, pause an expansion, pause a migration. Every one of those depends on the plane being able to decide, and running them against a plane that cannot is how a quiet fragility becomes a partition with nobody serving it.

  • What should an operator stop doing while the coordination plane cannot make decisions?
    Anything that requires a decision: rolling restarts, upgrades, node additions or removals, partition moves, and stream creation. Each of those deliberately removes or adds capacity and then relies on the plane to reassign work. Run them now and the cluster will take the node out and never reassign what it was serving.
  • Why can a client fail to reach a partition while every serving node is healthy?
    Because the metadata it was given is stale. A coordination plane that cannot update its state keeps answering with an assignment that has since changed, so the client connects to a node that no longer serves that partition. The symptom looks client-side, but the cause is the plane's inability to publish a new answer.

saying these in an interview costs you the question

  • Concludes the cluster is healthy because the request dashboard is green
  • Thinks no decision-maker and two decision-makers are equally bad
  • Proceeds with a rolling restart while the plane cannot decide
  • Forgets that coordination may be a separate system to monitor
  • Assumes a rented cluster exposes this reading at all
  • Believes every read consults the coordination plane