Should a cluster node that has lost contact with the coordination role stop serving writes on its own, or rely on later refusal?
answer
- two ends: step aside, or be refused
- both safe, different exposure
- the platform usually decides for you
- what you own is the client contract
basics
~20 sDecide by what a misleading answer costs. Self-demotion after a silence period shortens the time a client can believe it is writing successfully, at the price of stepping aside during blips; relying only on refusal keeps serving through blips but leaves work provisional for longer.
solid answer
~50 sThe two ends are a node that demotes itself after a period without contact, and a node that keeps going until something refuses its writes. Both are safe in the sense that a superseded leader's writes do not become durable — the difference is how long a client can be misled and how eagerly the cluster gives up a node. Self-demotion narrows the misleading period and makes behaviour easy to reason about in an incident, but a healthy node that is briefly isolated stops serving traffic it could have served. Refusal-only preserves availability across brief disturbances and puts the whole burden on what an acknowledgement means to the client. In practice the platform usually chooses for you, and a hosted one certainly does, so the policy you actually own is the client contract: an unacknowledged write is unknown, a locally accepted one is not durable, and no path into the data may skip the ownership check.
go deeper
The takeaway is that a node cut off from the rest of the cluster is in an awkward position, and platforms either have it stand down or have its writes refused later. Both keep the data correct.
Be able to state the trade in one line each: standing down bounds how long a client can be misled but gives up service during blips, while refusing later preserves service and leaves work provisional for longer.
Show where the trade lands in practice — what your writers do with an unacknowledged result, whether their retries re-resolve the owner, and which streams would actually suffer from work being provisional for a while.
Own the part that is still yours. The platform fixes the behaviour; you fix the contract, the per-stream tolerance, and the rule about whether anyone may force a change on a node that is merely unreachable.
## The decision, stated plainly A node leading a partition (or queue) loses contact with the **coordination role** — the part of the cluster that records membership and decides who leads. It does not know whether the rest of the world is gone or whether it is. Two policies are available, and they are ends of a spectrum rather than a binary: - **Self-demotion**: after a bounded period of silence, the node stops serving writes for the units it leads, on its own initiative. - **Refusal-only**: the node keeps serving, and correctness is preserved downstream — its writes carry an out-of-date generation number and are refused wherever they would have to become durable. Both are sound. Neither lets a superseded leader's records enter the accepted history. The choice is about **how long a client can be told something that later turns out not to have happened**, and about how readily the estate gives up capacity. ## What each one costs | | Self-demotion after silence | Refusal-only | |---|---|---| | Time a client can be misled | bounded by the silence period | until a write is refused or times out | | Behaviour during a brief disturbance | the node steps aside, and service for its units pauses | the node keeps serving and usually rides it out | | Failure mode to fear | a healthy node removing itself during a monitoring wobble | work accepted locally for a while and then discarded | | Ease of reasoning in an incident | high — a silent node is also an idle node | lower — a reachable node may be doing nothing useful | | Dependence on client discipline | lower | high | The asymmetry worth naming: self-demotion converts an uncertainty into a small, deliberate availability loss, while refusal-only converts it into a larger uncertainty that somebody else — the client author — has to handle correctly. ## What you can actually choose Be honest about the level at which this is decided: 1. **The platform's design** usually fixes it. Some products demote on lost contact; some rely on rejection at the replication path; designs with detached storage enforce it at the storage layer and the node's own opinion matters even less. 2. **A hosted offering** chooses and may not document it. What you can still hold it to is the outcome — a superseded owner's writes must not become durable. 3. **What is genuinely yours** is the contract you impose on everything that writes, and the paths you allow to exist at all. So a policy written as "our brokers shall self-demote within N seconds" is often unimplementable. A policy written as what teams must assume is always implementable. ## The policy worth writing - **Acknowledgement is the only evidence.** A write is durable when the cluster acknowledged it under current authority. Locally accepted is not durable, and a log line on the writer is not an acknowledgement. - **An unacknowledged write is unknown, not failed.** Every writing service must be safe when the record turns out to exist after all, and safe when it does not. - **Retry means re-resolve.** A retry that does not re-resolve which node owns the unit is a defect, not a tuning choice. - **No unstamped path into the data.** Repair tools, imports and restores bypass the ownership check by their nature; each needs a named owner, a reason, and ideally a way to run them only when the cluster is quiet. - **State the silence tolerance per stream class, not per estate.** A stream carrying payment instructions and one carrying page views do not deserve the same answer about how long a node may keep serving in the dark. ## Where a principal earns the title The interesting judgment is not picking a side; it is noticing that the question is usually settled by a purchase rather than by a standard, and then defending the part that is still in your hands. Two moves do most of the work. The first is separating the streams whose provisional writes would be expensive from those where they would not, and accepting different behaviour for each rather than one estate-wide number that satisfies nobody. The second is refusing to let an ambiguity be resolved by an operator under pressure — an incident is exactly the moment when somebody proposes to force a change on a node that is merely unreachable, and whether that is ever allowed, and who may authorise it, belongs in a standard written on a calm day. The failure this policy prevents is mundane and common: a team ships a writer that logs success on a local acceptance, an incident isolates a node for ninety seconds, and months later nobody can explain a gap in a stream that every dashboard says was written.
- Why is a single estate-wide silence tolerance usually the wrong answer?Because the cost of a provisional write is not uniform. A stream carrying financial instructions justifies stepping aside quickly, while a high-volume telemetry stream would rather ride out a disturbance than lose service to one. One number optimises for neither and is defended by nobody.
- If both policies are safe, why does the choice matter at all?Safety here means a superseded owner's records never enter the accepted history. It does not mean nobody is misled: under refusal-only a client can believe for some time that work landed. The choice sets who absorbs that ambiguity — the cluster, by standing down, or every writing service.
- What should the standard say about forcing a change on a node that is merely unreachable?It should say whether it is ever permitted, who may authorise it, and what evidence is required first — written before an incident, not during one. Pressure to act on a node nobody can reach is exactly when an unwritten rule becomes whatever the loudest person believes.
saying these in an interview costs you the question
- Claims self-demotion on lost contact costs nothing in availability
- Assumes refusal-only needs no discipline from the writing client
- Believes a hosted platform lets you choose this behaviour
- Treats a locally accepted write as an acknowledged one
- Sets one silence tolerance for every stream in the estate
- Says only one of the two policies is actually safe