Why does losing the state store behind a cluster's API hurt differently from losing a node that runs workloads?
answer
- hosts are fungible, the record is not
- every decision reads the store
- the fleet outlives the record
- a restore re-declares an older intent
- images and volume data are not in it
basics
~20 sA node is interchangeable: its work is placed elsewhere. The state store is the cluster's only record of what should exist and what was last reported, so losing it leaves the deciding half with nothing to act on, even while the containers keep running.
solid answer
~40 sHosts are designed to be fungible — losing one costs you its copies, and the deciding half places them somewhere else. The state store is not fungible in the same way: it is the durable record of every declared object and the last status reported about it, and every decision in the cluster reads or writes it. Lose it and the workloads keep serving, but the cluster no longer knows what it is supposed to be running. Recovery is a restore, and a restore is not a rewind: it re-declares an older intent, which the deciding half then enforces on a fleet that has moved on. That is why the store's backup and restore path is planned and rehearsed, while a lost host is just an event.
go deeper
Remember what the store is for: it holds what the cluster was told to run. Losing a host loses some copies; losing the store loses the knowledge of what should exist.
Explain that every decision reads and writes the store, so degrading it degrades change while leaving running workloads alone — a cluster that is fine until you try to alter it.
Show that recovery is a restore of intent, not a rewind: counts revert, recent objects are unknown, deleted things return, and the fleet is then moved toward that older record.
Decide where the authoritative copy of declared state lives across the estate, so a cluster's store is a cache of an intent you hold elsewhere rather than the only copy of it.
## Two very different losses | | Losing one host | Losing the state store | |---|---|---| | What is lost | The copies that were running on it | The record of what should be running anywhere | | What still serves | Everything on every other host | Everything, everywhere — for now | | Who repairs it | The deciding half, by re-placing the copies | An operator, from a backup | | Cost of the repair | Minutes, automatic | A restore, then a reconciliation of intent against reality | A host is one of many, and the whole point of declaring a copy count rather than naming machines is that any host can be replaced by any other. The store has no such peer in the model: it is the cluster's memory. Everything the deciding half does begins by reading it and ends by writing to it. ## Why the store sits on the critical path of every decision The store holds two things that look similar and are not: the **declared state** you wrote, and the **last reported status** the node agents sent in. Every placement, every copy-count change, every controller pass, and every status read passes through it. That has a consequence teams meet long before they ever lose the store: **a degraded store is a degraded cluster, in one specific direction.** Running workloads are untouched, but every change gets slow — applying a spec waits, controllers fall behind, newly declared work queues before anyone places it, and the status you read ages. The symptom is a cluster that seems perfectly healthy until you try to change something. (How the store's copies agree on a write among themselves is a distributed-systems subject in its own right; what matters here is that the deciding half is only as available as its record.) ## Restoring an older copy is a declaration, not a rewind This is the part that surprises people, and it is the best follow-up in this material. While the store was gone, the fleet kept running. When a backup from some hours earlier is restored, the cluster does not travel back in time — it acquires an *older intent*, and the deciding half immediately begins making reality match it. Concretely: 1. **Counts revert.** A workload scaled from four copies to twelve after the backup is declared as four again, and the extra copies are candidates for removal. 2. **Recent objects are unknown.** Workloads created after the backup are not in the restored record. Platforms differ on what happens to them — some tear down what they do not recognise, some leave them running but unmanaged — and it is worth saying that rather than guessing. 3. **Configuration and credentials go back.** Values changed after the backup are re-declared at their old contents, and workloads that read them fresh will pick the old ones up. 4. **Deleted things come back.** Something removed after the backup is declared again, and the deciding half dutifully recreates it. So a restore has to be treated as a change to the whole cluster's intent, not as a repair of one component. ## What is not in the store Another reason the two losses are different classes: the store does not hold your data or your images. - **Images** live in a registry outside the cluster, so a restore does not restore them and losing the store does not lose them. - **Workload data** lives in volumes and in the backing stores behind them, on their own lifecycle entirely. - **What is in the store** is intent and reported status — which is exactly what you cannot rebuild by looking at a running host. The practical corollary is that a cluster's restore plan and a database's backup plan are separate plans, and neither substitutes for the other. ## What teams do about it - **Keep the declared state somewhere you own outside the cluster**, so the record can be re-applied rather than only restored. A cluster whose intent exists only inside its own store has a single copy of something irreplaceable. - **Test the restore**, including the part after the restore: what the deciding half does to a running fleet once the older intent is in place. - **Watch the store's growth and write rate.** Very large numbers of objects, very large individual objects, and very chatty status writes all push it toward the degraded-but-alive state, which is the harder failure to diagnose. - **Treat the running fleet as evidence.** When the record and reality disagree after a restore, the fleet is what your users are actually talking to — reconcile deliberately rather than letting the older intent take effect unexamined.
- After a restore, the record and the running fleet disagree — which one wins?The record. The deciding half acts on declared state, so it moves the fleet toward the restored copy: counts revert, deleted objects come back. Workloads created after the backup are unknown to it, and platforms differ on whether those are torn down or simply left running unmanaged, so plan for both.
- Does a slow state store show up first as a workload problem?No. Running workloads are unaffected, because they need nothing from the store. What degrades is change: applying a spec waits, controllers fall behind, newly declared work queues, and status reads age. The signature is a cluster that looks healthy until someone tries to change it.
saying these in an interview costs you the question
- Says losing the state store stops the running workloads.
- Thinks restoring an old copy of the store rewinds the running fleet.
- Treats a slow store as a problem only for operators' own commands.
- Assumes workload data lives in the store and returns with it.
- Believes a node loss and a store loss are the same class of failure.