Why does a live workload's copy count, raised by hand during an incident, drop back to the declared number within a minute?
answer
- the platform is not running your command
- something keeps looking, not just once
- declared is the input, live the output
- your edit landed on the output
- next pass diffs and removes the extra
basics
~20 sA control loop runs continuously, reading the live state, comparing it with the declared spec and acting on the difference. A copy count raised by hand is exactly such a difference, so the next pass removes the extra copies.
solid answer
~40 sThe platform is not executing the command you typed; it is holding a declared state and closing the gap to it. A control loop observes the live state, diffs it against the declared spec, acts, and then repeats, passes seconds apart. Raising the copy count by hand changed the live side only — the declaration still says six — so the next pass sees eight running against six declared and stops two. The edit was never rejected; it was accepted by the live system and then undone. To make a change stick you edit the declaration, which is the loop's input, not the running copies, which are its output. The attempt itself usually survives only in the platform's change and event records, which age out.
go deeper
Recall the three steps — read the live state, compare it with the declared state, act on the difference — and that the loop repeats forever, which is why a change made by hand does not last.
Explain which side the edit landed on: the declaration is the loop's input and the running copies are its output, so changing the output is undone at the next pass rather than refused.
Show how you would actually add capacity during an incident — edit the declaration, or suspend reconciliation for that one workload where the platform offers it — and say where evidence of the attempt will still be an hour later.
Argue the operating rule this implies: if the declaration is the only durable truth, every emergency lever must write there, or the estate accumulates changes nobody can review or revert.
## What you handed the platform When you run a container yourself, you issue an instruction and the runtime carries it out **once**. A cluster platform works the other way round. You submit a **declared state** — a document saying what should be true: this ledger service runs six interchangeable copies of this image, each with this reservation, matched by this selector — and the platform takes on standing responsibility for making the world match it. Your submission is not an execution. It is a write to a record of intent. A **control loop** then does the work, and it does it forever: 1. **Observe** — read the live state: how many copies exist right now, on which hosts, in what condition. 2. **Diff** — compare that reading against the declared state and compute the difference. 3. **Act** — take the action that closes the difference, then discard everything it knew and start again. Passes run seconds apart, and platforms typically re-derive everything again on a slower full sweep. The loop is not waiting for anyone to tell it that something happened. ## Why the hand-made change did not survive During the incident the copy count was raised on the **live** side — the running objects — while the **declared** side still said six. The loop neither knows nor cares that a human produced the difference. It reads eight where six is declared, and stops two. Nothing rejected the change; the live system accepted it and the next pass undid it, which is a more confusing experience than an outright refusal. | What you edited | What the next pass computes | Outcome | |---|---|---| | The live copy count, raised to eight | live 8 against declared 6 | two copies stopped; the declaration wins | | The declared copy count, raised to eight | declared 8 against live 6 | two copies started; the change sticks | | Nothing, but a host went away | live 5 against declared 6 | one copy replaced elsewhere | The asymmetry is the whole lesson: **the declaration is the loop's input and the running copies are its output.** Editing an output is temporary by construction. ## Where the attempt survives After the revert, the declared document is byte-for-byte what it was before the incident — it has no memory of the excursion. What does survive is thinner and shorter-lived: - the platform's **event record** for that workload, usually saying that copies were created and then removed, which on most platforms ages out after hours; - the **audit record** of the write itself, where one is kept, naming who made it; - whatever your own change management captured outside the platform. So an incident timeline built only from the declared state will not show that anyone tried anything. This is a real argument for making emergency capacity changes as a declaration edit: the edit is durable, reviewable and revertible, and the loop will honour it until someone changes it back. ## Operating inside a loop - **Change the declaration, not the running copies.** Every durable change goes to the input. - **Expect a window.** The revert takes the pass interval plus the time the action needs. It is typically seconds and occasionally longer, so you may genuinely see the extra copies serve traffic before they are stopped — it is not proof the change was accepted. - **Anything else that writes to the live side is in the same fight.** A runbook script or a second automation that adjusts running copies will be reverted just as a human is, and the flapping that results is hard to diagnose from the declaration alone. - **Where the platform offers it, suspending reconciliation for one object is the honest emergency lever.** Designs differ: some let you pause the loop for a single workload, some do not. While it is paused, hand edits persist and nothing converges — which is exactly why it has to be undone deliberately. ## What the loop is not - It is **not a reviewer**. It has no opinion about whether the declared state is a good idea; it converges on a bad declaration as faithfully as on a good one. Review belongs on the declaration. - It is **not triggered by your edit**. It would have found the difference on its own pass anyway, which is why a change made while nobody was watching still gets reverted. - It does **not** promote the live state into the declaration. The live state is evidence, never intent — which is precisely why a hand-made change cannot make itself permanent.
- How quickly is a hand-made change reverted, and is that guaranteed?It takes one pass interval plus the time the corrective action needs, so usually seconds and sometimes a minute or two. There is no guarantee of instancy: passes are scheduled, and the platform may be busy. Seeing the extra copies serve traffic for a while is therefore not evidence that the change was accepted.
- What does the loop do if the declaration itself is wrong?It converges on the wrong thing, faithfully and repeatedly. The loop compares live state against declared state; it has no notion of whether the declared state is desirable. That is why review, approval and testing attach to the declaration rather than to the running copies — by the time something is running, the decision has already been made.
- Can you stop the loop from reverting an emergency change?On platforms that offer it, you can suspend reconciliation for a single object; while suspended, hand edits persist and no convergence happens for that workload, including corrections you would have wanted. Designs differ on whether this exists at all. Treat it as a deliberate, time-boxed override that someone must undo, not as a normal workflow.
saying these in an interview costs you the question
- Claims the platform rejected the manual change, rather than accepting and undoing it
- Thinks editing a running copy is how you make a change permanent
- Believes the loop only reports differences and waits for a human to act
- Assumes the loop fires on the edit, so an unobserved change would survive
- Says the declared state is rewritten to match whatever is running