When the cluster API stops answering, why does a node agent keep restarting a crashed container on its own host?
answer
- the assignment is already local
- enforcing, not deciding
- restart here needs no new decision
- new placement needs the deciding half
- status reports queue and age
basics
~20 sA node agent already holds its host's assignments and supervises them locally, so restarting a container it was told to run needs no new decision from the deciding half. Moving that work to a different host does.
solid answer
~40 sThe assignment arrived before the outage did. A node agent keeps a local record of what it was told to run on its host, and its supervision loop compares that record against what is actually running — entirely on the host, with no call outward. Restarting an exited container is therefore *enforcement* of a decision already made, not a new decision. What the agent cannot do alone is anything that changes the record: it cannot accept newly declared work, cannot raise a copy count, and cannot place a copy on another host. Meanwhile its status reports have nowhere to go, so the cluster's recorded view of that host ages while the host itself is fine.
code
pseudocode · 9 lines# node agent, one pass, on a single host
for assignment in agent.localRecord: # delivered by the cluster API before contact was lost
if not running(assignment) and assignment.restartOnExit:
runtime.start(assignment) # decided here; nothing is asked of the deciding half
if api.reachable():
api.report(agent.observedStatus) # status flows up, new assignments flow down
else:
agent.holdStatus() # nothing new arrives; localRecord does not changego deeper
Remember that the instructions for a host arrive before they are needed, so a host can keep doing its job while the cluster API is unreachable.
Explain the mechanism: the agent's supervision loop compares its local record against what is running and acts on the difference, which is enforcement rather than a fresh decision.
Show you know what degrades anyway — status ages, nothing new lands, credentials can expire — and that replacement on another host is a decision the host cannot make for itself.
The design question underneath is how much authority to push to the edge: a more autonomous agent survives longer alone, at the cost of acting on a record that may be badly out of date.
## The assignment arrives before the outage does The deciding half's output is an assignment: *this host should be running these containers, built from this image, with these settings*. Once that assignment has been delivered, the node agent holds it locally. From then on the agent's job is narrow and local — compare what it was told to run against what is actually running on this host, and act on the difference by asking the container runtime to start or stop something. Nothing in that loop requires the cluster API. That is the design, not an accident: if every container restart needed a round trip to a central service, the central service would be on the critical path of every process crash on every host in the fleet. ## What the agent settles by itself - **Restarting a container that exited**, where the assignment says it should be running. - **Noticing a container has stopped, or is failing its health check**, and acting on it within the scope of the assignment it holds. - **Stopping a container** that it was already told to stop before contact was lost. - **Reporting nothing upward**, and continuing anyway — a status report it cannot deliver does not change what it runs. ## What it cannot settle - **Accepting newly declared work.** A spec applied during the outage never reaches the API, let alone the agent. - **Changing a copy count.** The declared number lives in the state store; the agent holds only its own slice. - **Placing work on another host.** Placement is a decision, and agents do not make decisions for hosts other than their own — or for their own, for that matter. - **Replacing copies lost with a failed host.** Those copies had one agent, and it is gone with the host. ## Restart here against replacement elsewhere This is the line the question is really testing, and the two cases are easy to blur because both end with a container running again. | Event | Who acts | Needs the deciding half? | Outcome during an API outage | |---|---|---|---| | Container exits on a healthy host | That host's node agent, from its local record | No | It comes back on the same host | | Container is killed for exceeding its memory ceiling | That host's node agent | No | It comes back on the same host | | The whole host goes silent | Nobody local — the agent went with it | Yes | The copies stay gone until the API returns | | Declared copy count is raised | Control plane, then a chosen agent | Yes | The change is never recorded at all | ## What quietly degrades while the API is away The workloads are fine; the *management* of them is not. 1. **Status ages.** The agent's observations pile up locally or are dropped. The cluster's record of this host is whatever was last written, which means status read during an outage is evidence about the past. 2. **Nothing new arrives.** A configuration change, a new image reference or a scaling decision made during the hour simply does not exist from the host's point of view. 3. **Credentials and leases can expire.** Agents authenticate to the API, and some of what they hold is time-bounded. Platforms differ in how long an agent can stay usefully disconnected before reconnecting needs operator help, so a multi-hour outage is not just a longer version of a two-minute one. ## Where platforms differ Two behaviours genuinely vary, and a good answer says so rather than guessing: - **Whether the assignment survives the host rebooting.** Some designs persist the agent's record on local disk so the host comes back running the same workloads without asking anyone; others treat an empty host as empty until the deciding half tells it otherwise. If the interviewer asks, say that it depends and name both behaviours. - **How aggressively an agent gives up.** Some agents keep running disconnected indefinitely; others are built to stop trusting an assignment that has not been re-confirmed for a long time. ## The sentence to land *The agent is an enforcer, not a decider.* Everything it can do alone is enforcement of a decision that was made before the outage; everything it cannot do alone is a new decision. Answer in those terms and the follow-ups about failed hosts, stalled rollouts and stale status all fall out of the same rule.
- If the host itself reboots during the outage, do its workloads come back?It depends on the platform. Where the agent persists its assignments on local disk, the host boots and starts them again without asking anyone. Where it does not, the host comes up empty and stays empty until the deciding half can tell it what to run. Say both, rather than asserting one.
- What happens to a status change the agent could not report?It is held locally or dropped, and the cluster's recorded status stays at whatever was last written. Operators should read status from an outage window as evidence of the past, not the present — a workload shown as healthy may have crashed twice since, and one shown as failing may be serving fine.
saying these in an interview costs you the question
- Says every container restart is ordered by the control plane.
- Thinks an unreachable API stops containers that are already running.
- Assumes a node agent can place work on another host by itself.
- Believes a stale status view means the workload has stopped.
- Expects work declared during the outage to reach a host anyway.