skip to content

Orchestration Model

What a cluster scheduler decides for you: where a workload lands, how many copies run, and what happens when one dies. Interviewers ask because running a container and running a fleet differ.

on this pageshow

questions

page 1 of 2

Why does a live workload's copy count, raised by hand during an incident, drop back to the declared number within a minute?

level: juniorimportance: must knowfreq 68%

answer

  1. the platform is not running your command
  2. something keeps looking, not just once
  3. declared is the input, live the output
  4. your edit landed on the output
  5. next pass diffs and removes the extra

basics

~20 s

A control loop runs continuously, reading the live state, comparing it with the declared spec and acting on the difference. A copy count raised by hand is exactly such a difference, so the next pass removes the extra copies.

solid answer

~40 s

The platform is not executing the command you typed; it is holding a declared state and closing the gap to it. A control loop observes the live state, diffs it against the declared spec, acts, and then repeats, passes seconds apart. Raising the copy count by hand changed the live side only — the declaration still says six — so the next pass sees eight running against six declared and stops two. The edit was never rejected; it was accepted by the live system and then undone. To make a change stick you edit the declaration, which is the loop's input, not the running copies, which are its output. The attempt itself usually survives only in the platform's change and event records, which age out.

go deeper

for a junior

Recall the three steps — read the live state, compare it with the declared state, act on the difference — and that the loop repeats forever, which is why a change made by hand does not last.

for a middle

Explain which side the edit landed on: the declaration is the loop's input and the running copies are its output, so changing the output is undone at the next pass rather than refused.

for a senior

Show how you would actually add capacity during an incident — edit the declaration, or suspend reconciliation for that one workload where the platform offers it — and say where evidence of the attempt will still be an hour later.

for a principal

Argue the operating rule this implies: if the declaration is the only durable truth, every emergency lever must write there, or the estate accumulates changes nobody can review or revert.

## What you handed the platform When you run a container yourself, you issue an instruction and the runtime carries it out **once**. A cluster platform works the other way round. You submit a **declared state** — a document saying what should be true: this ledger service runs six interchangeable copies of this image, each with this reservation, matched by this selector — and the platform takes on standing responsibility for making the world match it. Your submission is not an execution. It is a write to a record of intent. A **control loop** then does the work, and it does it forever: 1. **Observe** — read the live state: how many copies exist right now, on which hosts, in what condition. 2. **Diff** — compare that reading against the declared state and compute the difference. 3. **Act** — take the action that closes the difference, then discard everything it knew and start again. Passes run seconds apart, and platforms typically re-derive everything again on a slower full sweep. The loop is not waiting for anyone to tell it that something happened. ## Why the hand-made change did not survive During the incident the copy count was raised on the **live** side — the running objects — while the **declared** side still said six. The loop neither knows nor cares that a human produced the difference. It reads eight where six is declared, and stops two. Nothing rejected the change; the live system accepted it and the next pass undid it, which is a more confusing experience than an outright refusal. | What you edited | What the next pass computes | Outcome | |---|---|---| | The live copy count, raised to eight | live 8 against declared 6 | two copies stopped; the declaration wins | | The declared copy count, raised to eight | declared 8 against live 6 | two copies started; the change sticks | | Nothing, but a host went away | live 5 against declared 6 | one copy replaced elsewhere | The asymmetry is the whole lesson: **the declaration is the loop's input and the running copies are its output.** Editing an output is temporary by construction. ## Where the attempt survives After the revert, the declared document is byte-for-byte what it was before the incident — it has no memory of the excursion. What does survive is thinner and shorter-lived: - the platform's **event record** for that workload, usually saying that copies were created and then removed, which on most platforms ages out after hours; - the **audit record** of the write itself, where one is kept, naming who made it; - whatever your own change management captured outside the platform. So an incident timeline built only from the declared state will not show that anyone tried anything. This is a real argument for making emergency capacity changes as a declaration edit: the edit is durable, reviewable and revertible, and the loop will honour it until someone changes it back. ## Operating inside a loop - **Change the declaration, not the running copies.** Every durable change goes to the input. - **Expect a window.** The revert takes the pass interval plus the time the action needs. It is typically seconds and occasionally longer, so you may genuinely see the extra copies serve traffic before they are stopped — it is not proof the change was accepted. - **Anything else that writes to the live side is in the same fight.** A runbook script or a second automation that adjusts running copies will be reverted just as a human is, and the flapping that results is hard to diagnose from the declaration alone. - **Where the platform offers it, suspending reconciliation for one object is the honest emergency lever.** Designs differ: some let you pause the loop for a single workload, some do not. While it is paused, hand edits persist and nothing converges — which is exactly why it has to be undone deliberately. ## What the loop is not - It is **not a reviewer**. It has no opinion about whether the declared state is a good idea; it converges on a bad declaration as faithfully as on a good one. Review belongs on the declaration. - It is **not triggered by your edit**. It would have found the difference on its own pass anyway, which is why a change made while nobody was watching still gets reverted. - It does **not** promote the live state into the declaration. The live state is evidence, never intent — which is precisely why a hand-made change cannot make itself permanent.

  • How quickly is a hand-made change reverted, and is that guaranteed?
    It takes one pass interval plus the time the corrective action needs, so usually seconds and sometimes a minute or two. There is no guarantee of instancy: passes are scheduled, and the platform may be busy. Seeing the extra copies serve traffic for a while is therefore not evidence that the change was accepted.
  • What does the loop do if the declaration itself is wrong?
    It converges on the wrong thing, faithfully and repeatedly. The loop compares live state against declared state; it has no notion of whether the declared state is desirable. That is why review, approval and testing attach to the declaration rather than to the running copies — by the time something is running, the decision has already been made.
  • Can you stop the loop from reverting an emergency change?
    On platforms that offer it, you can suspend reconciliation for a single object; while suspended, hand edits persist and no convergence happens for that workload, including corrections you would have wanted. Designs differ on whether this exists at all. Treat it as a deliberate, time-boxed override that someone must undo, not as a normal workflow.

saying these in an interview costs you the question

  • Claims the platform rejected the manual change, rather than accepting and undoing it
  • Thinks editing a running copy is how you make a change permanent
  • Believes the loop only reports differences and waits for a human to act
  • Assumes the loop fires on the edit, so an unobserved change would survive
  • Says the declared state is rewritten to match whatever is running
open as a page

In a container orchestration cluster, what does the control plane own and what does each host's node agent own?

level: juniorimportance: must knowfreq 62%

basics

~20 s

The control plane decides: an API, the state store behind it, and the scheduler and controllers that act on declared state. Each host's node agent runs and supervises the containers assigned to that host, alongside the workloads themselves serving traffic.

open as a page

When a scheduler places several containers as one co-located group, what do those members share?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Members of a co-located group share one host, one network address and port space, and a scratch area every member can mount. The group is also what the platform places, replaces and deletes — never one member on its own.

open as a page

At 3am one replica of a five-replica device-telemetry ingester exits and nobody is paged - what replaces it, and what does the replacement not inherit?

level: juniorimportance: must knowfreq 75%

basics

~20 s

The platform's control loop sees four copies where the spec declares five and starts a fifth. The replacement is a new instance built from the same spec: new address, empty writable layer, cold caches. Nothing from the dead copy is recovered.

open as a page

Your platform runs an always-on checkout service and a nightly report that must finish once — why are those different workload contracts?

level: juniorimportance: must knowfreq 72%

basics

~20 s

They disagree about what a process exiting means. A replicated long-running service must never exit, so the platform recreates any copy that does. A run-to-completion job is supposed to exit once; a clean exit is the whole point and ends it.

open as a page

What does a workload spec for an order service declare, and what does it deliberately leave out?

level: juniorimportance: must knowfreq 74%

basics

~20 s

A workload spec declares the end state: which image, how many copies, what each copy reserves and may not exceed, what storage it needs, and how it is reached. It leaves out the steps, the hosts and the current situation.

open as a page

A cluster's control loop stops for an hour, then resumes: why do workloads converge instead of staying stuck on missed changes?

level: middleimportance: must knowfreq 52%

basics

~20 s

Each pass re-derives everything from the current declared state and the current live state, so no notification needs to have been seen. A loop resuming after an hour reads the world as it is now and closes whatever difference it finds.

open as a page

When a cluster scheduler places one new replica, what are the two stages it runs over the candidate hosts?

level: middleimportance: must knowfreq 62%

basics

~20 s

A cluster scheduler first filters: it drops every host that cannot satisfy the workload's reservation, required attributes or hard placement rules. It then scores the survivors against weighted preferences and places the workload on the highest-scoring host.

open as a page

A scoring service with a large reservation stays pending while the cluster shows plenty of free capacity — why?

level: middleimportance: must knowfreq 70%

basics

~10 s

Placement is a per-host fit, never a cluster-wide one. That free capacity is a sum of small leftovers spread over many hosts, and the whole reservation has to fit inside one host's unreserved room.

open as a page

Two containers in one co-located group both listen on port 8080 — what happens when the group starts?

level: middleimportance: must knowfreq 50%

basics

~20 s

One of them fails. The group has a single address and therefore a single port space, so whichever member binds 8080 first keeps it and the other fails with the port already in use, then restarts into the same failure.

open as a page

A host running three device-telemetry ingester replicas stops reporting to the control plane - why does the platform wait before replacing them?

level: middleimportance: must knowfreq 60%

basics

~20 s

Silence is ambiguous: a host that stops reporting may be dead, or its agent may have crashed, or only the network between it and the control plane is broken. Platforms wait out an unreachable-node timeout before declaring its workloads gone and replacing them.

open as a page

Your team gets a named scope on a cluster six other teams share — what does that scope separate, and what stays shared?

level: middleimportance: must knowfreq 62%

basics

~20 s

A named scope separates object names, listings, and the access, budgets and defaults attached to it. It does not separate the hosts workloads land on, the kernel and capacity those hosts share, cluster-wide objects, or network reachability between scopes.

open as a page

A long-running queue consumer was declared under the run-to-completion contract and exits cleanly when the queue empties — what does the platform do?

level: middleimportance: must knowfreq 58%

basics

~20 s

Nothing. A clean exit is what the run-to-completion contract was told to expect, so the platform records a successful run and stops. The consumer is gone, no copy is missing, nothing failed, and no alert fires.

open as a page

Why does a stored workload object keep the declared spec and the live status as two separate blocks?

level: middleimportance: must knowfreq 60%

basics

~20 s

Because they have different authors and different jobs: the declared block is intent, written by the team; the status block is observation, written back by the platform. Keeping them apart is what makes the difference between them readable.

open as a page

A chat gateway's cluster API answers nothing for an hour while user sessions stay up — which operations stop and which keep working?

level: seniorimportance: must knowfreq 68%

basics

~20 s

Traffic keeps flowing and running copies keep serving: the node agents and containers need no new decisions. What stops is every change — applying a spec, scaling, placing new work, replacing a copy lost with its host — and reading status, which goes through the same API.

open as a page

Why must each pass of a reconciliation loop be written to set the live state, rather than to apply one more change?

level: middleimportance: should knowfreq 44%

basics

~20 s

Passes repeat, overlap and re-run after crashes, and each one has no memory of the last. An action phrased as 'add one more' is therefore applied again on every pass and overshoots; an action that sets the end state is harmless to repeat.

open as a page

When the cluster API stops answering, why does a node agent keep restarting a crashed container on its own host?

level: middleimportance: should knowfreq 50%

basics

~20 s

A node agent already holds its host's assignments and supervises them locally, so restarting a container it was told to run needs no new decision from the deciding half. Moving that work to a different host does.

open as a page

How does a setup container in a co-located group hand a data snapshot to the indexer that starts after it?

level: middleimportance: should knowfreq 48%

basics

~20 s

It is declared as a member that must run to completion first: it writes the snapshot into the group's shared scratch area and exits successfully, and only then does the platform start the long-running members, which read the same area.

open as a page

A device-telemetry ingester's spec carries a bad setting, so every replacement comes up faulty - what does self-healing actually restore?

level: middleimportance: should knowfreq 50%

basics

~20 s

Only the count, measured against the spec as written. The loop starts copies until the declared number exists; it never edits the spec, so a fault that lives in the spec is reproduced identically in every replacement, and a workload can sit at full count serving nothing.

open as a page

Your scope's aggregate reservation budget is full: a new copy is refused while the four running keep serving — why?

level: middleimportance: should knowfreq 50%

basics

~20 s

A scope budget is evaluated when an object is created or changed, against the sum of the reservations already declared in that scope. A request that would push the sum past the ceiling is refused, and workloads already admitted are never revisited, shrunk or killed.

open as a page

Why declare a per-host metrics agent as a one-copy-per-host workload rather than a replicated service with the copy count set to the number of hosts?

level: middleimportance: should knowfreq 50%

basics

~20 s

Because the requirement is coverage, not quantity. A one-copy-per-host workload derives its count from the set of matching hosts and guarantees exactly one copy on each. A copy count is just a number, and nothing stops it landing two copies on one host and none on another.

open as a page

Why does submitting the same workload spec twice converge on one set of copies instead of adding more?

level: middleimportance: should knowfreq 52%

basics

~10 s

Because the document is an absolute statement about a named workload, not an instruction. Four copies means there shall be four, so the second submission finds nothing to change and asks for nothing new.

open as a page

Why does declaring six copies when only four can be placed produce no error at all, just two copies left waiting?

level: seniorimportance: should knowfreq 37%

basics

~20 s

Accepting a declaration and satisfying it are separate steps: acceptance is a synchronous validation, satisfaction is an open-ended loop with no terminal failure. The two unplaceable copies stay a standing difference that the loop keeps retrying.

open as a page

Why does losing the state store behind a cluster's API hurt differently from losing a node that runs workloads?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A node is interchangeable: its work is placed elsewhere. The state store is the cluster's only record of what should exist and what was last reported, so losing it leaves the deciding half with nothing to act on, even while the containers keep running.

open as a page

A host sits at 20% measured CPU use, yet the scheduler will not place anything more on it — why?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Schedulers place against reservations, not measurements. Everything already on that host has reserved capacity it is not currently using, so the host is fully booked even though it looks idle, and no further reservation fits.

open as a page

Under a strict even-spread rule across three failure domains, why does a workload's fourth replica stay pending?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A strict spread rule is a filter, not a preference. With three failure domains holding one replica each and no unevenness permitted, placing a fourth anywhere would make the counts uneven, so every host is ruled out and the replica waits.

open as a page

A helper container in a running co-located group needs a new image — what does the platform actually replace, and at what cost?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The whole group. Changing any member's declared image produces a new group instance, so every member restarts, the group becomes reachable at a new address and its scratch area starts empty — the main process pays for the helper's release.

open as a page

An unreachable host's replicas were replaced; the host returns with its old copies still running and writing - what did the platform actually guarantee?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A floor, not a ceiling. The platform promises that at least the declared number of copies exists; it cannot promise at most that number, because declaring a silent host's workloads gone is a guess made without contact. Bounding writers must come from outside the scheduler.

open as a page

Your workloads declare no reservation or ceiling, so the scope's default is stamped onto each — what goes wrong later?

level: seniorimportance: should knowfreq 41%

basics

~20 s

A scope default is a placeholder chosen for the whole scope, not a measurement of your workload. Too low, and it is throttled or killed against a ceiling nobody sized for it; too high, and unused headroom drains the scope's budget while the hosts sit idle.

open as a page

Last night's reconciliation run is still going when the recurring time trigger fires again tonight — what are the platform's options, and how do you choose?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Three options: skip tonight's run, start it anyway alongside the old one, or cancel the old one and start fresh. Which is right depends on whether two runs can safely touch the same data and whether a skipped period is ever processed later.

open as a page

showing 1–30 of 38