skip to content

Why does declaring six copies when only four can be placed produce no error at all, just two copies left waiting?

level: seniorimportance: should knowfreq 37%

answer

  1. submitting is a write, not a run
  2. success means recorded, not running
  3. the loop has no give-up state
  4. the world may become satisfiable later
  5. watch declared against ready, and its age

basics

~20 s

Accepting a declaration and satisfying it are separate steps: acceptance is a synchronous validation, satisfaction is an open-ended loop with no terminal failure. The two unplaceable copies stay a standing difference that the loop keeps retrying.

solid answer

~40 s

Submitting a declaration is a write, not an execution. It is validated and recorded, and success means *recorded*, not *running*. A control loop then works on the difference, and it has no concept of giving up, because the world can change — capacity frees up, a host comes back — so it retries, usually with backoff, for as long as the difference exists. Six declared against four placed is therefore not an error anywhere: it is two copies in a waiting state plus an event trail saying each attempt could not be satisfied. That is a feature when the shortage is temporary and a trap otherwise, because nothing alerts unless you are watching the declared count against the count reported ready.

go deeper

for a junior

Recall that a declaration being accepted only means it was recorded as valid; whether the copies are actually running is a separate thing you have to look at.

for a middle

Explain the asynchrony: validation happens at submit time, satisfaction happens in a loop afterwards, and an unmet declared state is a standing difference rather than an error.

for a senior

Show the monitoring you would add — the gap between declared and ready copies, and how long it has stood — and note that the event trail explaining why expires, so the reason must be captured when first seen.

for a principal

Argue where responsibility lands: if declarations never fail, the platform owes every team a visible unmet-declaration signal, or teams will keep believing it is doing work it has accepted and cannot do.

## Two steps that people collapse into one The surprise here comes from expecting a command. A command either works or returns an error, and the error arrives while you are still looking at it. Declaring a state is not that. It is two separate steps with two different failure models. | Step | When it happens | What can fail | How you find out | |---|---|---|---| | Acceptance | synchronously, at submit time | malformed document, missing required field, a policy check refusing it, insufficient permission | immediately, as an error on the write | | Satisfaction | continuously, afterwards | anything about the world: no capacity, an image that cannot be fetched, a dependency absent | only by reading the object's state and events | So `accepted` is a statement about the **document**, not about the world. A declaration that can never be satisfied is a perfectly valid document, and rejecting it at submit time would require the platform to decide, at that instant, that no future world could satisfy it. It cannot know that, so it does not try. ## Why the loop has no failure state for this A control loop's job is to close the difference between declared and live state. Consider what "giving up" would mean: the loop would decide that a difference is permanent and stop looking. But the property that makes the difference unsatisfiable is a property of *the world right now*: - capacity is freed when another workload finishes or shrinks; - a host that was unavailable comes back or is added; - a constraint the workload depends on is fixed by someone else. Any of these can arrive a minute later, and when it does, the standing difference is satisfied with nobody resubmitting anything. That is the payoff of the design, and it is the same reason an outage in the deciding half strands nothing: the intent is durable, so the work resumes on its own. So the loop retries, typically with **backoff** so that a hopeless case does not burn resources at full speed. Backoff caps the *rate*, not the *number* of attempts. ## What the difference actually looks like Four copies running and two waiting is not a hidden state. It shows up in three places, in increasing order of detail: 1. **The counts.** The declared number and the number reported ready disagree, persistently. This is the signal to alert on. 2. **The waiting copies themselves.** Each unplaced copy exists as an object in a pending condition, so you can see there are two of them rather than inferring it. 3. **The event trail.** Each attempt records why it could not be satisfied. Events are time-bounded on most platforms, so a difference that has stood for days may have no events left explaining its origin — a common and frustrating dead end during an incident. ## The operational trap, and how to close it The failure mode is not the unplaced copies. It is **a team believing they are running at six**. A submission returned success, a rollout looked like it completed, dashboards show a healthy service — at two-thirds of its intended capacity, with no error anywhere in the chain. It survives until load arrives. What closes it: - **Alert on the gap, not on errors.** Declared count minus ready count, sustained beyond a few minutes, is the honest signal. Errors will never fire. - **Alert on the gap's age, not just its existence.** A difference of two for thirty seconds during a normal replacement is routine; the same difference for an hour is an incident. Without age, the check is either noisy or useless. - **Treat submission as a request, not a confirmation.** Any automation that submits a declaration and then reports success has to wait for the ready count to reach the declared one, or it is reporting on the document rather than on the service. - **Capture the reason early.** Because the event trail expires, an unmet difference should have its reason recorded by whatever watches it, at the moment it is first seen. ## Where this is genuinely the right behaviour It is worth defending the design rather than just warning about it. Under a one-shot model, a temporary shortage means a failed run, a human noticing, and a manual resubmission at some later moment when they guess capacity exists. Under a standing difference, the platform holds the intent and satisfies it the instant it becomes satisfiable, at three in the morning, without anyone awake. The cost of that is precisely that *nothing is ever an error*, which is why the monitoring has to be built on state rather than on outcomes.

  • How long does the loop keep retrying, and does anything slow it down?
    It retries for as long as the difference exists, because it has no terminal failure for an unmet declaration. What it does have is backoff: successive attempts are spaced further apart, so a hopeless case settles into an occasional cheap retry rather than a hot loop. Backoff limits the rate of attempts, never their total number.
  • What happens when the capacity finally appears?
    Nothing is resubmitted. The next pass reads the same declared state, finds the same difference, and this time the two copies can be placed. That is the payoff of a standing difference over a one-shot run: the intent outlives the condition that blocked it, so recovery needs no human at the moment it becomes possible.
  • Is an indefinitely retried declaration ever a problem in itself?
    Yes. A declaration nobody will ever satisfy is indistinguishable from one that will be satisfied in ten minutes, and it holds waiting objects, event volume and operator attention indefinitely. That is why the useful signal is the age of the unmet difference rather than its existence, and why someone has to own withdrawing declarations that are simply wrong.

saying these in an interview costs you the question

  • Expects the submission itself to fail because the copies cannot be placed
  • Says a declaration the loop cannot satisfy is eventually marked failed
  • Treats a successful submission as proof that the copies are running
  • Assumes someone is alerted without a declared-against-ready check existing
  • Thinks a fixed retry budget exists and is exhausted after some attempts
  • Expects the event trail to still explain a difference that has stood for days