skip to content

Why must each pass of a reconciliation loop be written to set the live state, rather than to apply one more change?

level: middleimportance: should knowfreq 44%

answer

  1. the pass keeps no memory
  2. actions land slower than passes run
  3. count work that is still starting
  4. same end state on the second run
  5. set the target, never add one more

basics

~20 s

Passes repeat, overlap and re-run after crashes, and each one has no memory of the last. An action phrased as 'add one more' is therefore applied again on every pass and overshoots; an action that sets the end state is harmless to repeat.

solid answer

~50 s

A reconciliation pass has no memory of the previous one, and it very often runs again before its last action has visibly landed. Say a ledger service declares six copies and four are running: a pass that computes the difference and issues *start two* is only safe if the next pass, seconds later, counts the two that are still starting. If it counts ready copies only, it reads four again, starts two more, and you finish with eight. The property you need is **idempotence** — running the pass twice leaves the same end state as running it once — and you get it by phrasing every action as *make it so* against a freshly read world, never as an increment. The same property is what makes a crash halfway through a pass safe: the next pass simply re-derives what is left to do.

code

pseudocode · 11 lines
pseudocode
every interval:
    declared = readDeclaredState(workload)        # replicas: 6
    observed = listCopies(workload)               # includes copies still starting
    diff = declared.replicas - observed.count

    if diff > 0:
        startCopies(workload, diff)               # pass 1: starts 2
    else if diff < 0:
        stopCopies(workload, -diff)
    # pass 2: observed.count is 6, diff is 0, neither branch fires
    # nothing here remembers what the previous pass did

go deeper

for a junior

Recall that the loop runs the same comparison again and again, so any action it takes has to be safe to take a second time.

for a middle

Explain the overshoot concretely: four running against six declared, two started, another pass before they are ready, and ten copies if the pass does not count work in flight.

for a senior

Show how you would build it: deterministic identities and ownership links so a repeated create is a no-op, in-flight work counted as existing, and retries that re-read rather than resume.

for a principal

Frame it as the contract every automated writer in the estate must meet, since several loops and pipelines will touch the same objects and none of them can assume it was the last writer.

## The pass is not the only pass The defining fact about a control loop is that its unit of work runs **over and over**, and every run starts from nothing. A pass does not know what the previous pass decided, whether the previous pass finished, or whether its own effects have landed yet. Three things follow immediately: 1. The same difference can be **read twice** before either reading's action takes effect. 2. A pass can **die halfway** through acting, leaving the world in a state nobody declared. 3. A pass can be **triggered spuriously**, by a duplicate notification or a periodic sweep, with nothing to do. An action that is safe under all three is called **idempotent**: applying it a second time produces the same end state as applying it once. It is not a nicety here; it is the only thing standing between the loop and a runaway. ## The overshoot, with the numbers Take a ledger service declaring **six** copies, with **four** currently running, and a pass interval of five seconds. Starting a copy takes roughly thirty seconds before it is ready. | Pass | What it counts as existing | Difference | Action | Copies after | |---|---|---|---|---| | 1 | 4 ready | 6 − 4 = 2 | start 2 | 4 running, 2 starting | | 2, ready-only | 4 ready | 6 − 4 = 2 | start 2 more | 4 running, 4 starting | | 3, ready-only | 4 ready | 6 − 4 = 2 | start 2 more | 4 running, 6 starting | | 2, counting starting copies | 6 existing | 6 − 6 = 0 | none | 4 running, 2 starting | The ready-only loop ends up with ten copies for a declaration of six, and would have kept going. The defect is **not** that the interval was too short — a longer interval only makes the bug rarer and harder to reproduce. The defect is in what the pass treats as *existing*: work already in flight is part of the live state, and a pass that ignores it double-counts the same difference. ## Two ways to make a pass repeatable - **Phrase the action as a target, not a delta.** `set count to declared` is repeatable by construction; `add one` is not. Where the platform's own operation is inherently incremental, the pass must compute the target and issue the smallest operation that reaches it from what it just read. - **Make individual creations naturally idempotent.** Give each created thing a deterministic identity derived from the declaring object, so that a repeated creation is a lookup that finds it already there rather than a second creation. Without this, the loop has to remember what it created, and remembered state is exactly what a restart destroys. Ownership links matter here too: if what the loop creates is tied back to the declaring object, a repeat can recognise its own earlier work instead of treating it as a stranger's. ## Why crashes stop being interesting In a system built this way, a pass that dies mid-action needs no rollback and no recovery routine. The world is left in some intermediate state; the next pass reads that state, computes the remaining difference, and continues. This is why control loops are usually written with no transaction spanning their actions at all — the retry *is* the recovery, and the cost of a partially applied pass is one more pass. The contrast with a one-shot run is sharp. A command that fails halfway leaves a human to decide whether re-running is safe, and that question only has a good answer if the command was idempotent to begin with. ## The boundaries of the property Two honest limits are worth stating, because interviewers probe them: - **Idempotence is about repetition, not concurrency.** Two passes running *at the same time* can both read the old state and both act, and no amount of idempotence in each one prevents that. Platforms keep at most one pass per object in flight and rely on conflict detection when writing, so a pass that lost a race simply runs again. - **Idempotence is about the end state, not about side effects.** Repeating a pass that sends a notification, bills someone, or appends to a ledger is only safe if those effects are themselves deduplicated. Keeping such effects out of the reconciliation path, or keying them, is part of designing a loop rather than an afterthought. Everything in this section is why the mental model for authoring loop logic is *a function from current state to actions*, evaluated fresh, rather than *a script that runs once*.

  • What if the action creates something that has no stable name?
    Give it one, derived deterministically from the declaring object, so that a repeat becomes a lookup that finds the thing already present. The alternative is for the loop to remember what it created, and that memory does not survive a restart — which is the failure the whole design is built to avoid.
  • Does idempotence mean two passes can safely run at once?
    No. Idempotence makes repeats safe in sequence; two concurrent passes can both read the stale state and both act on the same difference. Platforms therefore keep at most one pass per object in flight and detect conflicts on write, so the loser of a race simply runs again against fresher state.
  • How does this change how a pass handles a partial failure?
    It removes the need for rollback. A pass that dies halfway leaves an intermediate world; the next pass reads it, computes what is still missing, and continues. The retry is the recovery, which is why loop logic rarely wraps its actions in a transaction and why failing loudly mid-pass is acceptable.

saying these in an interview costs you the question

  • Thinks a longer interval between passes removes the need for idempotence
  • Treats copies that are still starting as non-existent, so they get started twice
  • Believes the loop remembers last pass's actions and resumes from there
  • Assumes a crash mid-pass must be rolled back before a retry is allowed
  • Phrases the action as an increment because the difference was computed correctly
  • Thinks idempotent actions also make concurrent passes safe