skip to content

In infrastructure-as-code, what does it mean for a tool to be convergent, and why is it normally safe to simply re-run the same configuration after a run failed halfway through?

level: juniorimportance: must knowfreq 66%

answer

  1. same end state, repeatable runs
  2. second run reports no changes
  3. closes the remaining gap only
  4. partial failure is safe to repeat
  5. unconditional commands break it

basics

~20 s

Convergence means each run moves the system toward the declared state and stops when it matches, so a second run with no changes does nothing. A half-finished run is safe to repeat: the tool redoes only the work still outstanding.

solid answer

~50 s

Convergent means every run drives the system toward the declared end state and does only the work still missing, so repeated runs settle on a fixed point instead of stacking up effects. The practical test is that a second run immediately after a successful one reports no changes. That is what makes a failed run recoverable by re-running rather than by hand-unwinding: the tool re-derives the gap between what is declared and what now exists, sees the resources the failed run already created, and finishes the rest. Two caveats worth raising. Convergence does not promise a *single* run reaches the end state — a dependency that is still initialising may need a second pass. And it only holds for the declarative parts: an escape-hatch step that unconditionally runs a command has whatever behaviour that command has, and is the usual reason a repeat run does damage.

go deeper

for a junior

Be ready to state the two-run test in one sentence: run it, run it again, the second run should report no changes. Say clearly that re-running after a failure finishes the missing work rather than duplicating what exists.

for a middle

Explain how the property is achieved — the tool derives actions from the current gap rather than replaying a script — and name the things that break it, especially unconditional commands and attributes the provider rewrites.

for a senior

Bring a real recovery story: a run cancelled or rate-limited mid-flight, and the judgment of when re-running is safe versus when a half-created object must be replaced instead. Also explain why a permanent no-op diff is an operational hazard.

for a principal

Frame convergence as what makes automated, unattended change viable at all, and set the standard: a plan that is not empty on a no-op run is treated as a defect, because reviewers who see noise every day stop reading the ones that matter.

## The property A convergent system has a fixed point. Point it at a declared end state and each run reduces the distance to that state; once the distance is zero, further runs do nothing. That last part is the observable test, and it is the one interviewers ask you to state: run the configuration, then run it again with nothing changed in between, and the second run must report no actions. If it proposes work every time, something is not convergent, and you have a bug — either in a resource whose observed form never matches what was declared, or in an escape hatch that runs unconditionally. ## Why it makes failure cheap The operational payoff is recovery. Infrastructure runs fail in the middle for boring reasons: an API rate limit, an expired credential, a quota, a network blip, someone cancelling the job. When that happens you are left in a partial state — some resources created, some not, possibly one half-configured. With a convergent tool the recovery procedure is: fix the cause, run it again. The run recomputes the gap from scratch. Resources the failed attempt created are now part of reality (and, in a stateful tool, part of its record), so they produce no action. Resources it never got to are still missing, so they are created. Nothing is duplicated, and no human has to reconstruct which of forty resources the crashed job had reached. The alternative is a script that executes a fixed sequence of steps. Re-running step one after it already succeeded is at best wasted work and at worst a second copy of something. Every such script eventually grows a wrapper of existence checks — which is the same idea as convergence, hand-rolled badly for one script. ## Convergence in one pass versus many A subtlety worth volunteering: convergent does not mean *one* run always suffices. Classic convergent configuration systems were explicitly designed around repeated passes — a run brings you closer, and a following run closes what the first could not. Modern infrastructure tools aim to finish in a single apply and will report an error rather than silently leave the gap open, but you still meet cases where a second run is the honest answer: a resource that reports ready before its dependent API accepts writes, an eventually-consistent lookup that has not caught up, a certificate still validating. The distinguishing behaviour is that the second run is *safe*, not that it is unnecessary. ## Where convergence breaks Three failure modes come up repeatedly. **Unconditional escape hatches.** Most tools let you run an arbitrary command as part of a change. That command is only as convergent as you make it. Appending a line to a file appends a line every run; adding a user with a create command fails the second time. Fixing it means either declaring the outcome instead of the command, or guarding the command with a condition so it is skipped when the outcome already holds. **Non-comparable attributes.** If the provider normalises, defaults or reformats a value — reordering a list, expanding a shorthand, filling in a field you left blank — the observed value never equals the declared one and the tool proposes a change on every single run. This shows up as a permanent, harmless-looking diff, and it is corrosive: people stop reading plans that always contain noise. **A genuinely failed partial action.** Some provider operations are not atomic. If a create half-succeeded and left an object in a broken state that the tool now considers present and correct, re-running converges to the wrong fixed point. Recognising this — and deliberately replacing the object rather than re-running — is the senior move. ## Saying it in an interview Define it, give the two-run test, then give the failure story: "we cancelled an apply mid-flight when the pipeline timed out; we did not unwind anything, we fixed the timeout and re-ran, and it created only the six resources that were still missing." That answer shows you have used the property rather than memorised it.

  • What does it mean if a run proposes the same change every single time, forever?
    It means the declared value and the observed value never compare equal, so the tool never reaches its fixed point. Usual causes are provider-side normalisation or defaulting, a list whose order is not stable, or a field the API rewrites. Either declare the value in the exact form the provider returns, or tell the tool to stop comparing that attribute. Leaving a permanent diff in place trains reviewers to ignore plans.
  • Is a convergent run always safe to re-run when the previous one failed?
    Almost always, but not blindly. The exception is a non-atomic operation that left an object present yet broken: the tool now sees it as existing and matching, so re-running converges to a wrong fixed point and looks green. When a create half-succeeded, replace the object rather than repeating the run, and check the provider's view before assuming the record is honest.

saying these in an interview costs you the question

  • Running it a second time will create duplicate resources
  • Convergence means the tool retries until the run succeeds
  • A failed run must be manually unwound before re-running
  • Convergence guarantees any single run reaches the end state
  • A shell step inside a declarative run is automatically convergent

context