skip to content

A configuration-management tool runs successfully against every host every night, yet two hosts that should be identical behave differently. Why does repeated in-place convergence still leave divergent servers, and how does replacing hosts instead of patching them remove that class of problem?

level: middleimportance: should knowfreq 44%

answer

  1. green means the declared subset matched
  2. deleting code does not delete its effects
  3. install date is a hidden input
  4. undeclared surface is invisible to the run
  5. no history, because there was no previous host

basics

~20 s

A convergence run only asserts what it currently declares. It does not remove what it stopped declaring, does not pin what was never pinned, and cannot undo manual edits, so hosts accumulate divergent history. Replacement discards that history on every change.

solid answer

~50 s

A green convergence run means the declared subset of the host matched, not that the host as a whole matches any other host. Three gaps produce divergence. First, removal: deleting a task from the code does not delete what it previously installed, so a host built last year still carries software no current host has. Second, time: anything not pinned to an exact version resolves to whatever was current on the day that host first ran it, so build date silently becomes part of the configuration. Third, everything undeclared — manual fixes, files written by an incident, kernel settings, leftover cron entries — is simply invisible to the run. Replacement removes the class of problem because a fresh instance starts from one artifact with no history, so its configuration is described by an image identifier and divergence is bounded by how long instances live.

go deeper

for a junior

Recall that a configuration run only checks what the code declares, so anything installed or edited outside it stays on the host unnoticed.

for a middle

Explain the three mechanisms precisely — no removal of undeclared state, unpinned versions resolving at install time, and the undeclared surface — and why replacement removes the accumulation rather than merely detecting it.

for a senior

Show how you would investigate a divergence in production: compare image identities and package manifests rather than diffing live file trees, then bound the problem with a replacement cadence instead of chasing individual hosts.

for a principal

Take a position on where the guarantee should live. Argue for reproducible builds and pinned inputs as the real control, and be explicit that immutability relocates the divergence risk into the pipeline rather than eliminating it.

## What "green" actually asserts A configuration-management run compares the host against the resources it *currently declares* and fixes any that differ. That is a genuinely useful guarantee, and it is much narrower than "this host matches that host". Green means: every declared resource is now in its declared state. It says nothing about the far larger surface of the machine that the code has never mentioned. Divergence lives entirely in that gap, and it arrives through three doors. ## Door one: the removal problem Deleting code does not delete its effects. If a role once installed a debugging package, wrote a config file and added a cron entry, removing those tasks from the repository stops *new* hosts from getting them — while every existing host keeps them forever, because nothing in the run asserts their absence. Converging tools are one-directional by default: they assert presence, not absence, unless you explicitly declare the removal. That leaves a trap. To actually clean up, you must add an explicit "absent" declaration and then keep it in the codebase indefinitely, because you cannot tell which hosts have already run it. Estates that have been managed this way for years accumulate a graveyard of removal tasks that exist only to undo things nobody remembers adding — and if one is deleted too early, the divergence comes back on the hosts that never saw it. ## Door two: time is an undeclared input Anything not pinned to an exact version resolves at the moment of execution. "Install the latest" on a host built in March and on a host built in September gives two different builds, both reported green. The same applies to anything fetched from a moving reference — a branch instead of a tag, a floating base image tag, an artifact endpoint that always serves "current". The host's build date becomes a hidden input to its configuration, which is why the phrase "but the code is identical" is never sufficient evidence that two hosts are identical. Partial failure compounds it. A run that fails halfway leaves a host in a state neither the old nor the new code describes, and if the failure is retried later the intermediate state may persist under the successful run that follows. ## Door three: everything nobody declared The hot-fix applied at 3am during an incident. A kernel parameter tuned by hand and never written back. A file left in place by a support engineer. A package pulled in as a transitive dependency of something else. A log directory that filled a disk on one host and not another. None of it is visible to a run that only inspects what it declares, and all of it can change behaviour. ## How replacement removes the class of problem The immutable answer does not detect divergence better — it removes the accumulation that causes it. A new instance boots from one artifact with no history. There is nothing left over from a previous configuration, because there was no previous configuration. Three properties follow: - **Identity becomes checkable.** "Are these two hosts the same?" collapses to comparing image identifiers, instead of diffing file trees on live machines. - **Divergence is time-bounded.** Whatever an operator does to a running host disappears the next time that host is replaced, so the maximum age of any undeclared change is the fleet's replacement cadence. - **Removal is free.** Deleting the line that installed a package genuinely removes it from the next image, because the image is rebuilt from scratch rather than converged forward from its own past. ## The honest caveat: the problem moves to the build Immutability does not delete the issue, it relocates it. If the image build itself is not reproducible, two builds of the same source produce different images and you have snowflake *images* instead of snowflake hosts — the same failure, one layer up, where at least it is visible in a pipeline rather than hidden on a machine. So the same discipline applies at build time: pin package and dependency versions, prefer immutable references (a digest or tag over a branch or a floating tag), and record a manifest of what actually went into each image so you can answer "what changed between these two builds?" without guessing. And note that replacement changes what the on-host tooling is *for*: instead of converging long-lived servers forever, the same code runs once during the image build, in a controlled environment, where a failure is a red build rather than a partially-configured production host.

  • Why does deleting a task from a role fail to remove what it previously installed, and what is the correct way to clean it up?
    Converging tools assert the presence of declared resources, not the absence of undeclared ones, so removing the code only stops new hosts from receiving it. To actually clean up you must add an explicit removal declaration and keep it in the codebase until every host that could have the old state is gone — which in practice means until the fleet has been fully replaced. That retention burden is one of the strongest practical arguments for rebuilding instead.
  • If immutability moves the problem into the image build, what makes a build reproducible enough to rely on?
    Pin every version you can: package versions, language dependencies via a lock file, and base images by digest rather than a floating tag. Fetch from artifact stores you control rather than whatever is current upstream. Record a manifest of what actually went into the image so two builds can be diffed. Without this you get snowflake images instead of snowflake hosts — the same defect, but at least visible in a pipeline.
  • Does immutability mean configuration-management tooling is no longer useful?
    No — its position changes. The same declarative code that converged long-lived servers is an excellent way to build the image, running once in a controlled environment where a failure is a red build instead of a half-configured production host. What goes away is the expectation that repeated runs against long-lived hosts will keep them identical to each other over time.

saying these in an interview costs you the question

  • Assumes a green convergence run proves two hosts are identical
  • Thinks removing a task from the code removes what it installed
  • Ignores that unpinned versions make build date part of the configuration
  • Believes immutability removes the need for reproducible builds
  • Claims manual hot-fixes are caught by the next scheduled run

context