skip to content

A shared infrastructure component is consumed by dev, staging and production. How do you roll a change to that component out across the environments safely?

level: seniorimportance: should knowfreq 50%

answer

  1. environments consume versions, not branches
  2. immutable publish, exact pin
  3. bump lowest environment first, then soak
  4. preview again in prod — same version, different reality
  5. revert re-pins config, not deleted data

basics

~20 s

Publish the component as an immutable version, have every environment pin an exact version, and promote by bumping the pin one environment at a time — dev, then staging, then production — with each bump a reviewed change carrying its own preview of what will change.

solid answer

~40 s

The rule is that environments consume *versions*, not a moving branch. I publish the component change as a new immutable version, then bump the pin in dev, apply, and let it soak while something actually exercises the infrastructure. Then the identical bump in staging, then production, each as its own reviewed change with its own preview. That gives three properties: the code that reached production is the code that ran in dev; the preview in each environment shows what the bump will actually do *there*, which is where surprises show up because prod has data and traffic; and back-out is re-pinning the previous version. If every environment tracked the main branch instead, a merge would silently change production's next apply, and there would be no soak and no distinct approval per environment.

go deeper

for a junior

Be ready to say that shared infrastructure code is versioned and that each environment pins a version, so a change reaches production only when someone deliberately bumps that pin rather than automatically on merge.

for a middle

Explain the promotion sequence and what semantic versioning signals to consumers. Know that the pin is an exact version, not a branch or a range, and that each bump is reviewed on its own.

for a senior

Show why the preview must be repeated in production even after clean lower environments, and give an honest account of rollback: re-pinning restores configuration but cannot restore data a replacement destroyed.

for a principal

Own the policy: who may publish a version, how breaking changes are announced and migrated, the maximum lag you tolerate between environments, and how you stop emergency fixes from creating a production that exists nowhere else.

## The failure mode this avoids A team keeps its shared components in a repository and every environment points at the latest commit on the main branch. Someone merges an improvement on Tuesday. Nothing happens — until Thursday, when an unrelated change triggers an apply in production and picks up Tuesday's component change as a side effect, in a diff nobody expected and at a moment nobody chose. The change may even be good; the problem is that it arrived without a decision. Version pinning turns "consuming a change" into an explicit, dated, reviewed act. ## The mechanics of promotion **1. Publish an immutable version.** The component gets a version identifier that never moves — semantic versioning is conventional: patch for a fix with no interface change, minor for a new optional input, major for anything that breaks callers or forces a resource replacement. "Immutable" is the load-bearing word: if a version can be republished with different content, every guarantee below evaporates. **2. Pin exact versions in each environment.** Not a range, not a branch, not "latest". A range means two environments can resolve differently on the same day, which reintroduces the drift you were trying to remove. If your tool caches resolved versions in a lock artifact, that artifact belongs in version control for the same reason. **3. Bump one environment at a time, lowest first.** The bump is a one-line change to the environment's pin, reviewed like any other change. Its preview is the interesting part — it shows what applying the new component version will do *in that environment*. **4. Soak.** "It applied cleanly" is a weak signal. A component change can apply cleanly and still break the thing it manages. Give dev enough time and enough real traffic — integration tests, someone using the environment — to surface it before staging. **5. Repeat identically upward.** The same version number moves to staging, then to production. Nothing is rebuilt, re-resolved or re-merged along the way. ## Why the preview differs per environment, even at the same version This is the part candidates miss, and it is the reason you cannot skip the preview in production just because dev was clean. The same component version behaves differently against different actual infrastructure: production has resources dev does not, larger data, existing values that were set by hand years ago, and resources whose replacement is destructive precisely because they contain data. A component change that computes an in-place update in dev can compute a **replace** in production because production's resource has an attribute that forces new. That is why the promotion is three previews, not one. ## Interface changes need a migration path When the change removes or renames an input, the bump is not mechanical any more — every caller must change at the same time. Options, roughly in order of preference: add the new input alongside the old one and accept both for a release; publish the breaking change as a major version and let consumers adopt on their own schedule; or, when you control every caller, do a coordinated change. Publishing a breaking change as a patch and letting environments discover it is how a shared component library gets a bad reputation and teams start forking it. ## Rolling back Re-pinning the previous version and applying is the mechanism, and it works for anything the component *configures*. It does not work for anything the component *destroyed*. If the bad version replaced a data store, pinning the old version will happily create a fresh empty one — the code is reverted, the data is not. This is why the preview is the gate: reverting configuration is cheap and reverting deletion is impossible, so the check has to happen before the apply, not after. ## The organisational half A promotion policy is only real if the environments cannot drift far apart. Two guardrails help: a maximum lag — no environment more than N versions behind production's target — and a rule that emergency fixes go in at the bottom and are promoted, never patched directly into production. A hotfix applied straight to prod produces a production environment running code that exists nowhere else, which is the exact state the whole scheme exists to prevent. ## What a weak answer sounds like "We merge to main and everything picks it up." That is a promotion policy in the same sense that gravity is a landing policy. The interviewer is listening for immutable versions, an explicit pin, a preview per environment, and an honest account of what rollback can and cannot undo.

  • Why is pinning a version range instead of an exact version a problem across environments?
    Because two environments resolving the same range on different days can land on different versions, so what you validated in staging is not necessarily what production gets. It also makes the change arrive without a decision or a review. Pin exactly, and commit any resolution lock artifact so the resolution is reproducible on every machine and runner.
  • The bump applied cleanly in dev and staging. Why still require a preview in production?
    Because the preview is computed against actual infrastructure, and production's differs — resources created before the component existed, attributes set by hand, data stores whose replacement is destructive. The same version can be an in-place update in dev and a replacement in production. The clean lower environments raise confidence; they do not remove the need to look.
  • How do you publish a change that removes an input the component previously accepted?
    As a major version, so no consumer picks it up implicitly. Where possible, ship a release that accepts both the old and the new input first, migrate callers, then remove the old one in the next major. Shipping a breaking interface change as a patch is what makes teams stop trusting the shared library and fork it.

saying these in an interview costs you the question

  • Every environment points at the main branch of the component repository.
  • It applied fine in dev, so skip the production preview.
  • Rolling back the version undoes everything the bad version did.
  • Hotfix production directly, then backport to dev when there is time.
  • Version ranges are fine; the tool always picks the same thing.

context