skip to content

A shared Terraform module is pinned at v1.4.0 by twenty root configurations, and you need to ship a breaking v2.0.0. How do you roll that out?

level: seniorimportance: should knowfreq 50%

answer

  1. Twenty consumers, twenty migrations
  2. Additive release, never in place
  3. The changelog predicts the plan
  4. Pilot on the smallest blast radius
  5. Somebody owns the version inventory

basics

~20 s

Publish v2.0.0 as a new immutable version without touching v1, keep v1 patchable during a deprecation window, then move consumers one at a time — pinning each, reading its plan, applying, and only then proceeding. Never edit a released tag.

solid answer

~50 s

Treat the upgrade as twenty independent migrations, not one release. Publish `v2.0.0` alongside `v1.4.x` — never rewrite the old tag — and keep the v1 line alive for patches during a stated deprecation window so nobody is forced to move on your schedule. Write a migration note that says exactly what a consumer's plan will show: which inputs were renamed, which defaults changed, which resources will be replaced. Then upgrade in order of blast radius: a sandbox consumer first, read the plan carefully for destroys, apply, and let it soak before moving to the next. Where v2 renamed resource addresses inside the module, ship refactor declarations so consumers see a move rather than a destroy-and-create. Keep the pin visible in each consumer's diff — a version bump should be a reviewed one-line change, which is also why a floating constraint would have made this rollout unauditable.

go deeper

for a junior

Know that upgrading a module means editing the pinned version in the caller and running init before plan, and that you read the plan before applying.

for a middle

Explain the mechanics: publish a new version rather than moving the old tag, why a major bump signals plan-visible breakage, and why each consumer needs its own plan review.

for a senior

Demonstrate the rollout judgment — pilot first, sequence by blast radius, keep the old line patchable during a deprecation window, and predict destroys in the release notes rather than discovering them in production.

for a principal

Own the module as a published API: deprecation policy and end-of-support dates, who maintains the inventory of consumers, and whether the organisation tolerates any pinning scheme that lets a version change without a reviewed commit.

## The rollout is per consumer, not per module The module repository has one release event; the estate has twenty. Each root configuration owns real resources, has its own state, its own maintainers and its own risk profile, and the only thing that tells you what v2.0.0 does to it is *that consumer's plan*. Any plan that treats "upgrade the module" as a single change is going to be wrong for at least one of the twenty. ## Publish additively, never in place The first rule is that `v1.4.0` keeps meaning what it meant. Force-pushing a tag or re-publishing a registry version silently changes infrastructure code for everyone who has not upgraded, and it destroys the only property pinning was supposed to buy. Publish `v2.0.0` as a new tag or a new registry version, keep a `v1` maintenance line for security and bug fixes during the deprecation window, and state when that line ends. ## Make the breaking change legible A module changelog for infrastructure is not a list of commits; it is a prediction about plans. Say which inputs were removed or renamed, which defaults changed, which outputs disappeared, and — most importantly — which resources will be **replaced**. If a consumer has to rename variables, give the mapping. If a resource will be destroyed and recreated, say so at the top, because that is the sentence that decides whether the upgrade happens on a Tuesday afternoon or during a maintenance window. Where v2 only *refactored* — the same resources under new addresses inside the module — the module author can ship refactor declarations (`moved` blocks) with the release so that consumers' plans show a move rather than a destroy and create. Shipping that in the module is far cheaper than twenty teams each performing state surgery. ## Sequence by blast radius A practical order: 1. **Pilot.** Pick a low-stakes consumer — a sandbox or an internal tool — pin it to `v2.0.0`, run `terraform init` then `plan`, and read every line. Apply, then let it run long enough to reveal anything the plan did not show. 2. **Fix in the module.** Anything surprising in the pilot is usually a module bug: ship `v2.0.1` rather than asking nineteen teams to work around it. 3. **Fan out by tier.** Non-production consumers next, then production, one at a time, each as its own reviewed pull request with the plan attached. 4. **Track.** Keep a simple inventory of which consumer is on which version. Twenty repositories with no inventory is how you discover a forgotten consumer eighteen months later, still on v1, when the v1 line is long dead. ## Why the pin shape matters here This rollout is only controllable because each consumer pins explicitly. With a Git source, `?ref=v2.0.0` is a one-line diff a reviewer can see. With a registry source and a constraint like `~> 1.4`, a *major* bump will not be picked up automatically — which is exactly why the pessimistic operator is the right default for shared modules — but a loose `>= 1.4` would have dragged consumers into v2 the next time any of them ran `init` on a clean machine, with no code change and no review. The upgrade discipline described here is what pinning is *for*; if a consumer's version can change without a commit, none of the sequencing above is enforceable. ## Talking about it in an interview The strong answer is not a tool trick — it is the recognition that a shared module is a published API with unknown blast radius behind each caller. The signals an interviewer is listening for: never mutate a released version; keep the old line alive so migration is opt-in on the consumer's schedule; predict plan behaviour in the release notes; pilot before fanning out; and treat "which consumers are on which version" as an inventory somebody owns rather than something you rediscover during an incident.

  • Why not just move every consumer to v2.0.0 in a single coordinated change window?
    Because the twenty plans are not the same plan. Each consumer has different inputs, different existing resources and a different tolerance for replacement, so a single window means twenty simultaneous unreviewed applies with no way to stop after the first surprise. Sequencing costs calendar time and buys the ability to abort cheaply.
  • A consumer's v2 plan shows the module's database being destroyed and recreated. What do you do before applying?
    Stop and find out whether that is intended. If v2 genuinely re-declares the resource under a new address, the fix belongs in the module as a refactor declaration so the plan becomes a move; if a changed input forces replacement, decide whether the consumer can keep the old value. Applying a destroy you did not intend to ship is not an upgrade, it is an outage.
  • How do you stop a consumer being forgotten on v1 forever?
    Publish an end-of-support date with the v2 release, keep an inventory of which root configuration is pinned where, and make the inventory something automation can produce — the pins are in code, so they are greppable across repositories. Then chase the stragglers before the v1 line stops receiving fixes, not after.

saying these in an interview costs you the question

  • Re-tags v1.4.0 to point at the new code
  • Upgrades all consumers in one change
  • Treats a breaking change as a minor bump
  • Assumes a plan for one consumer predicts the rest
  • Has no record of which consumer uses which version

context