skip to content

Your team has split a Terraform estate into about thirty state files, one per component. What does that granularity start costing you, and how do you judge whether it has gone too far?

level: principalimportance: nice to knowfreq 33%

answer

  1. no single plan for the whole change
  2. who sequences the applies
  3. two states always in the same pull request
  4. coordination overtakes blast-radius benefit
  5. merging back is state mv, not recreate

basics

~20 s

Every seam adds a handoff: no single plan shows a cross-cutting change, applies must run in a chosen order, and one edit ripples through several pipelines. Split for blast radius and ownership; stop when routine changes need three states coordinated.

solid answer

~50 s

Splitting keeps paying until the coordination cost overtakes the plan-time and blast-radius benefit, and past that point it reverses. The concrete costs are loss of a single plan — nobody can see the whole change before it happens — ordered multi-apply choreography with a half-applied window in between, cross-state read chains where one change propagates through several pipelines in sequence, and a fixed overhead per state of a backend key, credentials, a pipeline and an owner. My signals that it has gone too far: an ordinary change touches three or more states in a required order; the seams don't line up with team ownership, so one team routinely applies four states; someone builds a home-grown orchestrator to sequence applies; or the estate can no longer be recreated from scratch without a hand-written runbook. The fix is to merge the states that always change together — a `moved` block or `terraform state mv` moves resources without recreating them.

go deeper

for a junior

Know that infrastructure is usually split across several Terraform configurations, and that a change spanning two of them means two separate plans and applies.

for a middle

Explain what the seam costs in practice: ordering is manual because no dependency edge crosses states, and a value must be published on one side and read on the other.

for a senior

Diagnose an over-split estate from its symptoms — long ordered apply runbooks, read chains, a home-grown sequencer — and know that merging is a state-move exercise, not a rebuild.

for a principal

Own the curve: articulate where coordination cost overtakes blast-radius benefit, set the organisation's default estate shape against team ownership, and keep the cross-state dependency graph shallow and acyclic.

## The curve has two sides The case for splitting is real: shorter refresh and plan, a lock over fewer resources, credentials scoped to what the pipeline actually manages, and a mistake that cannot reach the whole estate. All four improve as states get smaller. What does *not* improve is coordination, and past some granularity coordination is the dominant cost. A principal-level answer is the shape of that curve, not a rule about how many states are correct. ## Cost 1 — you lose the single plan Inside one configuration, `terraform plan` is a complete preview: everything that will change, in dependency order, before anything happens. Across thirty configurations there is no such artifact. A change that spans four of them is four plans, reviewed separately, applied at different times, and no reviewer ever sees the whole thing. That is a genuine loss of the property IaC was adopted for, and it is the reason "just split more" is not free. ## Cost 2 — ordering becomes yours Terraform's graph does not cross state boundaries. Where one configuration would have derived the order, thirty require it to be written down and then executed correctly, every time, by a human or a pipeline. In between the applies the estate is in a state that no configuration describes. Rollback is worse: reverting a four-state change means reversing four applies in the opposite order, and any that already succeeded may not reverse cleanly. ## Cost 3 — read chains propagate When A publishes to B and B publishes to C, changing A means applying A, then B, then C — even though only A's code changed — because each downstream reads the upstream's recorded values. Chains longer than two are a strong smell: they mean the seam cut through a unit that actually changes as a whole. A useful discipline is to keep the dependency graph *between* states shallow and acyclic, and to notice that a cycle between two states is not a design, it is two states that must be one. ## Cost 4 — fixed overhead per state Each state costs a backend key, a lock entry, a CI pipeline, a set of variables and provider credentials, a set of version constraints and a lock file, and an owner who notices when it fails. Thirty of these is thirty of everything, including thirty places for provider versions to drift apart. ## Cost 5 — bootstrap and disaster recovery The honest test of estate shape is: can we build this environment from nothing? With a handful of states that is a short ordered list. With thirty and non-trivial cross-reads it becomes a runbook nobody has run since the last time it was wrong, which quietly means the answer to "can we rebuild the region" is no. ## The signals that it has gone too far - A routine, everyday change requires applying three or more states in a specific order. - Two states are always changed in the same pull request. They are one unit wearing two hats. - Someone has written an in-house orchestrator whose only job is to sequence `terraform apply` calls. That tool is a symptom. - Seams do not match ownership: one team routinely applies six states, or one state has three teams in it. - There is a dependency cycle between two states, worked around by applying one of them twice. - Plan time is already seconds, so further splitting buys nothing measurable while still adding a handoff. ## The signals it has not gone far enough The opposite pressure is just as real, and a good answer names both: plans in the many-minutes range, CI jobs queueing on a lock, a pipeline role that needs delete permission on the production database in order to tag an instance, or a review culture where nobody reads the plan because it is four hundred lines every time. ## Correcting it Merging is more work than splitting but it is a solved problem: bring the resources into one configuration and move them without recreating them — `moved` blocks for refactors inside a configuration, and `terraform state mv` with the `-state-out` form for relocating resources between state files, always against a backed-up copy of both states. Then delete the redundant seam: the output, its consumer read, the pipeline and the backend key. ## Where the answer usually lands Environment is the outer cut and is close to non-negotiable. Within an environment, a small number of states aligned to cadence, blast radius and team ownership — foundation, platform, and one per service team — is what most estates settle on. The target is that a team's normal work touches exactly one state, and that cross-state changes are rare enough to deserve a written plan when they happen.

  • Two of your states always change together in the same pull request. What does that tell you?
    That the seam cuts through a single unit of change. The split is buying nothing — the lock and the blast radius were never actually independent — while charging you two plans, an ordering constraint and a published handoff. Merge them: bring the resources into one configuration with `terraform state mv`, then delete the output, the consumer read and the redundant pipeline.
  • How do you merge two state files without destroying and recreating resources?
    Move the records, not the infrastructure. Back up both state files, then use `terraform state mv` with `-state-out` to relocate resource addresses from the source state into the target, with the corresponding configuration moved into the target root module. Run a plan afterwards and require it to show no changes — a proposed destroy means an address did not land where you thought.
  • What single question best distinguishes a healthy split from an over-split estate?
    "How many states does an ordinary change touch?" One is healthy. Two occasionally is fine. Three routinely, in a required order, means the seams do not match how the system actually changes — and the cost of that shows up as sequencing runbooks, an in-house apply orchestrator, and a disaster-recovery story nobody has rehearsed.
  • Why is a dependency cycle between two states an urgent problem rather than an inconvenience?
    Because nothing can resolve it. There is no ordering that satisfies both directions, so the workaround is always to apply one state, apply the other, then apply the first again — a sequence no plan describes and nobody can review. It also makes bootstrap and rebuild impossible to automate. Treat a cycle as proof that the two states are one.

saying these in an interview costs you the question

  • Assuming more state files are always better isolation
  • Ignoring that no single plan spans multiple states
  • Building an orchestrator instead of fixing the seams
  • Believing merged states require destroying and recreating resources
  • Splitting further when plans already run in seconds

context