skip to content

Splitting State and Remote State

One giant state file means slow plans and a blast radius that covers everything. Splitting it raises the next question: how do the pieces read each other's outputs without coupling too tightly?

part ofTerraformoverview, primer and where to startread it →
on this pageshow

questions

4

How do you decide where to split a Terraform estate into separate state files, and what does splitting actually buy you?

level: middleimportance: must knowfreq 68%

answer

  1. cadence and blast radius
  2. how much can one bad apply destroy
  3. plan time, lock contention, credentials
  4. every seam becomes a handoff you order yourself

basics

~20 s

Split along blast radius, change cadence and ownership: long-lived networking in one state, fast-moving application resources in another. The payoff is shorter plans, less lock contention, narrower credentials per state, and a mistake that cannot reach the whole estate.

solid answer

~50 s

I draw seams where the *rate of change* and the *cost of a mistake* differ. Networking, DNS zones and shared data stores change a few times a year and are catastrophic to destroy; application resources change several times a day and are cheap to recreate — keeping them in one state means every routine deploy runs a plan that could propose deleting the VPC. Ownership is the second axis: one state per team that can be granted its own backend key and its own provider credentials. Region and account boundaries are a natural third. The operational payoff is concrete: `terraform plan` refreshes only the resources in that state, so plans go from many minutes to seconds; the state lock is held over a much smaller set of resources, so two teams stop queueing behind each other in CI; and the credentials for the app state need no permission over the network resources. The cost is that every seam becomes a handoff you now have to design and order.

go deeper

for a junior

Know that a Terraform estate is normally many root configurations, not one, and that each has its own state file and its own backend location.

for a middle

Explain the three concrete wins — shorter refresh and plan, a lock over fewer resources, and credentials scoped to that state — and name cadence and blast radius as the axes you cut on.

for a senior

Show judgment about the cost side: every seam is a handoff you must publish, read and order, and you can describe how you would migrate an existing monolithic state without recreating resources.

for a principal

Own the estate shape: how seams map to team ownership and credential boundaries, what the organisation's default layout is, and how you stop it fragmenting into thirty states nobody can change together.

## The symptom that starts the conversation A single root configuration holding an entire environment behaves fine at twenty resources and badly at two thousand. Three things degrade together. `terraform plan` refreshes every managed resource, so plan time grows roughly with resource count and turns a one-line change into a coffee break. The state lock is taken for the whole configuration, so any apply blocks every other change in the environment — in a busy CI pipeline that becomes a queue. And the blast radius is total: the plan for a tag change on one instance is also, mechanically, a plan that could destroy the VPC if someone mis-edits a variable, and the credentials the pipeline runs with must be powerful enough to manage everything in the file. Splitting the estate into several root configurations, each with its own state, addresses all three at once. The design question is where the cuts go. ## Axis one: change cadence Group resources by how often they legitimately change. A rough tiering that survives most estates: - **Foundation** — accounts, VPCs, subnets, transit gateways, DNS zones. Changes measured in months. Reviewed carefully, applied rarely. - **Platform** — clusters, shared databases, load balancers, IAM roles. Changes weekly. - **Application** — services, task definitions, per-service queues and buckets. Changes hourly, often from an app team's own pipeline. Putting a monthly-cadence resource in the same state as an hourly-cadence one forces the slow thing to be re-planned on every fast change, and forces the fast pipeline to hold a lock over the slow thing. ## Axis two: blast radius Ask what a bad apply in this state can destroy. A state containing the production database and a state containing stateless compute deserve different review gates, different approvals and different credentials. Splitting is what makes those differences *enforceable* rather than a convention: the app pipeline's role simply has no `rds:DeleteDBInstance`, because it never manages that resource. `prevent_destroy` on a critical resource is a good belt, but the separate state is the braces. ## Axis three: ownership and permissions A state file is the smallest unit you can hand to a team. One backend key, one CI pipeline, one set of provider credentials, one on-call owner. If two teams routinely edit the same state, they will contend on the lock and review each other's unrelated diffs. If a state has no clear owner, nobody applies it and it drifts. ## Axis four: provider and account boundaries Separate cloud accounts, or separate regions with genuinely independent lifecycles, are natural seams — they usually also want separate credentials, and keeping them in one configuration pushes you into provider aliases for what could simply be two configurations. ## What splitting costs Every seam becomes a **handoff**. The app configuration needs the VPC id that the network configuration produced, so someone must publish it as an output and someone must read it — through `terraform_remote_state`, a provider data lookup by tag, or a value published to SSM Parameter Store. That read is a coupling: rename the output and the downstream plan breaks. The seam also breaks atomicity. Inside one configuration Terraform derives ordering from the dependency graph; across two, no such edge exists. A change that spans both is now two applies in a deliberate order, with a window in between where the estate is half-changed, and no single plan that shows the whole thing. On a fresh environment you must bootstrap in dependency order. And there is a fixed overhead per state: a backend key, a pipeline, a lock table entry, a set of variables, a place for credentials to live. ## How to answer in practice Start from the environment as the outermost cut — separate state per environment is close to non-negotiable, because it is what stops a staging apply from touching production. Inside an environment, cut on cadence and blast radius until each state is something one team can own and plan in under a minute or two. Stop there. Cutting further to "one state per resource type" trades the problems you had for coordination problems you did not, and the handoffs start costing more than the plans ever did. A useful sanity check: if a routine, ordinary change to your system requires applying three states in a specific order, the seams are in the wrong place.

  • Is a separate state per environment the same decision as splitting by cadence?
    No — it is the outer cut and it is about isolation, not performance. Separate state per environment is what stops a staging apply from proposing production changes and lets each environment carry its own credentials. Splitting *within* an environment by cadence and blast radius is a second, independent decision, and you normally do both.
  • How would you demonstrate the plan-time benefit rather than assert it?
    Measure it. Plan time tracks the number of resources refreshed in that state, so time the current monolith's `terraform plan`, then time a plan for the candidate subset. Teams are usually surprised: pulling a few hundred rarely-changing foundation resources out of the loop often takes a multi-minute plan down to seconds, and it removes them from the lock window too.
  • Your app team wants their own state but shares an RDS instance with two other teams. Where does the database go?
    In the shared platform state, owned by whoever is accountable for it — not duplicated into each app state, and not managed by whichever team applies last. The app states then consume its endpoint through the seam. Shared, high-blast-radius, low-cadence resources belong on the slow side of the cut.
  • Does splitting state remove the need for state locking?
    No. It narrows the contention — a lock now covers one component rather than the whole environment — but any single state can still be applied concurrently by two runners, and that still corrupts it. Every state file needs locking regardless of how finely the estate is cut.

saying these in an interview costs you the question

  • Splitting by resource type instead of cadence or ownership
  • Believing a bigger state file is inherently faster to plan
  • Treating the handoff between states as free
  • Thinking separate states remove the need for locking
  • Keeping production and staging in one state with a variable

context

open as a page

In Terraform, how does one root configuration read a value produced by a different configuration's state, and what has to exist on the producing side?

level: juniorimportance: should knowfreq 60%

basics

~10 s

The producing configuration must declare a root-level output. The consumer adds a terraform_remote_state data source pointing at the producer's backend and reads data.terraform_remote_state.NAME.outputs.KEY. Only declared root outputs are exposed, never arbitrary resource attributes.

open as a page

Compare the three ways one Terraform configuration can consume another's results: the terraform_remote_state data source, a normal provider data lookup, and a value published to a store such as SSM Parameter Store.

level: seniorimportance: should knowfreq 50%

basics

~20 s

terraform_remote_state reads the producer's entire state file — simple, but it needs broad read access. A provider data lookup queries the live API by tag or name — looser, but it relies on naming discipline. Publishing to SSM Parameter Store makes the contract explicit and narrowly scoped.

open as a page

Your team has split a Terraform estate into about thirty state files, one per component. What does that granularity start costing you, and how do you judge whether it has gone too far?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Every seam adds a handoff: no single plan shows a cross-cutting change, applies must run in a chosen order, and one edit ripples through several pipelines. Split for blast radius and ownership; stop when routine changes need three states coordinated.

open as a page