skip to content

Compare the three ways one Terraform configuration can consume another's results: the terraform_remote_state data source, a normal provider data lookup, and a value published to a store such as SSM Parameter Store.

level: seniorimportance: should knowfreq 50%

answer

  1. all three deliver the same string
  2. what does the consumer's IAM policy have to allow
  3. record versus live reality
  4. the tag is an unwritten contract
  5. publishing makes the interface explicit

basics

~20 s

terraform_remote_state reads the producer's entire state file — simple, but it needs broad read access. A provider data lookup queries the live API by tag or name — looser, but it relies on naming discipline. Publishing to SSM Parameter Store makes the contract explicit and narrowly scoped.

solid answer

~50 s

`terraform_remote_state` is the tightest coupling and the least work: the consumer reads the producer's state object directly, so it needs read on that whole file — every attribute the producer recorded, secrets included — and it sees the producer's *last applied* values rather than reality. A provider data lookup such as `data "aws_vpc"` filtered by tag queries the live API, so it is decoupled from how the other side is managed, works even when the producer is not Terraform, and needs only a read permission on that one resource type; the price is that the tag or name convention becomes an unwritten contract and an ambiguous filter fails the plan. Publishing to SSM Parameter Store — the producer writes an `aws_ssm_parameter`, the consumer reads `data "aws_ssm_parameter"` — is the most deliberate: the parameter path *is* the contract, permission scopes to that path, and non-Terraform consumers can read it too. It costs one more resource to own and a stale value if nobody rewrites it.

code

hcl · 27 lines
hcl
# 1. read the producer's state file directly
data "terraform_remote_state" "network" {
  backend = "s3"
  config = {
    bucket = "acme-tfstate"
    key    = "network/prod/terraform.tfstate"
    region = "eu-west-1"
  }
}

# 2. ask the live API, by convention
data "aws_vpc" "main" {
  tags = { Name = "prod" }
}

# 3. read an explicitly published contract
data "aws_ssm_parameter" "vpc_id" {
  name = "/network/prod/vpc_id"
}

output "three_paths_to_the_same_id" {
  value = [
    data.terraform_remote_state.network.outputs.vpc_id,
    data.aws_vpc.main.id,
    data.aws_ssm_parameter.vpc_id.value,
  ]
}

go deeper

for a junior

Know that a value can cross between configurations in more than one way, and be able to name reading the other state versus looking the resource up by tag.

for a middle

Explain the mechanics of each: which reads a stored state file, which queries the live API, and which reads a parameter someone deliberately wrote.

for a senior

Lead with the permission and freshness trade — that a remote state read grants the whole file, and that a live lookup fails loudly rather than serving a stale record — and pick per trust boundary.

for a principal

Own it as an interface standard: which handoff is the organisation's default, how cross-team values get published and versioned, and how a rename is rolled out when nothing across the seam is compile-checked.

## The seam is where the coupling lives Once an estate is split, the interesting design decision is not the split — it is how the pieces read each other. All three options deliver the same string, and they differ entirely in what they *couple*: what permission the consumer needs, what breaks when the producer changes, and whether the value reflects the record or reality. ## Option 1 — the terraform_remote_state data source ```hcl data "terraform_remote_state" "network" { backend = "s3" config = { bucket = "acme-tfstate" key = "network/prod/terraform.tfstate" region = "eu-west-1" } } # data.terraform_remote_state.network.outputs.vpc_id ``` **Coupling:** on the producer's *implementation*. The consumer must know which backend, which bucket and which key — move the state and every consumer breaks. Rename an output and every consumer's plan errors. Only root outputs are visible, which is the one healthy constraint here: the producer chooses its surface. **Permission:** all-or-nothing. `s3:GetObject` on that key plus `kms:Decrypt` gives read of the *entire* state file, not just the outputs. State records resource attributes verbatim, and plenty of those are secrets in practice. This is why it is a poor choice across a trust boundary — a team that should see one VPC id ends up able to read everything the producing team manages. Sensitivity marks propagate into the consumer's plan output, but they do not restrict who can read the file. **Freshness:** it reflects the producer's last apply. If the producer has an unapplied change, or someone edited the resource in the console, the consumer plans against the stale record. No lock is taken, so a read during someone's apply can even catch a partially written view. **Verdict:** the default when both sides are Terraform, owned by the same team, in the same trust boundary. In HCP Terraform the equivalent is the workspace's remote-state-sharing setting, or the `tfe_outputs` data source, both of which at least make the grant explicit per workspace rather than per bucket key. ## Option 2 — a provider data lookup ```hcl data "aws_vpc" "main" { tags = { Name = "prod" } } data "aws_subnets" "private" { filter { name = "vpc-id" values = [data.aws_vpc.main.id] } tags = { Tier = "private" } } ``` **Coupling:** on a *convention*, not on an implementation. The consumer never learns where the producer's state lives, or whether a producer exists — this works identically if the VPC was made by CloudFormation, by another team's Terraform, or by hand. The unwritten contract is the tag or name, and it is unwritten precisely because nothing enforces it; someone retagging "prod" to "production" breaks a plan in a repository they have never opened. **Permission:** narrow and natural — `ec2:DescribeVpcs` and friends, which the consumer's role almost certainly already has. No access to anyone's state file. **Freshness:** it reads *live* reality at plan time, which is the honest answer to "what is actually there". The failure mode is the opposite of stale: if the resource does not exist yet, or if the filter matches zero or several resources, the plan fails outright. It cannot wait for a producer. **Verdict:** best across team and tooling boundaries, and the only option when the producer is not Terraform. Make the convention real — enforce the tag with policy-as-code rather than a wiki page. ## Option 3 — an explicit published contract ```hcl # producer resource "aws_ssm_parameter" "vpc_id" { name = "/network/prod/vpc_id" type = "String" value = aws_vpc.main.id } # consumer data "aws_ssm_parameter" "vpc_id" { name = "/network/prod/vpc_id" } # data.aws_ssm_parameter.vpc_id.value ``` **Coupling:** on a named path that exists on purpose. The producer is deliberately publishing; the parameter path is a versioned, greppable interface, and moving the producer's state or retagging the VPC changes nothing. **Permission:** the tightest of the three — `ssm:GetParameter` on `/network/prod/*` and nothing else. Non-Terraform consumers (an application at boot, a deploy script, a Lambda) read the same path, which the other two options cannot offer. **Freshness:** exactly as fresh as the last producer apply, with the added failure mode that the parameter and the real resource can diverge if the parameter is written outside the producer's Terraform. Keep the `aws_ssm_parameter` resource next to the resource it describes so they move together. **Verdict:** worth the extra resource when the consumer set is wide, crosses trust boundaries, or includes things that are not Terraform. ## Choosing Same team, both Terraform, same trust boundary — `terraform_remote_state`, and don't over-engineer. Different team or different tool — a provider lookup on an enforced tag. Many consumers, or consumers outside Terraform, or a permission boundary you must not blow through — publish it explicitly. Whichever you pick, the values that cross the seam are an API: name them deliberately, change them with notice, and keep the list short.

  • Why is terraform_remote_state a poor fit across a trust boundary?
    Because the grant is the whole file. Reading one output requires read on the state object, and state records every managed resource's attributes verbatim — connection strings, generated passwords, key material that providers persist. You cannot scope it down to "just vpc_id". Across a boundary, either publish the value explicitly or let the consumer look it up live through the provider.
  • The tag-based lookup matched two VPCs after someone copied an environment. What happens, and how do you prevent it?
    The plan fails — `aws_vpc` requires exactly one match and errors on multiple results, which is at least loud rather than silently picking one. Prevent it by making the filter genuinely unique (add an environment and owner tag, or filter on a naming convention plus tag) and by enforcing the tagging rule with policy-as-code so an untagged or duplicated resource never reaches the account.
  • Which option copes best when the producing side is not Terraform at all?
    The provider data lookup, followed by the published parameter. A live API query does not care what created the resource, and an SSM path can be written by any tool. `terraform_remote_state` is the only one that structurally requires the producer to be Terraform, because it parses a Terraform state file.
  • How do you evolve a value that many configurations already consume?
    Treat it as an API change: add the new output or parameter path alongside the old one, migrate consumers, then remove the old one in a later change. Because there is no compiler across the seam, nothing tells you who reads it — so keep an owner list, grep the estate before renaming, and prefer a small, stable published surface over exposing many outputs.

saying these in an interview costs you the question

  • Thinks reading one output grants access to only that output
  • Assumes terraform_remote_state reflects live infrastructure
  • Treating a tag convention as enforced because it is documented
  • Publishing every output rather than a deliberate small surface
  • Believing an SSM handoff requires no producer-side resource

context