skip to content

In a Terraform root module, a `data "aws_subnets"` lookup filtered by a tag on a VPC that the same configuration creates fails on the first apply in a fresh account but succeeds when you re-run. Why, and how do you fix it?

level: middleimportance: should knowfreq 38%

answer

  1. ordering comes from references
  2. a literal filter creates no edge
  3. read at plan time, before creation
  4. works on the second run — the tell
  5. reference it, do not look it up

basics

~20 s

The filter is a literal string, so Terraform sees no dependency on the VPC and reads the data source during plan, before the VPC exists. The re-run succeeds only because the partial first apply left the objects behind. Reference the created resources directly instead of looking them up.

solid answer

~50 s

Terraform builds its ordering from references. A data source whose filter is a hard-coded tag value contains no reference to `aws_vpc.main`, so there is no edge in the graph and the read runs at plan time — before anything is created. In a fresh account the lookup matches nothing and fails, or returns an empty list that breaks something downstream. Re-running appears to fix it because the first apply already created the VPC and subnets, so the second read finds them; that "works the second time" symptom is the tell for this bug. The real fix is not to look up what you create: reference `aws_subnet.private[*].id` directly, which creates a proper dependency and always resolves. If the lookup has to stay, `depends_on = [aws_vpc.main]` on the data block defers the read to apply, at the cost of making its results unknown at plan time.

code

hcl · 13 lines
hcl
resource "aws_vpc" "main" {
  cidr_block = "10.0.0.0/16"
}

resource "aws_subnet" "private" {
  count      = 2
  vpc_id     = aws_vpc.main.id
  cidr_block = cidrsubnet(aws_vpc.main.cidr_block, 8, count.index)
}

resource "aws_lb" "app" {
  subnets = aws_subnet.private[*].id
}

go deeper

for a junior

Know that Terraform decides ordering from references between blocks, and that a hard-coded string in a filter is not a reference — so the lookup can run before anything is created.

for a middle

Explain the full chain: no edge in the graph, read during plan, empty match in a fresh account, and a second run that succeeds only because the partial first apply left objects behind. Name both fixes and their tradeoff.

for a senior

Recognise the failure signature — works everywhere except a clean bootstrap — and push the fix upstream: reference what you own, and put genuine cross-boundary lookups behind a separate root module applied first. Argue against reflexive depends_on.

for a principal

Own the convention that a configuration must be provable from an empty account, and back it with a periodic from-scratch build in CI. Decide where ownership boundaries sit so lookups only ever read things another pipeline stage guarantees exist.

## Terraform orders by reference, not by proximity Terraform builds a dependency graph from the references in your configuration. `aws_subnet.private.vpc_id = aws_vpc.main.id` creates an edge because the expression names the VPC. A string literal creates nothing: ```hcl resource "aws_vpc" "main" { cidr_block = "10.0.0.0/16" tags = { Name = "platform" } } data "aws_subnets" "private" { filter { name = "tag:Network" values = ["private"] # a literal — no edge to anything } } ``` As far as the graph is concerned, `data.aws_subnets.private` has no dependencies at all. All of its arguments are known, so Terraform reads it during plan — which happens before a single API call has created anything. In a fresh account the read matches zero subnets. Depending on the data source, that is either a hard error (data sources that require exactly one match, like `aws_vpc` or `aws_ami`, fail with "no matching resources found") or an empty result that breaks a downstream expression such as an index or a `length()` assumption. ## Why the re-run "fixes" it The first apply usually gets partway through — Terraform created the VPC and subnets before something else failed, or the failure came from a later expression. Those objects now exist and are in state. On the second run the plan-time read finds them, everything resolves, and the apply succeeds. The configuration looks fine forever afterwards, until someone builds a new environment and hits the same wall. This is the signature of the bug: **it only reproduces on a clean bootstrap.** It survives code review, it survives every plan on the existing environments, and it fails exactly when you are standing up disaster recovery or a new region under time pressure. ## The fix, in order of preference **1. Do not look up what you create.** This is the actual answer, not a workaround. If the subnets are declared in this configuration, refer to them: ```hcl resource "aws_lb" "app" { subnets = aws_subnet.private[*].id } ``` The reference creates the dependency edge, the value resolves during apply in the right order, and the configuration works on a clean account by construction. It also documents intent: these are *my* subnets, not somebody's. **2. If the lookup must stay, add `depends_on`.** A `data` block accepts `depends_on`, and it exists for exactly this situation: ```hcl data "aws_subnets" "private" { filter { name = "tag:Network" values = ["private"] } depends_on = [aws_subnet.private] } ``` When the listed dependency has changes pending in the current plan, Terraform postpones the read until apply, after the dependency is created. Note the price: a deferred read has unknown attributes, so everything downstream renders as `(known after apply)` and the plan stops being a readable preview. You have traded review quality for ordering. **3. Split the configuration.** If the lookup is genuinely about a foundation layer — a network built by a different team or a different pipeline stage — then it belongs in a separate root module applied first, and the lookup here reads something that always already exists. That is a boundary decision rather than a syntax fix, and it is often the right one when the same tag-based lookup keeps appearing. ## Why `depends_on` is not the default advice Engineers who learn this trick tend to sprinkle `depends_on` onto data blocks defensively. That is a mistake: every one of them pushes a read to apply time whenever its dependency has pending changes, and the cascade of unknown values makes plans progressively less reviewable. `depends_on` on a data source is a deliberate statement that a read is meaningless until a change lands — not a general safety measure. ## The diagnostic habit When a Terraform configuration fails only on the first run in a new environment, look for reads that should have been references. Ask of every data block: could this object have been created by this same configuration? If the answer is yes, the block is probably wrong regardless of whether it currently works.

  • What is the cost of the `depends_on` fix on a data block?
    When the listed dependency has pending changes, the read is deferred to apply, so the data source's attributes are unknown at plan time and every value derived from them renders as `(known after apply)`. The apply is still correct, but the plan a reviewer approves shows far less about what will happen.
  • How would you catch this class of bug before someone builds a new environment?
    Stand up a throwaway environment from scratch in CI periodically — a clean-account bootstrap is the only test that exercises the empty-lookup path. Reviewing every data block for objects the same configuration creates catches it statically, but only a fresh apply proves it.
  • The lookup is tag-based and matches subnets created by another team's configuration. Is it still wrong?
    No, that is a legitimate use — the objects live outside this configuration's ownership, so a lookup is the honest expression of the boundary. Make the match precise enough that it cannot silently return the wrong set, and be aware that the other team's tag rename becomes your plan failure.

saying these in an interview costs you the question

  • Says data sources always run after resources in the same apply
  • Adds `depends_on` to every data block as a precaution
  • Concludes the first failure was a transient API error
  • Thinks a tag filter creates a dependency on the tagged resource
  • Uses a lookup for objects the same configuration creates

context