skip to content

In Terraform, when is a data source actually read, and why does a plan sometimes show a data source and everything downstream of it as `(known after apply)`?

level: middleimportance: must knowfreq 56%

answer

  1. read during plan, not a separate step
  2. known arguments or no read
  3. unknown-ness is contagious
  4. deferred read means unknown attributes
  5. `depends_on` forces the late read

basics

~20 s

Terraform reads a data source during plan whenever all of its arguments are already known. If any argument depends on a value that does not exist yet, the read is deferred to apply, its attributes become unknown, and every value derived from them renders as (known after apply).

solid answer

~50 s

Data sources are read as part of the plan, not as a separate step, so a plan normally shows concrete values from them. That only works when the data block's arguments are fully known at plan time — literals, variables, locals built from those. The moment an argument references an attribute of a resource that has not been created yet, or an attribute that this plan is about to change, Terraform cannot perform the read and postpones it until apply. A deferred data source has no known attributes, so anything computed from it is unknown too, and the unknown-ness cascades: names, ARNs, policy documents, entire downstream resources render as `(known after apply)`. The plan is still correct and still applies, but it has stopped being a useful preview. The fix is to keep data source arguments dependent only on inputs rather than on things this run creates, or to split the work into stages.

code

bash · 17 lines
bash
terraform plan

# Terraform will perform the following actions:
#
#   # data.aws_subnets.other will be read during apply
#   # (config refers to values not yet known)
#  <= data "aws_subnets" "other" {
#       + id  = (known after apply)
#       + ids = (known after apply)
#     }
#
#   # aws_lb.app will be created
#   + resource "aws_lb" "app" {
#       + arn      = (known after apply)
#       + dns_name = (known after apply)
#       + subnets  = (known after apply)
#     }

go deeper

for a junior

Know that a plan actually calls the provider to read data sources, and that (known after apply) means Terraform cannot compute the value yet — it is not an error.

for a middle

Explain the known/unknown rule: the read happens at plan time only if every argument is already known, otherwise it is deferred to apply and its attributes become unknown. Be able to trace how that unknown-ness spreads downstream.

for a senior

Treat a plan full of (known after apply) as a degraded review gate and say what you would change to restore it — feeding lookups from inputs, referencing created resources directly, or splitting the configuration into stages.

for a principal

Own the consequence for the approval process: if plans routinely cannot be read, the human gate is theatre. Decide where configuration boundaries go so that every plan a reviewer sees is concrete, and where you accept unknowns as the cost of a bootstrap.

## Reading happens inside the plan Terraform's plan walks the dependency graph and, for every `data` block it reaches, calls the provider's read operation. This is why `terraform plan` needs working provider credentials and network access even though it changes nothing: it is refreshing managed resources and reading data sources at the same time. The read result feeds the rest of the graph, so by the time Terraform renders the diff it can print real values — an actual AMI id, an actual account number. ## The condition: are the arguments known? A data source can only be read if Terraform can compute its arguments. Terraform tracks, for every value in the graph, whether it is *known* or *unknown*. Known values are literals, input variables, locals derived from known values, and attributes of resources that already exist and are not being changed in a way that affects them. Unknown values are attributes of objects that this plan is about to create, or attributes that will be recomputed. ```hcl variable "vpc_id" { type = string } data "aws_subnets" "private" { # read during plan: vpc_id is known filter { name = "vpc-id" values = [var.vpc_id] } } resource "aws_vpc" "new" { cidr_block = "10.0.0.0/16" } data "aws_subnets" "other" { # deferred: aws_vpc.new.id is unknown filter { name = "vpc-id" values = [aws_vpc.new.id] } } ``` The first read runs at plan time. The second cannot: the VPC does not exist yet, so its `id` is unknown, so the filter is unknown, so the read is postponed to apply. ## Why the whole plan goes unknown Unknown-ness is contagious in a purely mechanical way. If `data.aws_subnets.other.ids` is unknown, then `length(...)`, any string interpolated from it, any resource argument that references it, and every attribute of those resources that the provider computes rather than echoes back, is unknown too. The plan renders each of these as `(known after apply)`. That matters for review culture more than for correctness. The apply will still do the right thing — Terraform reads the data source first, then proceeds with real values. But the artifact a human approves has stopped saying what will happen. A reviewer looking at twenty resources whose every attribute is `(known after apply)` cannot tell an intentional change from an accident, and cannot tell that a resource is about to be *replaced* rather than updated in place. "Forces replacement" annotations do still appear, which is often the only signal left. ## `depends_on` on a data block A data block accepts `depends_on`. When you set it, and the thing it depends on has changes pending in this plan, Terraform defers the read to apply — even if every argument of the data source is a literal. This is deliberate: `depends_on` on a data source exists to say "this read is only meaningful after that change lands", and the only way to honour it is to read late. So `depends_on` is itself a cause of `(known after apply)` cascades, which is why it should be added on purpose and not by reflex. ## Reducing the cascade The practical moves, in order of preference: 1. **Do not look up what you create.** Reference the resource attribute directly. The plan then knows the reference is a dependency, and the diff shows `(known after apply)` for that one attribute rather than routing an entire subtree through a deferred read. 2. **Feed data sources from inputs.** Filters built from variables, tags, or naming conventions are known at plan time and read immediately. 3. **Split the run.** If a lookup genuinely depends on something created in the same configuration, that is usually a sign of two stages: create the foundation, then build on it in a separate root module that reads the foundation as a known input. 4. **Accept it deliberately.** Sometimes a first-time bootstrap really does have unknowns everywhere. Know that the second plan, once the foundation exists, is clean — and do not treat the first noisy plan as the normal review artifact. ## A related trap Because reads happen every plan, a data source that matches multiple objects or zero objects fails at plan time, not apply time. That is usually good news — the failure is cheap — but it also means an unrelated change in someone else's account can break your plan without anything in your repository changing.

  • If the plan is full of `(known after apply)`, is a saved plan file still safe to apply?
    Yes — the saved plan records the intended actions and Terraform performs the deferred reads during apply. The risk is not correctness, it is review: the humans who approved it could not see what the values would be, so an unintended replacement can pass approval unnoticed. Treat a heavily unknown plan as a weak gate, not a broken one.
  • Does adding `depends_on` to a data block always defer the read to apply?
    It defers the read whenever the dependency has changes pending in that plan. If the dependency is already created and unchanged, Terraform can still read during plan. That conditional behaviour is why the same configuration produces a noisy plan on the bootstrap run and a clean one afterwards.
  • Why does `terraform plan` need credentials if it changes nothing?
    Because planning includes refreshing managed resources and reading every data source whose arguments are known. Both are provider API calls. A plan is read-heavy by design, which is also why a plan can fail on permissions, throttling, or a lookup that suddenly matches zero objects.

saying these in an interview costs you the question

  • Thinks data sources are all read during apply, never plan
  • Says `(known after apply)` means the plan is broken or will fail
  • Adds `depends_on` to data blocks routinely to be safe
  • Believes unknown values only affect the resource that references them
  • Thinks refresh and data source reads are the same operation

context