skip to content

When would you use Terraform's `external` or `http` data sources, and what are the risks of shipping one in a production configuration?

level: seniorimportance: nice to knowfreq 24%

answer

  1. last-resort generic reads
  2. stdin JSON in, string map out
  3. runs on every plan, everywhere
  4. result lands in state in plaintext
  5. reads must have no side effects

basics

~20 s

They are escape hatches for facts no provider exposes: external runs a program and reads a JSON object of strings from its stdout, http performs a request and exposes the response. Both run on every plan, on every machine that plans, and their results land in state.

solid answer

~50 s

I reach for them only when there is genuinely no provider data source for the fact I need — an internal service that owns an id, a registry that has no Terraform provider. `external` from the hashicorp/external provider runs a program, passes the `query` map to it as JSON on stdin, and expects a single flat JSON object of strings on stdout, exposed as `result`. `http` performs a request and exposes `response_body`, `status_code` and `response_headers`. The risks are what make them a last resort: the read runs on every plan, so every CI runner needs the interpreter, the binary, the network path and the credentials; a transient failure breaks `plan`, not just `apply`; and the result is written to state in plaintext, so anything sensitive leaks. They must also be side-effect free — they are reads, not a way to run a step. Where possible I prefer having the pipeline compute the value and pass it in as a variable.

code

hcl · 10 lines
hcl
data "http" "release_meta" {
  url = "https://internal.example.com/releases/current"
  request_headers = {
    Accept = "application/json"
  }
}

locals {
  release = jsondecode(data.http.release_meta.response_body)["version"]
}

go deeper

for a junior

Know that these two data sources exist for facts no provider exposes: external runs a program and reads a flat JSON string map from stdout, http fetches a URL. Both are reads and neither creates anything.

for a middle

State the external contract exactly — query as JSON on stdin, one flat JSON object of strings on stdout, stderr on failure — and explain that reads run on every plan, so the tooling and network they need become a plan-time dependency.

for a senior

Argue the tradeoff: plan-time failure surface, results recorded in state in plaintext, and the hard requirement that the program be side-effect free because Terraform makes no promise about read counts. Offer pipeline-supplied variables as the first alternative.

for a principal

Decide the estate-wide policy on escape hatches: where a real provider or a published parameter is worth the investment, and which repositories may take a plan-time dependency on local tooling at all. Every such block is a shared constraint on every runner that plans that code.

## What they are Most data sources are typed reads offered by a provider that knows the API it is talking to. Two generic ones exist for when nothing fits. **`external`** (hashicorp/external provider) invokes a program: ```hcl data "external" "account_meta" { program = ["python3", "${path.module}/lookup.py"] query = { environment = var.environment } } # data.external.account_meta.result["cost_centre"] ``` The contract is narrow and strict. Terraform passes `query` to the program as a JSON object on stdin. The program must print **one flat JSON object whose values are all strings** on stdout and exit zero. Nested objects, numbers and booleans are not part of the contract. On a non-zero exit, Terraform surfaces stderr as the error message, which is the only diagnostic channel the program has. **`http`** (hashicorp/http provider) performs an HTTP request and exposes `response_body`, `status_code` and `response_headers`; `request_headers` lets you send headers. It does not fail on a 404 — a non-2xx status is a value you have to check, not an automatic error. ## When they are justified - The fact lives in a system with no Terraform provider: an internal service registry, a homegrown IPAM, a CMDB. - A provider exists but lacks the specific read, and you need it now rather than after upstreaming a patch. - You are computing something purely local and deterministic from inputs that HCL functions cannot express. That list is short on purpose. If the value comes from a system that *does* have a provider, use its data source; if it comes from your own pipeline, pass it in as a variable. ## The risks **They run on every plan.** Not on apply, not once — every single plan, including the ones a reviewer opens and abandons. A `plan` in CI now depends on the runner having Python (or whatever interpreter), the script being present at `path.module`, outbound network access, and credentials for whatever the script calls. Every one of those is a new way for `terraform plan` to fail on a change that has nothing to do with it. Container-based runners and developer laptops are different environments; a script that works locally and not in CI is a classic. **Failures are plan-time failures.** A flaky endpoint no longer risks an apply — it blocks review entirely. There is no retry policy and no timeout you control on `external`. **Results are recorded in state.** Whatever the program printed, or the endpoint returned, is stored in the state file as the read result. If it is a token, a password or personal data, it is now in state in plaintext and inherits whatever protection the backend gives it. That is a strong reason never to use `http` to fetch a secret. **They must be side-effect free.** Because they run on every plan, a program that mutates something will mutate it on plans that are never applied, on plans run concurrently by different engineers, and on the plan a reviewer generated to look at a diff. Terraform makes no promise about how many times it will read. `external` is not a mechanism for running steps — that is what provisioners and `terraform_data` exist for, with their own tradeoffs. **Type poverty.** `result` is `map(string)`, so everything comes back as a string and must be converted. `http` gives you a body you have to parse yourself, typically with `jsondecode`, with no schema and no validation. **Unknown propagation.** If `query` depends on a value created in the same run, the read defers to apply and everything downstream becomes unknown in the plan — the same cascade any deferred data source causes, but now with an opaque program at the centre of it. ## Better alternatives 1. **Compute it in the pipeline** and pass it as a `TF_VAR_`-style input. The value is then visible in the run's inputs, versioned in the pipeline definition, and the plan does not depend on a script. 2. **Use a provider that already reads the store** — for example a parameter or secret store data source, which at least gives you typed access, provider-level auth and no local interpreter dependency. 3. **Write a real provider** if the lookup is permanent and used across many configurations. It is more work once, and it removes the runner dependency for everyone. ## How to answer this in an interview Show that you know the contract precisely, then show that you treat it as a last resort. The strong signal is not that you can use `external`; it is that you can say what it costs and what you would try first.

  • What exactly must a program used by the `external` data source write to stdout?
    A single JSON object whose values are all strings, and nothing else — no logging, no banner text. It reads the `query` map as JSON on stdin, and on a non-zero exit Terraform reports stderr as the error. Anything richer than a flat string map has to be encoded and decoded by hand.
  • Why is it dangerous for that program to have side effects?
    Because reads happen on every plan, including plans that are abandoned, run concurrently, or generated by a reviewer just to look at the diff. Terraform gives no guarantee about how often it reads, so a mutating program executes an uncontrolled number of times outside any apply.
  • You need a value from an internal service in ten different configurations. What would you do instead of ten `external` blocks?
    Publish the value into something with a real provider — a parameter store entry or a well-known object — written by the service's own pipeline, and read it with a typed data source. If the lookup is permanent and non-trivial, writing an actual provider removes the runner dependency for every consumer at once.

saying these in an interview costs you the question

  • Uses `external` to run a setup step during apply
  • Expects the program to return nested JSON or numbers
  • Assumes the read happens once and is then cached
  • Fetches a secret over `http` and forgets it lands in state
  • Thinks a non-2xx status from `http` fails the plan automatically

context