When would you use Terraform's `external` or `http` data sources, and what are the risks of shipping one in a production configuration?
answer
- last-resort generic reads
- stdin JSON in, string map out
- runs on every plan, everywhere
- result lands in state in plaintext
- reads must have no side effects
basics
~20 sThey are escape hatches for facts no provider exposes: external runs a program and reads a JSON object of strings from its stdout, http performs a request and exposes the response. Both run on every plan, on every machine that plans, and their results land in state.
solid answer
~50 sI reach for them only when there is genuinely no provider data source for the fact I need — an internal service that owns an id, a registry that has no Terraform provider. `external` from the hashicorp/external provider runs a program, passes the `query` map to it as JSON on stdin, and expects a single flat JSON object of strings on stdout, exposed as `result`. `http` performs a request and exposes `response_body`, `status_code` and `response_headers`. The risks are what make them a last resort: the read runs on every plan, so every CI runner needs the interpreter, the binary, the network path and the credentials; a transient failure breaks `plan`, not just `apply`; and the result is written to state in plaintext, so anything sensitive leaks. They must also be side-effect free — they are reads, not a way to run a step. Where possible I prefer having the pipeline compute the value and pass it in as a variable.
code
hcl · 10 linesdata "http" "release_meta" {
url = "https://internal.example.com/releases/current"
request_headers = {
Accept = "application/json"
}
}
locals {
release = jsondecode(data.http.release_meta.response_body)["version"]
}go deeper
Know that these two data sources exist for facts no provider exposes: external runs a program and reads a flat JSON string map from stdout, http fetches a URL. Both are reads and neither creates anything.
State the external contract exactly — query as JSON on stdin, one flat JSON object of strings on stdout, stderr on failure — and explain that reads run on every plan, so the tooling and network they need become a plan-time dependency.
Argue the tradeoff: plan-time failure surface, results recorded in state in plaintext, and the hard requirement that the program be side-effect free because Terraform makes no promise about read counts. Offer pipeline-supplied variables as the first alternative.
Decide the estate-wide policy on escape hatches: where a real provider or a published parameter is worth the investment, and which repositories may take a plan-time dependency on local tooling at all. Every such block is a shared constraint on every runner that plans that code.
## What they are Most data sources are typed reads offered by a provider that knows the API it is talking to. Two generic ones exist for when nothing fits. **`external`** (hashicorp/external provider) invokes a program: ```hcl data "external" "account_meta" { program = ["python3", "${path.module}/lookup.py"] query = { environment = var.environment } } # data.external.account_meta.result["cost_centre"] ``` The contract is narrow and strict. Terraform passes `query` to the program as a JSON object on stdin. The program must print **one flat JSON object whose values are all strings** on stdout and exit zero. Nested objects, numbers and booleans are not part of the contract. On a non-zero exit, Terraform surfaces stderr as the error message, which is the only diagnostic channel the program has. **`http`** (hashicorp/http provider) performs an HTTP request and exposes `response_body`, `status_code` and `response_headers`; `request_headers` lets you send headers. It does not fail on a 404 — a non-2xx status is a value you have to check, not an automatic error. ## When they are justified - The fact lives in a system with no Terraform provider: an internal service registry, a homegrown IPAM, a CMDB. - A provider exists but lacks the specific read, and you need it now rather than after upstreaming a patch. - You are computing something purely local and deterministic from inputs that HCL functions cannot express. That list is short on purpose. If the value comes from a system that *does* have a provider, use its data source; if it comes from your own pipeline, pass it in as a variable. ## The risks **They run on every plan.** Not on apply, not once — every single plan, including the ones a reviewer opens and abandons. A `plan` in CI now depends on the runner having Python (or whatever interpreter), the script being present at `path.module`, outbound network access, and credentials for whatever the script calls. Every one of those is a new way for `terraform plan` to fail on a change that has nothing to do with it. Container-based runners and developer laptops are different environments; a script that works locally and not in CI is a classic. **Failures are plan-time failures.** A flaky endpoint no longer risks an apply — it blocks review entirely. There is no retry policy and no timeout you control on `external`. **Results are recorded in state.** Whatever the program printed, or the endpoint returned, is stored in the state file as the read result. If it is a token, a password or personal data, it is now in state in plaintext and inherits whatever protection the backend gives it. That is a strong reason never to use `http` to fetch a secret. **They must be side-effect free.** Because they run on every plan, a program that mutates something will mutate it on plans that are never applied, on plans run concurrently by different engineers, and on the plan a reviewer generated to look at a diff. Terraform makes no promise about how many times it will read. `external` is not a mechanism for running steps — that is what provisioners and `terraform_data` exist for, with their own tradeoffs. **Type poverty.** `result` is `map(string)`, so everything comes back as a string and must be converted. `http` gives you a body you have to parse yourself, typically with `jsondecode`, with no schema and no validation. **Unknown propagation.** If `query` depends on a value created in the same run, the read defers to apply and everything downstream becomes unknown in the plan — the same cascade any deferred data source causes, but now with an opaque program at the centre of it. ## Better alternatives 1. **Compute it in the pipeline** and pass it as a `TF_VAR_`-style input. The value is then visible in the run's inputs, versioned in the pipeline definition, and the plan does not depend on a script. 2. **Use a provider that already reads the store** — for example a parameter or secret store data source, which at least gives you typed access, provider-level auth and no local interpreter dependency. 3. **Write a real provider** if the lookup is permanent and used across many configurations. It is more work once, and it removes the runner dependency for everyone. ## How to answer this in an interview Show that you know the contract precisely, then show that you treat it as a last resort. The strong signal is not that you can use `external`; it is that you can say what it costs and what you would try first.
- What exactly must a program used by the `external` data source write to stdout?A single JSON object whose values are all strings, and nothing else — no logging, no banner text. It reads the `query` map as JSON on stdin, and on a non-zero exit Terraform reports stderr as the error. Anything richer than a flat string map has to be encoded and decoded by hand.
- Why is it dangerous for that program to have side effects?Because reads happen on every plan, including plans that are abandoned, run concurrently, or generated by a reviewer just to look at the diff. Terraform gives no guarantee about how often it reads, so a mutating program executes an uncontrolled number of times outside any apply.
- You need a value from an internal service in ten different configurations. What would you do instead of ten `external` blocks?Publish the value into something with a real provider — a parameter store entry or a well-known object — written by the service's own pipeline, and read it with a typed data source. If the lookup is permanent and non-trivial, writing an actual provider removes the runner dependency for every consumer at once.
saying these in an interview costs you the question
- Uses `external` to run a setup step during apply
- Expects the program to return nested JSON or numbers
- Assumes the read happens once and is then cached
- Fetches a secret over `http` and forgets it lands in state
- Thinks a non-2xx status from `http` fails the plan automatically