HashiCorp's documentation calls Terraform provisioners a last resort. What are the concrete reasons for that advice, and what should you reach for instead?
answer
- outside the state/diff loop
- invisible to plan reviewers
- runs on create only, never again
- needs a route and a key from CI
- bake it, cloud-init it, or hand it off
basics
~20 sProvisioners are opaque to Terraform: their effects are not in state, not visible in the plan, and not idempotent, and they need network reachability plus credentials from wherever Terraform runs. Prefer a baked image, cloud-init/user_data, or a config-management handoff.
solid answer
~50 sTerraform's model is that the provider knows the current state of a resource and can compute a diff. A provisioner breaks all of that. Whatever the script does is invisible to Terraform — it is not recorded in state, so nothing detects drift if someone undoes it, and it never shows up in `terraform plan`, so a reviewer approving the plan cannot see what will actually execute. Provisioners are not idempotent either: they run only on create (or destroy), so an edit to the script changes nothing on already-existing resources. And `remote-exec` adds an operational dependency — the machine running Terraform now needs a route to the instance, an open port, and SSH credentials, which is painful from a CI runner. Preferred alternatives, in order: bake the software into the image; pass a cloud-init or `user_data` script the platform runs itself; or hand the machine to a config-management tool such as Ansible after apply. Provisioners remain reasonable for genuinely one-shot local glue.
code
hcl · 20 linesdata "aws_ami" "base" {
most_recent = true
owners = ["amazon"]
filter {
name = "name"
values = ["al2023-ami-*-x86_64"]
}
}
resource "aws_instance" "web" {
ami = data.aws_ami.base.id
instance_type = "t3.micro"
user_data = templatefile("${path.module}/init.tftpl", {
app_port = 8080
})
user_data_replace_on_change = true
}go deeper
Know the headline: provisioners are a documented last resort, and the usual replacement is passing a startup script through user_data or cloud-init instead of SSHing in.
Explain the mechanics behind the advice — no state record, no plan diff, no re-run on change — and describe how moving the script into a resource attribute restores all three.
Demonstrate the operational cost: credentials and network paths opened purely for the tool, and a transient bootstrap failure tainting an instance into replacement. Have a concrete migration path ready.
Own the standard for where machine configuration lives across the estate — image build versus boot-time script versus a separate convergence tool — and the review guarantee that any shell-out inside the graph quietly breaks.
## The model provisioners break Terraform's whole value proposition is a three-way comparison: what the configuration says, what the state file recorded, and what the provider reports the real resource looks like. From those it computes a plan, you approve it, and it applies. A provisioner sits outside that loop entirely. It is an imperative shell-out bolted onto a resource's create step, and Terraform has no idea what it did. That single fact generates all five concrete objections. ## 1. Nothing lands in state If your `remote-exec` installs nginx, Terraform records no attribute saying "nginx is installed". The state file only knows the instance exists. If somebody SSHes in and uninstalls nginx, the next `terraform plan` says `No changes` — Terraform cannot detect drift in something it never tracked. ## 2. No plan visibility The plan shows `aws_instance.web will be created` plus its attributes. It does not show the commands the provisioner will run. A reviewer approving a plan in a pull request is approving a script they cannot see the effect of. For teams whose control story is "every prod change is a reviewed plan", that is a real hole. ## 3. Not idempotent, and only run on create Provisioners run when the resource is created and, for `when = destroy`, when it is destroyed. They do not run on subsequent applies. So changing the script body is a no-op for every existing instance — the fleet silently splits into machines built with the old script and machines built with the new one. Worse, the script itself usually isn't idempotent (`apt-get install` twice is fine; `echo >> file` twice is not), so any workaround that does force a re-run can corrupt the box. ## 4. Connectivity and credentials `remote-exec` and `file` need the Terraform host to reach the resource. On a laptop with a public-IP instance that is easy. In CI it is often impossible: the runner sits in a different network, the instance is in a private subnet, and the SSH private key would have to be handed to the pipeline. You end up widening a security group or shipping a key just so a provisioner can run — infrastructure changed to accommodate the tool, which is backwards. ## 5. Failure semantics are harsh A failed creation-time provisioner marks the resource **tainted**: the object exists in the cloud, but Terraform records it as unusable and plans to destroy and recreate it on the next apply. A transient network blip during bootstrap therefore costs you a full instance replacement. ## The sanctioned alternatives **Bake the image.** Build an AMI or machine image ahead of time with an image builder such as Packer, then have Terraform reference the finished image. Boot time drops, the content is versioned and testable, and every instance is identical by construction. This is the immutable-infrastructure answer. **Use the platform's own bootstrap channel.** Almost every compute service accepts a startup script — on EC2 that is `user_data`, consumed by cloud-init. The script becomes an *attribute of the resource*: it is in the plan diff, in state, and rendering it with `templatefile` keeps it in the repo. The AWS provider even exposes `user_data_replace_on_change`, so editing the script can force a replacement instead of silently diverging. ```hcl resource "aws_instance" "web" { ami = data.aws_ami.base.id instance_type = "t3.micro" user_data = templatefile("${path.module}/init.tftpl", { port = 8080 }) user_data_replace_on_change = true } ``` **Hand off to config management.** Let Terraform create infrastructure and stop there; a separate config-management run (Ansible, for example) converges the machine's software state afterwards, on its own schedule and with its own idempotency guarantees. Terraform's job ends at the resource boundary. **Use a provider resource instead of a shell command.** A surprising share of `local-exec` calls to a CLI have a real resource or data source available. If a provider models the thing, use it — you get state, diff and drift detection for free. ## What provisioners are still fine for HashiCorp's own carve-out is glue that genuinely has no declarative equivalent and no consequence if it drifts: writing a local file for another tool, triggering an external process, running a one-off command on a resource the platform gives you no other way to configure. `local-exec` on `terraform_data` for a single build step is a legitimate use. The rule of thumb interviewers want to hear: if the provisioner is configuring the *inside* of a machine, it is the wrong tool; if it is gluing Terraform to something outside its world, it may be the only tool.
- If a provisioner's effects are invisible to Terraform, why not just re-run it on every apply?Terraform has no way to know whether the work is already done, so re-running would demand that every script be perfectly idempotent — a guarantee shell scripts rarely offer. It would also make every apply a mutation of running machines, which is exactly what the immutable-infrastructure model exists to avoid. Terraform therefore ties provisioners to create and destroy only.
- You inherit a repo whose instances are bootstrapped by remote-exec. What is your migration path?Move the script body into `user_data` rendered with `templatefile` so the platform's cloud-init runs it, verify a single instance boots correctly, then delete the provisioner and connection blocks. Because the script now lives in a resource attribute, changing it produces a visible plan diff; set `user_data_replace_on_change` if edits must roll the fleet. Longer term, bake the stable parts into the image.
- Is there a case where local-exec is genuinely the right answer?Yes — one-shot glue between Terraform and something it does not model: writing a generated file for another tool, invoking a CLI with no corresponding provider resource, or triggering an external job after a deploy. The test is whether the command configures the inside of a managed resource (wrong tool) or connects Terraform to the outside world (acceptable).
saying these in an interview costs you the question
- Claiming provisioners re-run to converge configuration
- Saying provisioner output is stored in state
- Treating remote-exec as equivalent to user_data
- Believing the plan shows what a provisioner will do
- Widening a security group so CI can SSH in