skip to content

During terraform apply an instance is created successfully but its remote-exec provisioner fails halfway through. What state is that resource left in, and what does the next apply do?

level: middleimportance: must knowfreq 55%

answer

  1. the object exists, Terraform distrusts it
  2. a flag in state, not a deletion
  3. next plan shows -/+ replacement
  4. on_failure has exactly two values
  5. destroy-time failures behave differently

basics

~20 s

The apply errors and Terraform marks the resource tainted in state: the real object still exists, but Terraform treats it as unusable. The next plan schedules it for destroy and recreate, so the whole provisioner sequence runs again on a fresh object.

solid answer

~50 s

Terraform's assumption is that a creation-time provisioner is part of creating the resource, so a failed provisioner means the resource is not correctly built. The object was really created — you are billed for it — but Terraform flags it as tainted in state and the apply exits non-zero. The next `terraform plan` shows it as `must be replaced` with the reason `tainted`, and the next apply destroys and recreates it, re-running `file` and `remote-exec` from scratch. That is usually what you want, because the half-configured machine is not trustworthy. If you would rather keep going, set `on_failure = continue` on the provisioner block; the default is `on_failure = fail`. The dangerous middle case is a partially-executed script that left the box in a state nobody reasoned about — which is one of the reasons bootstrapping via `user_data` or a baked image beats provisioners.

code

hcl · 20 lines
hcl
resource "aws_instance" "web" {
  ami           = var.ami_id
  instance_type = "t3.micro"

  connection {
    type        = "ssh"
    user        = "ubuntu"
    private_key = var.ssh_key
    host        = self.public_ip
  }

  provisioner "remote-exec" {
    inline = ["sudo /opt/bootstrap.sh"]
  }

  provisioner "local-exec" {
    command    = "./notify-slack.sh ${self.id}"
    on_failure = continue
  }
}

go deeper

for a junior

Recall the word tainted and what it implies: the instance really exists, but Terraform will destroy and recreate it on the next apply rather than reuse it.

for a middle

Explain the ordering — state is written after creation and before provisioning — and describe on_failure with its fail and continue values and what each does to the taint flag.

for a senior

Weigh the operational cost: a transient bootstrap error escalating into an instance replacement, and the risk of a half-executed script that nobody can characterise. Say what you would change to avoid it.

for a principal

Own the policy on suppressing failures — when untaint or on_failure = continue is ever acceptable — and push the bootstrap out of the apply path so a package-mirror blip cannot become an infrastructure change.

## What "tainted" means Tainted is a per-instance flag recorded in Terraform state. It says: this object exists in the real world, but Terraform believes it is not in a usable condition, so it should be destroyed and recreated rather than reused. It is not a separate object and not a deletion — the cloud resource is still running and still costing money. A creation-time provisioner failure is the *automatic* way a resource becomes tainted. Terraform's reasoning is straightforward: you declared the provisioner as part of creating this resource, so if it did not complete, the resource was not fully created. ## The apply sequence 1. The provider creates the object and returns its attributes; Terraform writes them to state immediately. 2. Terraform runs the resource's provisioners, in declaration order. 3. A provisioner exits non-zero. Terraform records the error, marks the instance tainted, writes state, and fails the apply with a non-zero exit code. Step 1 matters: state is written *before* provisioning, which is why the object is not orphaned. If Terraform had not recorded it, you would be left with an untracked instance running in the account. ## What the next plan shows ``` # aws_instance.web is tainted, so must be replaced -/+ resource "aws_instance" "web" { ``` The `-/+` marker means destroy-then-create. On apply, the tainted object is destroyed, a fresh one is created, and the full provisioner sequence runs again from the top. There is no partial resume — provisioners have no checkpointing. ## on_failure Every provisioner block accepts `on_failure`, with two values: - `fail` (the default) — error out and taint the resource. - `continue` — log the error, ignore it, and carry on with the rest of the apply. The resource is **not** tainted. ```hcl provisioner "remote-exec" { inline = ["/opt/optional-metrics-agent.sh"] on_failure = continue } ``` Note it is a bare keyword, not a quoted string. Use `continue` only for genuinely optional side work; using it to paper over a flaky bootstrap gives you machines that report success while being half-configured — strictly worse than a loud failure. ## Destroy-time failures behave differently A `when = destroy` provisioner that fails does not taint anything — instead Terraform errors and the resource stays in state, and the provisioner will be attempted again on the next apply that tries to destroy it. That is why the docs warn destroy-time provisioners must be safe to run more than once. ## Clearing or forcing the flag manually Since Terraform 0.15.2 the recommended way to force replacement is the plan option `-replace=ADDRESS` rather than the older `terraform taint` command, which is deprecated. To go the other way and tell Terraform the object is actually fine, `terraform untaint ADDRESS` clears the flag — appropriate only when you have verified by hand that the machine is correctly configured, for example because the provisioner failed on its very last, harmless step. ## Why this behaviour argues against provisioners Consider the operational picture. A transient DNS hiccup or a package mirror timeout during bootstrap now costs a full instance replacement — and if that instance is in an autoscaling-adjacent or stateful position, the blast radius is real. Compare with `user_data`: if a cloud-init script fails, the instance still exists and boots, Terraform's apply succeeds, and you debug the box by reading `/var/log/cloud-init-output.log`. The failure is contained to the machine rather than escalating into an infrastructure change. The deeper problem is the half-executed script. Terraform knows only that the command exited non-zero; it has no idea which of the eight lines in your `inline` list ran. If someone reaches for `on_failure = continue` or `terraform untaint` to get the pipeline green, the fleet quietly acquires a machine nobody can reason about. Destroy-and-recreate is the safe default precisely because it refuses to reason about partial state. ## What to say in an interview Name the flag (tainted), state that the real object still exists and is billed, describe the destroy/recreate plan, mention `on_failure` and its two values, and close with the judgment: the harshness of this failure mode is one of the strongest practical arguments for putting bootstrapping in the image or in `user_data` instead.

  • Why does Terraform write the resource into state before running its provisioners?
    So a provisioner failure cannot orphan the object. The provider has already created a real, billable resource; if Terraform discarded that record on error, the instance would keep running with nothing tracking it. Writing state first means the failed apply leaves a tainted-but-tracked object that the next apply can destroy and recreate cleanly.
  • When is it legitimate to run terraform untaint instead of letting the replacement happen?
    Only when you have verified by hand that the object is genuinely in the desired state — for example the provisioner failed on a final, cosmetic step, or on a log-shipping command with no bearing on the workload. Untainting to make a pipeline green without that verification leaves an unreasoned machine in the fleet.
  • Does on_failure = continue leave the resource tainted?
    No. `continue` tells Terraform to log the provisioner error and proceed; the resource is treated as successfully created and is not marked tainted, so no replacement is planned. That is exactly why it should be reserved for genuinely optional side effects rather than used to suppress a flaky bootstrap.

saying these in an interview costs you the question

  • Thinking the failed apply deletes the created instance
  • Believing the next apply resumes the script mid-way
  • Saying tainted means the resource is missing from state
  • Using on_failure = continue to hide a flaky bootstrap
  • Assuming a destroy-time provisioner failure also taints

context