skip to content

A Terraform config with a remote-exec provisioner works from a developer's laptop but times out when the same code runs on a CI runner. What is going wrong, and how would you fix it properly?

level: seniorimportance: should knowfreq 42%

answer

  1. the session starts on the Terraform host
  2. laptop and runner sit in different networks
  3. route, then port, then credential
  4. bastion_host is a patch, not a fix
  5. let the machine configure itself

basics

~20 s

remote-exec connects from wherever Terraform runs, so the CI runner must have a network route to the host, an open port, and valid credentials. The laptop had all three; the runner usually does not. The proper fix is to stop provisioning over the network.

solid answer

~50 s

The `connection` block is evaluated on the Terraform host, not in the cloud. On a laptop, the developer is often on a VPN or hitting `self.public_ip` through a security group that allows their office address, with the SSH key sitting in `~/.ssh`. The CI runner has a different source address, no key, and frequently no route at all if the instance is in a private subnet. Diagnose it in that order: route, port, credential. You can patch it — `bastion_host` and its companion arguments to hop through a jump box, a security-group rule for the runner's egress range, and the private key injected as a pipeline secret — but every one of those widens exposure to satisfy a tool. The correct answer is to remove the dependency: move the script into `user_data` so the platform runs it at boot, bake it into the image, or let an agent-based config-management run converge the host afterwards. Then Terraform never needs to reach the machine at all.

code

hcl · 12 lines
hcl
resource "aws_instance" "web" {
  ami                    = var.ami_id
  instance_type          = "t3.micro"
  subnet_id              = var.private_subnet_id
  vpc_security_group_ids = [var.app_sg_id]

  user_data = templatefile("${path.module}/bootstrap.sh.tftpl", {
    app_port = 8080
  })

  user_data_replace_on_change = true
}

go deeper

for a junior

Know that remote-exec connects from the machine running Terraform, so a CI runner needs its own network path and its own copy of the credentials — the developer's ~/.ssh key is not there.

for a middle

Walk the checks in order: does host resolve to a reachable address, is the port open through the security group, is the user present on the image, and is sshd up yet.

for a senior

Show that each workaround — bastion, wider ingress, a key in the pipeline — buys success with exposure, and propose eliminating the inbound dependency by bootstrapping at boot instead.

for a principal

Own the rule that pipelines get no standing inbound path or long-lived host key into workload networks, and make boot-time or pull-based configuration the estate default so the question stops arising.

## Where the connection is made from The most common misconception here is that Terraform somehow asks the cloud to run the commands. It does not. `remote-exec` and `file` open an SSH or WinRM session **from the process running `terraform apply`** to the address in `connection.host`. Everything follows from that. Terraform Cloud/Enterprise workers, a container in a build cluster, and a laptop on the corporate VPN are three completely different network positions, and the same configuration behaves differently in each. ## The three preconditions, in diagnostic order **1. A route to the host.** If `host = self.public_ip` and the instance has no public IP (private subnet, no `associate_public_ip_address`), the attribute is empty and the connection goes nowhere. If it is `self.private_ip`, the runner needs to be inside the VPC — via VPN, Direct Connect, or peering — which a hosted runner is not. **2. An open path on the port.** Port 22 (or 5985/5986 for WinRM) must be reachable: security group ingress, network ACLs, and any host firewall in the image. Laptop success often means the security group allows the office CIDR, and the runner's egress address is not in it. **3. Credentials valid on the booted image.** `private_key = file("~/.ssh/id_rsa")` reads a file that exists on the developer's machine and does not exist on the runner. `agent = true` relies on an SSH agent the runner has not started. And the `user` must already exist in the image — a freshly booted AMI has one default account, not your developer usernames. ## The timing trap on top of those Even when all three hold, the resource is often "created" from the provider's point of view before the operating system has finished booting and sshd is listening. Terraform retries the connection until `connection.timeout` elapses (default five minutes), so this shows up as an intermittent failure that is worse on cold, larger images — and, per the failure semantics of provisioners, a timeout taints the instance and forces a replacement. ## The patches, and what they cost ```hcl connection { type = "ssh" user = "ubuntu" private_key = var.ssh_private_key host = self.private_ip bastion_host = var.bastion_public_ip bastion_user = "ec2-user" bastion_private_key = var.bastion_key timeout = "10m" } ``` That works. It also means: a jump box exists and is exposed; the pipeline holds a long-lived SSH private key as a secret; and a security group now permits the runner's address range. You have made three security concessions so that one imperative step can run inside the resource graph. An interviewer asking this question wants to hear you notice that. ## The proper fixes **Move the script to `user_data`.** The platform's own metadata service delivers it and cloud-init executes it locally at first boot. No inbound path, no key, no timing race. It also becomes a resource attribute, so a change shows in the plan diff — and on AWS you can set `user_data_replace_on_change = true` to make edits roll the fleet. **Bake it into the image.** Build the image ahead of time and reference it. Boot is faster and every instance is identical by construction. **Invert the direction.** Instead of Terraform pushing to the host, let the host pull: an agent, an instance-profile-authenticated fetch, or a separate config-management run after apply. Outbound connections from the instance need no inbound rules and no shared private key. **Use an agentless remote-command service where the platform offers one.** Where a managed remote-command channel exists, it authenticates with the same cloud credentials the pipeline already has and needs no open SSH port — the general shape being "reach the machine through the control plane, not through the network". ## How to answer Structure it as diagnosis then design. Diagnosis: connection originates on the Terraform host; check route, then port, then credential, then the sshd-readiness race. Design: every patch that makes it work in CI widens the attack surface, so the durable fix is to eliminate the inbound dependency entirely by having the machine configure itself at boot. Closing with "and this is exactly why HashiCorp calls provisioners a last resort" ties it back to the principle rather than leaving it as a networking anecdote.

  • The connection sometimes succeeds and sometimes times out on identical infrastructure. What causes that?
    A race between the provider reporting the resource created and sshd actually accepting connections. Terraform retries until `connection.timeout` (five minutes by default) expires, so slower boots — large images, cold starts, contended hosts — tip it over. Raising the timeout masks it; moving the work to boot-time `user_data` removes the race entirely.
  • What do bastion_host and its companion arguments actually do in a connection block?
    They tell Terraform to open an SSH session to the bastion first and tunnel the real connection to `host` through it, using `bastion_user`, `bastion_private_key` and `bastion_port`. It solves reachability into a private subnet, but at the cost of an exposed jump box plus a second key in the pipeline's secret store.
  • Why is pulling configuration from the instance safer than Terraform pushing it?
    An outbound fetch needs no inbound firewall rule and no shared private key: the instance authenticates with its own attached identity, which is short-lived and scoped. Pushing requires an open port, a route from the runner, and a long-lived credential stored in CI — three standing exposures that exist only so a build step can reach the box.

saying these in an interview costs you the question

  • Thinking the cloud provider executes remote-exec for you
  • Blaming the AMI when the real problem is a missing route
  • Opening SSH to 0.0.0.0/0 to make CI work
  • Assuming the key in ~/.ssh exists on the runner
  • Treating a raised timeout as the fix for the boot race

context