skip to content

When would you run Terraform through a purpose-built run service such as Atlantis or HCP Terraform instead of hand-written plan and apply jobs in your existing CI system?

level: principalimportance: nice to knowfreq 34%

answer

  1. the pull request becomes the control surface
  2. who holds the cloud credentials
  3. duplication drifts across many repos
  4. audit history as a product feature
  5. private networks need agents

basics

~20 s

Adopt a run service when the number of Terraform repositories makes hand-rolled jobs a maintenance burden, and when you want cloud credentials, state, run history and approvals owned by one system instead of scattered across CI configurations.

solid answer

~50 s

Hand-rolled jobs are the right answer while you have a handful of root modules: they are transparent, use the CI system everyone already knows, and cost nothing extra. The case for a run service appears at scale. Atlantis and HCP Terraform both make the pull request the control surface — the plan is posted automatically, and an approval or a comment triggers the apply — and both hold the credentials and serialise runs per project or workspace themselves, so your CI system no longer needs write access to the cloud at all. HCP Terraform additionally hosts state and run history, giving you one audit log and one place to attach policy checks. The costs are real: another system to operate or pay for, a queue that becomes a bottleneck when it stalls, and network reachability for private infrastructure, which means running agents inside your network. I would move when duplicated pipeline YAML across dozens of repositories starts drifting, not before.

go deeper

for a junior

Know that Terraform can be driven by a dedicated service that plans on the pull request and applies on a comment or approval, rather than by scripts in the CI system.

for a middle

Be able to describe what each tool does concretely: comment-driven runs and project locks in Atlantis, hosted state, speculative plans and queued workspace runs in HCP Terraform.

for a senior

Argue the operational tradeoffs — who holds credentials, what breaks when the service is down, how agents reach private networks, and what bespoke pipeline steps become awkward.

for a principal

Own the adoption decision and its timing: the duplication and audit thresholds that justify the move, the vendor coupling you accept, the cost model as the estate grows, and the migration path if you change your mind.

## The three shapes **Hand-rolled CI jobs.** Your existing CI system checks out the repo, installs Terraform, and runs plan and apply. Credentials are federated to the job. The pipeline definition lives in the repository, next to the code it applies. **Atlantis.** A self-hosted service that listens to pull-request webhooks. On a PR it runs plan for the affected projects and posts the output back as a comment; engineers drive it by commenting `atlantis plan` and `atlantis apply` on the PR, and it holds its own lock on a project and workspace until the PR is merged or closed, so a second PR touching the same project is told which PR is blocking it. It executes Terraform itself, so your CI system never touches cloud credentials. **HCP Terraform** (the hosted service formerly called Terraform Cloud). It goes further: it hosts the state, connects directly to the VCS repository, produces a speculative plan on every PR, queues runs per workspace, holds the variables and credentials as workspace configuration, and keeps a durable run history with the plan, the approver and the result attached. Applies wait for confirmation in its own UI or API. Where infrastructure is not internet-reachable, agents run inside your network and poll for work. ## What you are actually buying 1. **A credential boundary.** In the hand-rolled shape, the CI system is the thing with the power to change production, so its supply chain — every action, every plugin, every fork trigger — is in scope for that power. A run service moves the credentials out, leaving CI with only the ability to ask for a run. 2. **Uniformity without copy-paste.** Forty repositories with hand-written pipelines become forty slightly different pipelines. Their drift is invisible until the one that skipped the validation step ships something bad. A run service applies the same lifecycle everywhere by construction. 3. **Serialisation and locking as a product feature.** Both tools queue per project or workspace out of the box, which removes the concurrency design work you would otherwise do per pipeline. 4. **Audit and evidence.** "Show me every production change last quarter, who approved it, and the diff" is a query against one run history rather than an archaeology exercise across CI logs with expired retention. 5. **A hook for guardrails.** Policy evaluation attaches to the run itself rather than being re-wired into each pipeline. ## What you pay - **An operational dependency.** Atlantis is a service you run, upgrade, secure and make highly available; when it is down, nobody ships infrastructure. HCP Terraform is somebody else's uptime and a per-resource bill that scales with the size of your estate. - **Reduced flexibility.** Bespoke steps — a pre-plan schema migration, a custom cost check, an unusual credential broker — are trivial in a general-purpose pipeline and awkward inside an opinionated run lifecycle. - **Network placement.** A hosted service cannot reach a private VPC endpoint. That means agents inside the network, which is another fleet to run, and one that holds the credentials. - **Coupling to the tool's model.** Workspace-per-environment, VCS-driven runs and the service's variable handling become how your team thinks. Migrating away later is a project. ## The decision I would use hand-rolled jobs while the estate is small, one team owns it, and the pipeline definition fits on a screen — the transparency is worth more than the features. I would move to a run service when at least two of these are true: pipeline YAML is duplicated across many repositories and has begun to drift; the audit question is asked by someone outside the team and the current answer is unconvincing; the CI system's blast radius has become the security concern rather than the cloud permissions themselves; or engineers outside the platform team need to make infrastructure changes and need a guided path rather than a set of conventions. Between the two, the axis is control versus completeness. Atlantis keeps state and backend choices yours and costs only the effort of running it, which suits an organisation that already has strong platform capability and dislikes vendor coupling. HCP Terraform takes over state and run history too, which is exactly what an organisation without that capability wants — and exactly the coupling the first one is avoiding. Whichever you choose, the pipeline shape underneath does not change: plan is a preview a human reads, apply is bound to reviewed code, credentials are short-lived, and runs against one state are serialised. The service automates that shape; it does not replace it.

  • What specifically does a run service remove from your CI system's blast radius?
    The ability to change infrastructure. In a hand-rolled pipeline the CI system holds credentials that can write production, so every third-party action, plugin and trigger path is a route to those credentials. With a run service, CI can at most request a run; the credentials live with the service, whose surface is far narrower and whose access is auditable in one place.
  • Atlantis locks a project until the pull request is merged or closed. Why is that different from a state lock?
    A state lock lives for the seconds or minutes of a single operation and exists to stop concurrent writers. Atlantis's lock is a workflow lock over the whole review cycle: it stops a second pull request from planning against a project whose change is still under review, so two people cannot both be looking at plans that assume they merge first. It is coordination between humans, not protection of a file.
  • Your infrastructure sits in private subnets with no inbound internet access. What does that imply for a hosted run service?
    The hosted control plane cannot reach your providers' endpoints, so you deploy agents inside the network that poll the service for work and execute runs locally. That gives you a fleet to operate and patch, and those agents now hold the credentials — so some of the operational and security burden you were outsourcing comes back.

saying these in an interview costs you the question

  • A run service means you no longer need plan review
  • Atlantis and HCP Terraform are the same product
  • Adopt one on day one, before there is any duplication
  • Hosted runs work fine against private-only networks
  • The run service replaces state locking entirely

context