skip to content

CI/CD Integration

In a real team Terraform runs in a pipeline: plan on the pull request, apply on merge, with short-lived credentials. Describing that flow — including how the reviewed plan reaches the apply job — is a standard system-level question.

part ofTerraformoverview, primer and where to startread it →
on this pageshow

questions

5

In a team's Terraform pipeline, why does terraform plan run on the pull request while terraform apply runs only after the change merges to the main branch?

level: juniorimportance: must knowfreq 78%

answer

  1. review the effect, not the intent
  2. cheap veto before the irreversible half
  3. merged branch describes production
  4. read-only on the PR, write on main
  5. one commit, one audit trail

basics

~20 s

A plan on the pull request turns the proposed change into a reviewable diff before anything is touched. Apply runs from main so only merged, reviewed code ever changes real infrastructure, giving one source of truth and one audit trail.

solid answer

~50 s

The pull request is where a human still has a cheap veto, so that is where `terraform plan` belongs: it renders exactly what would be created, changed or destroyed, and the pipeline posts that output back onto the PR so the reviewer approves a diff rather than approving HCL and guessing. The plan job needs no permission to change anything. Apply is the irreversible half, so it is triggered by the merge to main and runs only from main's commit — that keeps a single branch as the description of what production should be, makes `git log` the change record, and means nobody can apply from a fork or a feature branch. In automation both commands run with `-input=false` so a missing variable fails the job instead of blocking on a prompt, and the apply job usually sits behind an approval gate as well.

go deeper

for a junior

Be able to say plainly that plan is the safe, read-only preview run on the pull request and apply is the irreversible step run after merge, and that the reviewer is approving the plan output, not just the code.

for a middle

Explain the mechanics: -input=false so the job fails instead of hanging, a pinned Terraform version and committed lock file so the plan is reproducible, and separate credentials for the plan and apply jobs.

for a senior

Show you have operated this. Talk about plan/apply skew when pull requests queue up, break-glass applies putting state ahead of main, and detecting which root modules a shared-module change actually affects.

for a principal

Own the policy: which environments get an approval gate and who holds it, how emergency access is granted and reconciled afterwards, and how the same pipeline shape stays affordable across dozens of root modules and teams.

## The pipeline in one paragraph A team Terraform repository normally has exactly two automated paths. On a pull request, CI checks out the branch, runs `terraform init`, then `terraform fmt -check`, `terraform validate` and `terraform plan`, and publishes the plan output where the reviewer can read it — usually as a comment on the PR. Nothing is created. On a merge to `main`, CI checks out `main`, initialises again and runs `terraform apply`, frequently behind a manual approval. Everything else — nightly drift runs, ad-hoc destroys of ephemeral environments — is an extra path bolted onto that spine. ## Why plan belongs on the pull request HCL is not self-evident. A three-line change to a variable default, a module version bump, or an edit to a list can expand into a plan that replaces a database. Reviewing the code alone means reviewing intent; reviewing the plan means reviewing effect. The plan makes the effect concrete: how many resources are added, changed, destroyed, and which ones. The reviewer's real job is to look for the destroys and the replacements, and for a resource count that does not match the size of the diff. Running it in CI rather than on a laptop matters for a second reason: reproducibility. The pipeline runs a pinned Terraform version, a committed provider lock file and the same backend every time, so the plan the reviewer reads is the plan the machine will act on — not one produced by whatever versions happen to be on one engineer's machine. ``` terraform init -input=false terraform plan -input=false -no-color -out=tfplan ``` `-input=false` is the flag that separates automation from interactive use. Without it, a variable with no value makes Terraform prompt on stdin, and a CI job with no stdin simply hangs until its timeout. `-no-color` strips ANSI escapes so the output is readable when pasted into a PR comment. Terraform also honours the `TF_IN_AUTOMATION` environment variable, which suppresses the "run terraform apply next" style hints that make no sense in a pipeline. ## Why apply belongs to main Apply is the half that cannot be undone by closing a browser tab. Tying it to the merge commit gives three properties: 1. **One source of truth.** Whatever is on `main` is the claim about what production looks like. If apply could run from feature branches, two branches could fight over the same state and `main` would describe nothing. 2. **An audit trail for free.** "Who changed prod" becomes "which commit was applied", answerable from Git history plus the run log, without anyone reconstructing console activity. 3. **A credential boundary.** The plan job can hold read-only cloud credentials; only the apply job, which runs on a protected branch, holds credentials that can write. A pull request from an untrusted contributor therefore cannot reach anything that can be changed. The apply job typically also requires a human approval before it proceeds, because merge approval and change approval are not always the same decision — the reviewer approved the code, and someone still has to decide that now is the time to touch production. ## The failure modes this shape has The most common one is **plan/apply skew**. The plan on the PR was computed against the base branch and the state as they were at plan time. By the time the PR merges, another PR may have merged first, or someone may have changed the same resources by hand. Whatever the pipeline does at apply time — re-plan on `main`, or carry the reviewed plan file forward as an artifact — the reviewed plan is a snapshot with an expiry date, and a long-lived PR should be re-planned before it is trusted. The second is **the escape hatch that becomes the habit**. Everyone keeps credentials for emergencies, and every emergency apply from a laptop puts state ahead of `main`, so the next pipeline plan shows a diff nobody wrote. If break-glass access exists, applying from it should be loud, logged, and followed by a commit that makes `main` match reality again. The third is **partial coverage**. A pipeline that plans only the directories a PR touched will miss the case where a change to a shared module affects a root module the PR did not touch. Either plan everything, or make the affected-directory detection follow module dependencies rather than the file paths in the diff.

  • The pull request was opened last Monday and its plan is a week old. What should the pipeline do before applying it?
    Treat the old plan as advisory and re-plan. Other pull requests may have merged and other people may have changed resources directly, so the diff computed a week ago no longer describes what apply would do. Either re-run plan on the merge commit and gate the apply on that fresh output, or require the PR to be updated against the base branch, which forces a new plan run.
  • Why do the Terraform commands in a CI job pass -input=false?
    Because a CI job has no interactive stdin. Without `-input=false`, a variable with no supplied value makes Terraform prompt for it, and the job hangs there until the runner's timeout kills it — which looks like an infrastructure problem rather than a missing variable. With the flag, Terraform fails immediately with a clear error naming the variable.
  • Where would you put the human approval in this pipeline, and what does it protect against?
    On the apply job, after the merge, showing the plan being applied. Merge approval says the code is correct; the apply approval says now is an acceptable time to change production — during a freeze, an incident, or a risky replacement those are different answers. The approver must see the actual plan, otherwise the gate only records a click.

saying these in an interview costs you the question

  • Apply straight from the feature branch to test it faster
  • The reviewer reads the HCL diff and skips the plan output
  • Plan needs write credentials because it talks to the cloud
  • A plan from last week is still valid at merge time
  • Emergency applies from a laptop need no follow-up commit

context

open as a page

In a Terraform pipeline, what cloud credentials should the plan job hold compared with the apply job, and why is a single long-lived access key stored in CI a poor choice for both?

level: middleimportance: should knowfreq 48%

basics

~20 s

The plan job should get read-only access to the managed resources; only the apply job needs permission to create, change and destroy. Both should receive short-lived credentials issued per run through federation, not one static key sitting in CI forever.

open as a page

Two pipeline runs for the same Terraform root module start at the same time and the second one fails because it cannot acquire the state lock. How should the pipeline be designed so this stops being a problem?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Serialise runs per state in the pipeline itself: one concurrency lane per root module, queueing rather than cancelling, so a second run waits its turn instead of racing. The state lock is a last-resort safety net, not a scheduler.

open as a page

In a CI pipeline, why is the plan job's saved plan file handed to a separate apply job as a build artifact, and what problems does storing that artifact introduce?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Passing the saved plan forward guarantees the apply performs exactly the change a human approved, with no re-plan in between. The costs are staleness once state moves on, and an artifact that contains sensitive values in the clear.

open as a page

When would you run Terraform through a purpose-built run service such as Atlantis or HCP Terraform instead of hand-written plan and apply jobs in your existing CI system?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Adopt a run service when the number of Terraform repositories makes hand-rolled jobs a maintenance burden, and when you want cloud credentials, state, run history and approvals owned by one system instead of scattered across CI configurations.

open as a page