How would you run a scheduled Terraform drift-detection job in CI, and how should the job decide whether to raise an alert?
answer
- cron plan, non-interactive, per root module
- exit codes 0, 1, 2 mean three different things
- drift or unapplied merge — say which
- the job competes for the state lock
- detect and alert; never auto-remediate
basics
~20 sRun terraform init and a non-interactive plan on a schedule against the deployed branch, and branch on -detailed-exitcode: 0 means no changes, 2 means the world differs from code, 1 means the run itself failed. Alert on 2, page on repeated 1.
solid answer
~40 sOn a cron schedule, the job checks out the branch that is actually deployed, runs `terraform init -input=false`, then `terraform plan -detailed-exitcode -input=false -lock-timeout=<n>` and inspects the exit code: `0` no changes, `2` changes present, `1` Terraform errored. Exit 2 is the drift signal — post the plan output to a channel or open a ticket, never auto-apply it. Two things make the difference between a useful alarm and one everyone mutes. First, exit 2 conflates real drift with merged-but-unapplied code, so run it against the exact revision that was last applied, or use `plan -refresh-only` when you want state-versus-reality only. Second, this job competes for the state lock with real applies, so give it a `-lock-timeout` and treat lock failures as retry-later, not as drift. Credentials should be scoped to reads.
code
bash · 15 lines#!/usr/bin/env bash
set -euo pipefail
terraform init -input=false -no-color
set +e
terraform plan -detailed-exitcode -input=false -no-color -lock-timeout=2m > plan.txt 2>&1
code=$?
set -e
case "$code" in
0) echo "OK: infrastructure matches configuration" ;;
2) echo "DRIFT OR UNAPPLIED CHANGE"; cat plan.txt; exit 1 ;;
*) echo "PLAN FAILED (exit $code)"; cat plan.txt; exit "$code" ;;
esacgo deeper
Know that plan can run on a schedule to check whether reality still matches the code, and that -detailed-exitcode is what lets a script tell 'no changes' from 'changes found'.
Be able to write the job: init and plan non-interactively, branch on exit codes 0, 1 and 2, and explain why a plain successful plan's exit code alone is useless for detection.
Show the operational judgment — exit 2 conflates drift with unapplied commits, the job contends for the state lock, refresh multiplies API calls across modules, and remediation stays human.
Own the programme: which modules are worth monitoring, what response time drift warrants, who is paged, and how you keep the alert credible instead of the thing everyone filters to a dead channel.
## Why the job exists A plan run on a pull request tells you what your change will do. It says nothing about the weeks in between, when someone clicks in the console during an incident, an operator widens a security group "temporarily", or a service mutates a field on your behalf. Terraform will notice all of it — but only the next time somebody happens to plan that root module, which for a stable module might be months. A scheduled drift job turns "we find out eventually" into "we find out tomorrow morning". ## The mechanics The job is an ordinary Terraform run with the interactivity stripped out: ```bash set -euo pipefail terraform init -input=false -no-color set +e terraform plan -detailed-exitcode -input=false -no-color -lock-timeout=2m > plan.txt 2>&1 code=$? set -e case "$code" in 0) echo "no drift" ;; 2) echo "DRIFT"; cat plan.txt; exit 1 ;; *) echo "terraform failed"; cat plan.txt; exit "$code" ;; esac ``` `-detailed-exitcode` is the whole trick: without it, a successful plan exits `0` whether or not it found anything, and the job has to scrape human-readable output. With it, plan returns `0` for no changes, `1` for an error, `2` for changes present — the only three values, and each means something different to your alerting. ## Deciding what to alert on **Exit 2 is not automatically drift.** A normal plan diffs configuration against refreshed state, so exit 2 means "the world does not match this code", which can be either real drift *or* a commit that was merged and never applied. Two ways to keep the signal honest: - Run against the revision that was last successfully applied, not the tip of the branch, so unapplied code cannot appear in the diff. - Or run `terraform plan -refresh-only`, whose diff is state-versus-reality only and therefore isolates out-of-band change specifically. Most teams run the normal plan, because "deployed infrastructure does not match main" is worth knowing regardless of which of the two causes it is — but the alert text should say both are possible so the responder does not go hunting for a console culprit that never existed. **Exit 1 is an operations signal, not a drift signal.** Expired credentials, a throttled provider, a removed upstream module, a lock held by someone else. Treat it separately: a single failure is noise, the same module failing every night is a broken pipeline nobody is watching. ## Operating it without making enemies - **State locks.** A drift plan takes the same lock as a real apply. Pass `-lock-timeout` so a plan that collides with a deploy waits instead of failing instantly, and classify a lock timeout as "retry later", never as drift. - **Refresh cost.** Every run re-reads every resource in state. Multiply that by every root module in the estate on the same cron minute and you will find your provider's rate limiter. Stagger the schedule. - **Credentials.** The job only needs to read, so give it a read-scoped role — ideally through short-lived federated credentials rather than a static key sitting in the scheduler. - **Never auto-apply.** "Detected drift, so I reverted it" is how a scheduled job destroys the emergency capacity someone added at 2am. Detection and remediation are separate decisions with a human in between. - **Redaction.** Plan output can contain values you would rather not paste into a chat channel. Post a link to the job, or a summarised resource list, rather than the raw diff. ## Making the output usable A wall of plan text at 6am gets muted within a fortnight. What works: name the root module, list the addresses that changed, and route to the team that owns that module. `terraform show -json` over a saved plan gives you machine-readable output if you want to build a proper summary rather than shipping raw text. ## What to say out loud Schedule a non-interactive plan per root module, branch on `-detailed-exitcode`, alert a human on `2` with the resource addresses, keep `1` in a separate operational bucket, hold a `-lock-timeout` so you never fight a real deploy, give the job read-scoped credentials, and never let it apply anything by itself.
- Your drift job fires exit code 2, but nobody touched the console. What else explains it?A commit merged to the deployed branch that was never applied — a normal plan diffs config against refreshed state, so unapplied code looks identical to drift. Also: a provider version bump changing how an attribute is read, or a resource whose value the platform legitimately mutates. Pin the job to the last-applied revision, or use `plan -refresh-only`, to isolate true out-of-band change.
- Should the scheduled job automatically apply the plan to correct drift?No. Drift is sometimes the only thing keeping production up — capacity added during an incident, a firewall rule opened for a migration. An unattended apply reverts it at 3am with nobody watching. Detect, attribute and notify; let a human choose between reverting, adopting the change into code, or leaving it alone.
- How do you stop the drift job from fighting real deploys over the state lock?Pass `-lock-timeout` so the plan waits rather than failing the moment a deploy holds the lock, schedule it outside peak deploy windows, and classify a lock-acquisition failure as a retryable operational error rather than a drift alert. Never let it call `force-unlock`; that exists for a genuinely orphaned lock, not for a scheduler in a hurry.
- Why give the drift job read-scoped credentials when plan does not change anything?Because it runs unattended on a schedule with no reviewer, which makes its credentials an attractive target and an easy mistake — a copy-pasted apply step would then be able to change production. Refresh only needs read APIs, so scope the role to reads and issue short-lived federated credentials instead of a static key stored in the scheduler.
saying these in an interview costs you the question
- Scrapes 'No changes' from plan text instead of using -detailed-exitcode
- Thinks exit code 1 means changes were detected
- Auto-applies the drift plan to 'self-heal' production
- Ignores the state lock and force-unlocks when the job collides with a deploy
- Assumes any exit-2 plan proves someone edited the console