skip to content

Refresh and Drift Handling

Every plan quietly re-reads the real world first, and I can run that step on its own to see what changed behind my back. This is Terraform's concrete answer to the drift question.

part ofTerraformoverview, primer and where to startread it →
on this pageshow

questions

5

What does Terraform do during the refresh step of `terraform plan`, and what are you trading away when you run `terraform plan -refresh=false`?

level: middleimportance: must knowfreq 68%

answer

  1. three-way comparison, not two
  2. one provider read per stateful object
  3. plan does not persist the refreshed state
  4. speed and rate limits versus truth
  5. stale state means a wrong diff

basics

~20 s

Refresh re-reads every resource recorded in Terraform state from the provider API, so the plan diffs real infrastructure against your config. Passing -refresh=false skips those reads — faster and easier on rate limits, but the diff then trusts possibly stale state.

solid answer

~40 s

A Terraform plan compares three things: your configuration, the recorded state, and reality. Before it computes the diff, `terraform plan` refreshes — it calls the provider's read operation for every resource instance already in state and updates its in-memory copy with what the API returns. That is why a plan can show a change you never wrote: someone edited the console and refresh noticed. Since Terraform 0.15.4 a plain `plan` does not persist that refreshed state to the backend; it is used only to build the diff. `-refresh=false` skips the read calls entirely, so the plan is config-versus-last-known-state. It is genuinely faster on a large root module and avoids hammering provider APIs, but it hides drift, and the apply that follows can surprise you or fail because the world is not what state claims.

code

bash · 4 lines
bash
# Fast feedback on a draft PR: skip the provider reads
terraform init -input=false
terraform plan -refresh=false -input=false -lock-timeout=5m -out=tfplan
terraform show -no-color tfplan

go deeper

for a junior

Know that Terraform re-reads the real resources before showing you a diff, which is why a plan can list a change nobody wrote in code. Say plainly that this reading step never modifies infrastructure.

for a middle

Be ready to explain the three-way comparison of config, state and reality, that refresh issues one provider read per resource instance in state, and exactly what -refresh=false trades away.

for a senior

Show the operational judgment: refresh cost scales with state size and collides with provider rate limits, so the durable fix for slow plans is splitting root modules, not defaulting to -refresh=false in the pipeline.

for a principal

Own the policy. Decide where in the estate a fast unrefreshed plan is acceptable, where a fully refreshed plan is mandatory before apply, and how state sizing keeps refresh affordable as the estate grows.

## The three inputs to every plan Terraform's plan is a three-way comparison, not a two-way one: 1. **Configuration** — the `.tf` files: the desired state you wrote. 2. **State** — the JSON record of what Terraform believes it created, including each resource's attributes and the provider-assigned ID. 3. **Reality** — what the provider's API says exists right now. State alone is not trustworthy, because anything outside Terraform can change reality: a colleague in the console, an auto-scaling controller, a support engineer resizing a database during an incident, or a service that mutates a field on your behalf. So before Terraform computes any diff, it reconciles state against reality. That reconciliation is the **refresh**. ## What refresh actually does For every managed resource instance in state, Terraform asks the owning provider to read that object by its stored ID. The provider issues one or more API calls and returns the current attributes, and Terraform updates its in-memory state entry with them. Three outcomes matter: - The object matches state — nothing changes. - The object exists but some attributes differ — state is updated to the real values, and the subsequent diff will propose bringing them back in line with the configuration. This is what surfaces as `Note: Objects have changed outside of Terraform` at the top of the plan output. - The object is gone (deleted out of band) — Terraform removes it from its in-memory state, and the plan then proposes to create it. Data sources are read here too, since their results feed the diff. A detail people get wrong in interviews: since Terraform 0.15.4, `terraform plan` **does not write** the refreshed state back to the backend. The refreshed values live only for the duration of that plan. `terraform apply` does persist state, including the refreshed values, because it is writing state anyway. Older Terraform, and the now-deprecated standalone `terraform refresh` command, did write state as a side effect of what looked like a read-only operation — one reason the standalone command was deprecated in favour of `terraform apply -refresh-only`. ## What refresh costs Refresh cost scales with the number of resource *instances* in state, not with the size of your diff. A root module holding 900 resources issues roughly 900 read calls (more, for resources the provider reads in several calls) on every single plan — including a plan for a one-line tag change. On a busy account, with several pipelines planning concurrently, that is the fastest way to meet a provider's rate limiter and start seeing throttling errors, retries, and multi-minute plans. ```bash terraform plan # refresh, then diff (the default) terraform plan -refresh=false # diff config against last-known state ``` ## What `-refresh=false` trades `-refresh=false` buys wall-clock time and API quota, and pays for it with truth. The plan you get is "what would change if state were still accurate", which means: - **Drift is invisible.** An out-of-band change simply does not appear. - **The plan can be wrong.** If a resource was deleted in the console, the plan says "no changes" and the apply may fail, or a dependent resource may be built against something that no longer exists. - **Values in state may be stale**, so computed attributes fed into other resources or outputs can be out of date. Sensible uses: a fast feedback plan on an early draft PR, a large module where you already know reality is untouched, a targeted emergency change where the refresh itself is timing out, or a re-plan seconds after a plan you just ran. It is a speed knob, not a default. If your team is reaching for it constantly, the real fix is usually splitting an oversized root module so each plan refreshes fewer objects. ## The related knob you should not confuse it with `-refresh=false` turns refresh **off**. `-refresh-only` turns everything *else* off: it refreshes and then proposes only state updates, never infrastructure changes. They are opposite ends of the same axis, and mixing them up in an interview is a common tell. ## What to say out loud Refresh is Terraform's answer to "how does it know what really exists": it re-reads state's objects through the provider on every plan, which is what makes drift visible at all, and it costs one read per object. `-refresh=false` skips that for speed and accepts a diff computed against a record that may already be wrong.

  • Does a plain `terraform plan` write the refreshed values back to the state file?
    No. Since Terraform 0.15.4 the refresh performed during `plan` is in-memory only and is discarded when the command ends. `terraform apply` persists state, including refreshed values, because it writes state anyway. That change is why the standalone `terraform refresh` command — which silently mutated state during what looked like a read — was deprecated.
  • What happens in the plan if a resource in state was deleted in the console?
    Refresh asks the provider to read it, the provider reports it is gone, and Terraform drops it from its in-memory state. The plan then proposes to create it, so a normal apply recreates it. With `-refresh=false` the plan would report no changes at all, and you would only discover the deletion when something downstream failed.
  • Your plans take twelve minutes and half of that is refresh. What is the real fix?
    Split the root module. Refresh cost is proportional to the number of resource instances in a single state, so a 900-resource root module pays that on every plan regardless of what changed. Breaking it into smaller states with narrower blast radius cuts refresh time, plan noise and lock contention at once. `-refresh=false` only hides the symptom.

Refresh is taking inventory before writing the restock order: skip the count and you order from last month's ledger, which is faster and occasionally very wrong.

saying these in an interview costs you the question

  • Thinks plan compares config to state only, never to reality
  • Says terraform plan modifies infrastructure during refresh
  • Believes -refresh=false makes plans safer rather than blinder
  • Confuses -refresh=false with -refresh-only
  • Claims plan always writes the refreshed state to the backend

context

open as a page

In Terraform, what do `terraform plan -refresh-only` and `terraform apply -refresh-only` do, and when would you reach for them?

level: middleimportance: should knowfreq 52%

basics

~20 s

A refresh-only run re-reads real infrastructure and proposes updates to the Terraform state file only — never to infrastructure. Plan shows what drifted; apply writes those observed values into state, leaving the configuration and the real resources untouched.

open as a page

After `terraform apply -refresh-only` records a colleague's console change into Terraform state, what does the very next plain `terraform plan` propose, and why?

level: seniorimportance: should knowfreq 38%

basics

~20 s

It proposes changing the resource back to whatever the configuration says. Refresh-only updates state, never the .tf files, so the plan now compares an accurate state against an unchanged config and sees a difference it intends to correct.

open as a page

How would you run a scheduled Terraform drift-detection job in CI, and how should the job decide whether to raise an alert?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Run terraform init and a non-interactive plan on a schedule against the deployed branch, and branch on -detailed-exitcode: 0 means no changes, 2 means the world differs from code, 1 means the run itself failed. Alert on 2, page on repeated 1.

open as a page

You own hundreds of Terraform root modules, and refresh on every plan is now saturating provider API rate limits. How do you keep refresh and drift detection affordable?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Treat refresh cost as proportional to resources per state and runs per hour, then attack both: keep root modules small, tier how often each is fully refreshed, stagger schedules so runs do not collide, and reserve unrefreshed plans for cheap feedback rather than for gated applies.

open as a page