skip to content

In an infrastructure-as-code workflow, what does it mean for infrastructure to have "drifted", and what are the most common ways drift appears in a live cloud estate?

level: juniorimportance: must knowfreq 72%

answer

  1. reality moved, the code did not
  2. 3am console fix nobody put back
  3. not only humans — autoscalers and controllers
  4. provider defaults and normalised values
  5. the next apply silently reverts it

basics

~20 s

Drift is divergence between what your infrastructure code declares and what actually exists at the provider. Common causes: console or emergency edits, autoscalers and other controllers changing fields, provider-applied defaults, and a second tool managing the same resource.

solid answer

~50 s

Drift means the real infrastructure no longer matches the definition your code claims is true. The classic source is a human: someone opens the cloud console during an incident, widens a rule or bumps an instance size, and never puts it back in the repo. But plenty of drift is not a mistake — an autoscaler moves a desired capacity, a Kubernetes controller or a security-remediation bot writes a tag, a managed service rotates something on its own schedule, or a second IaC repo/tool is quietly managing the same object. There is also drift the provider creates for you: defaults it fills in, values it normalises, or fields it changes on a version upgrade. The important reflex in an interview is to say drift is *expected* in any real estate, so the workflow has to notice it rather than assume the code is reality.

go deeper

for a junior

Be ready to define drift in one sentence and give a concrete example, such as someone widening a firewall rule in the console during an incident. Say plainly that the code no longer matches reality.

for a middle

Explain that drift is not only human: autoscalers, other controllers, provider defaults and value normalisation all create differences. Distinguish drift from a pending, deliberate code change you have not applied yet.

for a senior

Show you expect drift in any real estate and treat a report as triage input, not an alarm. Point out that the next unrelated apply will silently revert a manual fix, and that phantom drift from normalisation destroys trust in the report.

for a principal

Own the framing that drift is a process signal: a rising drift rate says console access, on-call tooling or ownership boundaries are wrong. Argue about who may touch the console at all, rather than about individual resources.

## What drift is Infrastructure-as-code rests on one claim: the files in the repository describe the infrastructure that exists. **Drift** is any case where that claim has stopped being true — the live resource differs from what the code declares, and nobody updated the code. The word covers a spectrum. A security group with an extra ingress rule that a human added is drift. A load balancer whose target count moved because an autoscaler did its job is also drift, in the narrow mechanical sense, even though nothing is wrong. Part of maturing on this topic is learning that "drift detected" is a signal to triage, not automatically an incident. ## Where drift comes from **Humans in the console.** This is the canonical case and the one interviewers reach for. Production breaks at 3am, the on-call engineer has console access, and the fastest fix is three clicks. The fix works, the incident closes, and the repository still describes the broken configuration. Nothing surfaces the divergence until someone runs the tool weeks later and sees a change nobody expected. **Support and vendor action.** Cloud support engineers, managed-service automation and account-level policy enforcement can all mutate objects you believe you own. **Other automation that legitimately owns a field.** An autoscaler adjusts capacity. A Kubernetes controller writes annotations or provisions a load balancer. A tag-enforcement or auto-remediation bot rewrites tags or closes an open port. A certificate service rotates an identifier on the resource. None of these are human error; they are systems doing their job on a field your code also names. **A second tool or a second repository.** Two IaC codebases that both declare the same object will fight each other forever, each seeing the other's writes as drift. The same happens when a resource is managed by code *and* by a click-ops runbook. **The provider itself.** Providers fill in defaults you did not specify, normalise values (reordering a JSON policy document, lower-casing an identifier, expanding a shorthand), and occasionally change behaviour across API or provider-plugin versions. This produces *phantom* drift: the tool reports a difference that has no operational meaning. Phantom drift is worth naming in an interview because it is the main reason teams stop trusting their drift reports. **Deletion.** A resource removed by hand is drift too, and it is the variety that breaks tools hardest — a stateful tool will want to recreate it, and a stateless one may not notice the absence at all. ## Why it matters Drift undermines every promise IaC makes: - **Reproducibility.** If the code is not the truth, rebuilding the environment from the code produces something different from what you had — which is discovered at the worst moment, during a disaster recovery. - **Review as a control.** Pull-request review is the audit trail. A console change bypasses it, so a change reaches production with no reviewer, no ticket, and no record of intent. - **Safety of the next deploy.** The next routine apply will silently revert the manual fix, because the tool's job is to converge reality onto the code. A team that does not know about the drift will re-break production while deploying something unrelated. - **Compliance.** "The code is the control" is only a valid statement to an auditor if the running estate provably matches it. ## Drift versus a pending change Beginners conflate two different divergences. If you edit the code and have not applied yet, the code differs from reality — that is a **pending change**, deliberate and about to be applied. Drift is the opposite direction: reality moved out from under code nobody edited. Tools that keep a recorded state can tell these apart, because the record shows what the tool last believed it created. ## Configuration drift on servers There is a sibling notion of drift that is about long-lived machines gradually diverging from each other — hand-patched hosts, one-off installs, the "snowflake server". That is the same word for a related problem, and the standard answer to it is different: rebuild from an image rather than reconcile field by field. When an interviewer says drift in an IaC context, they usually mean the resource-level divergence described above; it is worth one sentence to show you know both senses exist. ## What a good answer sounds like Define drift in one sentence, give the 3am console edit as the human source, immediately add that machines cause drift too (autoscalers, controllers, provider defaults), and close on the consequence: the next apply will quietly undo it unless someone triages the difference first.

  • Is drift always a problem that needs fixing?
    No. Some drift is a system doing its job — an autoscaler moving capacity, a controller writing a tag, a provider normalising a value. That drift is triaged by deciding which system owns the field, not by reverting. Real problems are the unreviewed human change and the second tool competing for the same object; those need a decision and, usually, a code change.
  • Why is drift especially dangerous just before a routine deployment?
    Because the tool converges reality onto the code. An unrelated apply will also undo the manual fix that is currently holding production together, and the person deploying has no reason to look at that resource. This is why a diff should be read in full before approval, not skimmed for the resource you meant to change.
  • How does drift show up as a disaster-recovery problem?
    DR assumes you can rebuild the environment from the repository. If months of undocumented console changes are keeping production healthy, the rebuilt environment is the configuration from before those fixes — it will fail in the same ways, at the moment you can least afford to debug it. Drift turns a recovery plan into an untested guess.

saying these in an interview costs you the question

  • Claiming drift only happens if someone is careless
  • Assuming all drift is human-caused console clicking
  • Thinking a pending code change is the same as drift
  • Saying drift is impossible once you use IaC
  • Treating every reported difference as a real incident

context