skip to content

An engineer fixed a production outage by changing a resource directly in the cloud console overnight, and your infrastructure code no longer matches. Walk through how you decide what to do about it.

level: seniorimportance: must knowfreq 66%

answer

  1. ask why before you converge
  2. is production still leaning on it?
  3. revert, codify, or adopt
  4. new object means adopt, not recreate
  5. doing nothing is silently reverted later

basics

~20 s

First establish what was changed and why, and whether the change is still load-bearing. Then choose one of three honest outcomes: revert to the code, codify the change into the code and apply, or adopt an unmanaged resource into the tool's record. Never re-apply blindly.

solid answer

~50 s

I start by finding out what actually changed and whether production currently depends on it — the answer to that decides everything else. If the manual change was a stopgap that is no longer needed, I revert by letting the code win, in a reviewed deploy during business hours, not as a side effect of somebody else's release. If it is the correct configuration and we simply learned something during the incident, I codify it: write the change into the repository, open a pull request so it gets the review it never had, and apply so the tool now sees no difference. If the engineer created something new by hand, there is nothing to revert to — I adopt it, bringing the existing object under management without recreating it, because recreating it would cause a second outage. The failure mode to name is the fourth option nobody chooses deliberately: leaving it, so the next unrelated apply silently reverts the fix.

go deeper

for a junior

Know that the right first step is asking what changed and why, not re-running the tool. Be able to say that the code and reality must be brought back together deliberately.

for a middle

Explain the three outcomes — revert, codify, adopt — and which situation each fits. Be clear that a hand-created resource needs adopting, because declaring and applying it would create a duplicate.

for a senior

Show the triage judgment: check the audit log for intent, establish whether production still depends on the change, watch for immutable attributes that make a revert destructive, and treat the revert itself as a reviewed production change.

for a principal

Frame recurring drift as an authorisation and process problem: decide whether standing console write access should exist, define break-glass with expiry and an automatic codify-afterwards obligation, and make drift rate a signal you track.

## Triage before remediation The instinct to "just run the tool and let the code win" is what makes this an interview question. Reverting is one valid outcome of three, and picking it without asking why the change exists risks re-creating the outage it fixed. Three questions come first: 1. **What changed, exactly?** Read the diff attribute by attribute, and cross-check the provider's audit log for who made the change and when. The audit trail is the only place the *intent* survives, since the change bypassed pull-request review by definition. 2. **Is production currently depending on it?** A widened timeout that is holding a payment flow together is very different from a debug flag someone forgot to remove. 3. **Is the resource managed at all?** A modified managed resource and a brand-new hand-created resource look similar on a dashboard and have completely different remedies. ## The three honest remediations ### Revert — the code is right, reality is wrong You decide the declared configuration is correct and let the tool converge reality back onto it. This is the default when the manual change was a temporary probe, an experiment, or something that violates policy — an over-permissive network rule, a disabled log destination. The discipline is *how*: do it as its own reviewed, announced change, during hours, with the person who made the manual edit in the loop. A revert is a production change with real blast radius, and some reverts are destructive — if the manual edit touched an immutable attribute, converging back may replace the resource rather than update it. Read the diff before approving. ### Codify — reality is right, the code is stale The incident taught you something true: the instance really did need more memory, the health-check threshold really was too tight. Write it into the repository, open a pull request, get it reviewed, and apply. The apply itself becomes a no-op because reality already matches — which is exactly the point. You have converted an undocumented emergency change into a reviewed, version-controlled decision, and the next disaster-recovery rebuild will now reproduce the working configuration. This is the outcome most teams under-use, because it feels like paperwork after the incident is over. It is the step that keeps the repository trustworthy. ### Adopt — the object exists but the tool has never heard of it If the engineer created a whole new resource by hand, there is no prior declaration to revert to. Writing a declaration and applying it naively creates a *second* object, and deleting the hand-made one to let the tool rebuild it is a self-inflicted outage. The correct move is to adopt it: write the configuration to match what exists, then bring the existing object into the tool's record so it is managed from now on without being recreated. Every mature IaC tool has a mechanism for this, and it is worth naming as the reason the mechanism exists. Adoption is fiddly in practice — you must reconstruct the declaration accurately enough that the first run after adoption reports no changes. Treat "the comparison is clean" as the acceptance criterion. ## The fourth outcome: scoping the attribute out There is a fourth path, applicable when the changing field genuinely belongs to another system rather than to a person: declare the attribute out of scope so your tool stops trying to own it. That is the right answer for autoscaled capacity or a field another controller writes. It is the *wrong* answer to a human console edit, because it removes the only control that would have caught the next one. Reach for it deliberately, never as a way to silence a noisy report. ## The non-choice The outcome to name as the failure mode is doing nothing. The difference stays, everyone forgets, and weeks later an unrelated deploy — a tag change, a new subnet — converges the whole configuration and silently undoes the fix. The person deploying had no reason to look at that resource, so the outage returns with no obvious cause. "Drift is remediated by the next apply whether you decided or not" is the sentence that shows you have lived it. ## Closing the loop A senior answer does not stop at the resource. Two follow-through items belong in the incident review: was console write access necessary, or could the fix have gone through the pipeline with an expedited approval; and if console access genuinely was necessary, is there a break-glass path that expires and automatically raises a ticket to codify the change afterwards. Drift is a symptom; recurring drift on the same team is a process defect.

  • Why is naively declaring a hand-created resource in code and applying it a dangerous move?
    Because the tool has no record linking your declaration to the existing object, so it creates a second one. You end up with duplicate infrastructure, and cleaning up means deleting a live resource that traffic may already be using. Adoption exists precisely to attach the declaration to the object that is already there.
  • When is reverting the wrong first response even though the code is the source of truth?
    When the manual change is currently keeping production healthy. Converging back re-creates the outage, at a time nobody chose. It is also wrong when the manual edit touched an immutable attribute, because convergence would replace rather than update the resource. Establish dependency and read the diff before letting the code win.
  • What should come out of the incident review beyond fixing this one resource?
    Two things: whether console write access was actually needed, or an expedited pipeline path would have been fast enough; and if it was needed, whether break-glass access expires and automatically raises a follow-up to codify the change. One drift event is an incident detail; a pattern of them is a process defect in how emergency change is authorised.
  • How do you make sure an adoption actually worked?
    The acceptance criterion is a clean comparison: after adopting, running the tool must report no changes to make. If it still wants to modify or replace the resource, your reconstructed declaration does not match reality yet, and applying would mutate a live production object. Iterate on the declaration until the difference is empty.

saying these in an interview costs you the question

  • Immediately re-running the tool to make the code win
  • Assuming the manual change was a mistake without checking
  • Writing a new declaration for a hand-made resource and applying it
  • Scoping the attribute out just to silence the report
  • Fixing the resource and never touching the repository

context