A resource in your infrastructure code shows drift on the same attribute after every deployment because another system — an autoscaler or a controller — legitimately writes that field. How do you resolve the conflict?
answer
- it is an ownership bug, not drift
- two writers, one field
- who has the better information?
- narrowest possible exclusion, documented
- every exclusion is a blind spot
basics
~20 sGive the field exactly one owner. Either your code stops managing that attribute and hands it to the other system, or you disable the other writer and keep it declared. Two writers on one field produce endless flapping and eventually ignored drift reports.
solid answer
~50 sThis is not drift to remediate, it is an ownership bug. Two systems both believe they set that attribute, so each run undoes the other and the report is never clean. The resolution is to pick a single writer per field. Usually the other system wins — an autoscaler reacting to live load knows more about the right capacity than a value committed in a repository last quarter — so I declare that attribute out of scope in my configuration and let the tool stop watching it. Occasionally my code should win, in which case I turn the other writer off rather than fight it. What I avoid is scoping out the whole resource, or scoping out fields just to make a noisy report go quiet: everything I stop managing becomes an invisible blind spot where a real misconfiguration will never be flagged. I scope narrowly, record why in a comment, and re-review it periodically.
go deeper
Know that some drift is caused by other automation doing its job, not by a person making a mistake, and that the fix is deciding which system should own the field.
Explain the single-writer principle and both resolutions: stop managing the attribute so the other system owns it, or disable the other writer and keep it declared. Say why keeping both is unstable.
Show what the exclusion costs — the attribute is no longer covered by that control — and the discipline around it: narrowest scope, a comment naming the owner, periodic review, and compensating detection elsewhere.
Own the boundary design across the estate: decide which layer creates and owns each class of object, avoid co-managing one resource from two systems at all, and hold a review process that stops exclusions accumulating unexamined.
## Recognise the shape Drift that reappears on the same attribute after every reconciliation is qualitatively different from drift caused by a human. Nobody is making a mistake. Two authorities are writing the same field, and each sees the other's write as a deviation to correct. The symptom is a report that is never clean and a deploy that always shows a change nobody asked for. Common instances: - **Autoscaling.** Your code declares a capacity; the autoscaler moves it in response to load. This is the textbook case. - **A platform controller.** A Kubernetes controller provisions a cloud load balancer and writes annotations, tags, or listener settings on an object your IaC also declares. - **Security or tag auto-remediation.** A compliance bot rewrites tags, closes a port, or attaches a policy. - **A managed service.** A rotation service updates an identifier or version on the resource. - **A second IaC repository.** The worst variant, because both writers are your own code and neither is obviously subordinate. ## The single-writer principle The fix is a design decision, not a tooling trick: **every mutable field should have exactly one system that owns it.** Once you accept that, there are only two resolutions. ### Hand the field to the other system Declare the attribute out of scope for your tool, so it stops comparing that field and stops trying to converge it. Your declaration supplies the *initial* value at creation, and thereafter the other system owns it. This is the right choice whenever the other writer has information your repository cannot have — live load, live threat signal, a rotation schedule. An autoscaler is reacting to reality; a number in a file is a guess made at commit time. A related, coarser variant is to stop managing the resource's *existence* under one tool entirely — for example, letting the controller create the load balancer and having your IaC only reference it. That is cleaner than co-managing a shared object attribute by attribute, when the tooling allows it. ### Take the field back Sometimes your code should win: the other writer is a leftover script, a scheduled function someone wrote, a bot whose policy is wrong for this resource. Then the answer is to disable that writer, not to keep two of them and pick a winner per run. "We just re-apply after the bot runs" is not a resolution; it is a race you will lose during an incident. ## The cost of scoping something out Every attribute you stop managing becomes a place where your controls do not reach. If you exclude a security group's rules because one automated system adds a rule, you have also stopped detecting the day a human adds an over-permissive rule to the same group. This is why scoping out is a decision with a security consequence, not a way to reduce noise. Practical discipline: - **Scope the narrowest thing that works** — one attribute, not the resource, and never a broad wildcard across many resources. - **Write down why**, in a comment next to the exclusion, naming the system that owns the field. A future engineer will otherwise find an unexplained exclusion and either delete it (reintroducing the fight) or copy it (spreading it). - **Put an expiry on it in review.** Exclusions accumulate; a periodic sweep asking "does this owner still exist?" keeps the blind-spot surface from growing silently. - **Consider compensating detection.** If your IaC no longer watches the field, something should — a policy scan on live infrastructure, or a provider-side rule. ## Distinguish it from phantom drift Repeating drift also comes from the provider normalising a value rather than from a competing writer — a reordered policy document, a reformatted duration. The signature differs: phantom drift shows a difference that is semantically identical, and no system claims to have written anything. That is fixed by expressing the value in the provider's canonical form or by a fix in the tool's comparison, not by handing ownership away. Diagnosing which of the two you are looking at is worth doing before you reach for an exclusion, because excluding a field to hide a formatting bug costs you real detection for no reason. ## Why interviewers ask this It separates people who treat drift reports as noise from people who treat them as information. The weak answer is "tell it to ignore that field" with no discussion of what is being given up. The strong answer names the ownership principle, chooses a single writer with a reason, scopes the exclusion tightly, documents it, and acknowledges that the excluded attribute is now unmonitored by this control.
- Why is the autoscaler usually the right owner of a capacity field rather than your repository?Because it holds information the repository cannot: current load, current failures, current time of day. A committed number is a guess frozen at review time. Letting code win means every deploy resets capacity to that guess, potentially in the middle of a traffic peak. The declaration supplies an initial value; the autoscaler owns it thereafter.
- How do you tell a competing-writer problem apart from phantom drift caused by provider normalisation?Check whether the difference is semantically meaningful. A reordered policy document or reformatted duration means the same thing and no system claims the write — that is normalisation, fixed by writing the canonical form. A genuinely different value, with an audit entry or an obvious automated actor behind it, means a second writer and needs an ownership decision.
- What is the risk of accumulating attribute exclusions over years?They become undocumented blind spots. Each one removes a field from your only automated comparison, and nobody notices when the system that justified it is decommissioned. The exclusions outlive their reason, the covered surface grows, and eventually a genuinely dangerous change lands in a field nothing watches. Periodic review of exclusions is the counterweight.
saying these in an interview costs you the question
- Just ignoring the field to make the report quiet
- Excluding the whole resource instead of one attribute
- Keeping both writers and re-applying after the other one runs
- Treating autoscaler-induced changes as human error
- Forgetting that an excluded attribute is no longer monitored