You are asked to set up continuous drift detection across a large estate of infrastructure-as-code repositories. How do you design it so it produces action rather than noise?
answer
- signal before audience
- cadence by blast radius, not uniformly
- read-only scanning identity
- route to the owning team, not a firehose
- detect and notify; do not auto-revert
basics
~20 sScan on a cadence matched to blast radius using read-only credentials, route each finding to the team that owns the code rather than a central channel, suppress known-owned attributes deliberately, and treat drift rate as a process metric rather than auto-reverting.
solid answer
~50 sI would run scheduled read-only comparisons per repository, on a cadence set by blast radius — production networking and identity daily, low-risk environments weekly — because a full refresh costs real provider API calls and will hit throttling if I scan everything constantly. Findings route to the team that owns that code, in their own channel with a link to the diff, never to a central firehose that nobody reads. Before any of that is useful I have to fix the noise: attributes legitimately owned by other systems get scoped out explicitly, and normalisation false positives get corrected, or the whole programme is ignored within two sprints. I do not auto-revert. Detection tells a human that something changed; converging production on a schedule with no approval is a way to cause an outage at 4am. And I track drift rate over time as a process signal — rising drift usually means console access or the emergency-change path is wrong, not that engineers are careless.
go deeper
Know that drift detection can be run on a schedule rather than only when someone deploys, and that the scan is a read of live infrastructure, not a change to it.
Explain why scanning has real cost — one or more provider API calls per managed resource — and why noise from provider normalisation and other-owner attributes must be removed before anyone will read the reports.
Show operational judgment: cadence tied to blast radius, staggered schedules away from deploy windows, read-only scanning credentials, findings routed to owning teams with the audit entry attached, and severity graded by the kind of attribute.
Own the programme: refuse blanket auto-remediation, define the narrow opt-in exceptions, treat drift rate as a metric about the change process, and set the access policy — read-only by default with time-limited break-glass that files its own codify-afterwards obligation.
## Detection is the easy half Running a comparison on a timer is a small amount of work. Making anyone act on the output is the actual problem, and it is why this is a design question rather than a tooling question. Most drift programmes die the same way: the scan is switched on, the first report contains hundreds of differences that are mostly meaningless, everyone mutes the channel, and six months later a real change goes unnoticed inside the noise. So the sequencing matters. **Fix the signal before you widen the audience.** ## Getting the signal clean Before alerting anyone: - **Eliminate phantom drift.** Provider normalisation — reordered documents, reformatted values, defaults filled in — produces differences with no meaning. Express the values canonically or fix the comparison. This is usually the single largest source of noise. - **Scope out the fields other systems own,** explicitly and with a written reason. Autoscaled capacity and controller-written tags will otherwise appear in every report forever. - **Baseline once and drive to zero.** Start with one repository, get it to a clean comparison, and only then add the next. A programme that begins with a hundred repositories at once never reaches a clean state anywhere. ## Cadence and cost A full comparison reads every managed resource, which is at least one provider API call each. Across a large estate that is expensive, slow, and a reliable way to trip rate limits — and the throttling will also affect the deploys running at the same time. So cadence is a risk decision: - High blast-radius, security-relevant scope — networking, identity, data stores, production accounts — daily or more. - Ordinary production services — daily to a few times a week. - Development and ephemeral environments — weekly, or not at all, since they are rebuilt anyway. Stagger the schedules rather than firing everything at midnight, and keep scans away from deploy windows: a scan that contends with a deploy for the same coordination lock delays real work, and the deploy is more important than the report. ## Safety of the scanning identity The scanner needs credentials to every account it inspects, which makes it a high-value target and a high-blast-radius job. Give it **read-only** permission. Detection is fundamentally a read; nothing about a scheduled scan requires the ability to mutate infrastructure. If your tooling wants to write a refreshed record as a side effect of comparing, treat that write as a mutation to think about — and prefer a mode that reports without persisting anything. ## Routing: ownership, not a firehose A finding is only actionable by the people who own the code. Route to them: the team channel that owns the repository, with the resource, the attribute, the before/after values, and the audit-log entry naming who changed it if the provider records that. Enriching with "who" is disproportionately valuable, because it turns a report into a conversation instead of an investigation. Severity should not be uniform. A drifted tag is a ticket. A security-group rule that opened, a logging destination that got disabled, an encryption setting that changed, or a deletion-protection flag that was cleared is a page. Classify by the *kind* of attribute, not by the number of differences. ## Do not auto-remediate The tempting next step is to close the loop: detect drift, converge automatically. Resist it, at least for production. Converging is a production change with a real diff, and the whole content of the earlier triage — is production leaning on this change, is the manual edit actually correct, would convergence replace rather than update the resource — is exactly what an automated reverter cannot evaluate. An unattended job that reverts a load-bearing emergency fix at 4am has caused an outage in order to satisfy a comparison. There are narrow exceptions where auto-revert is defensible: a small, explicitly enumerated set of security-critical attributes, in an environment where the code is unambiguously authoritative, with a loud notification. Treat that as an opt-in list, never the default. ## Drift rate is a process metric The most valuable output of a mature programme is not the individual finding; it is the trend. Drift concentrated in one team, or rising month over month, is telling you something about your change process: standing console write access that nobody needs, an emergency path too slow to use under pressure, or an ownership boundary that forces people to work around the pipeline. That leads to the organisational half of the answer, which is what distinguishes a lead-level response. Decide who may write in the console at all. The mature shape is read-only by default, with **break-glass** elevation that is time-limited, logged, announced, and carries an automatic obligation to codify the change afterwards — a ticket opened by the elevation itself, not by the memory of a tired engineer. Then drift detection becomes the verification that the obligation was met, rather than the only thing standing between you and an unknown production configuration.
- Why is a fully automated detect-and-revert loop risky in production?Because reverting is a production change and the decision needs context the job does not have: whether the manual edit is currently holding production up, whether it is actually the correct configuration, and whether converging would replace rather than update the resource. An unattended reverter can undo an emergency fix in the middle of the night and cause the outage it was meant to prevent.
- What would make you page someone about drift rather than open a ticket?The kind of attribute, not the volume. A security-group rule opening, encryption or logging turned off, deletion protection cleared, or a public-access setting flipped are all pages, because they are either an active exposure or a sign of an intrusion. Instance sizes, tags and capacity are tickets, handled in normal working hours.
- What does a rising drift rate in one team tell you?Almost always that the sanctioned change path is not usable under pressure, or that people hold standing console write access they do not need. It is a process finding, not an individual one. The response is to make the pipeline fast enough to use during an incident and to replace standing access with time-limited break-glass that files its own follow-up.
- Why does a scheduled scan need only read-only credentials?Because detection is a comparison of observed values against declared and recorded ones — entirely a read. A scanning identity holds access to every account you inspect, which makes it one of the most dangerous credentials in the estate, so it should have the narrowest possible power. Any workflow needing writes belongs to a separate, human-approved job.
saying these in an interview costs you the question
- Scanning everything continuously and ignoring API cost
- Sending every finding to one central alerts channel
- Auto-reverting drift in production without a human
- Giving the scanner write access to every account
- Measuring success by number of findings rather than actions