An AWS CloudFormation drift detection run reports a stack IN_SYNC, but you have good reason to believe someone changed one of its resources in the console. Give the reasons CloudFormation can miss a real change.
answer
- support is per resource type
- coverage is per property too
- NOT_CHECKED does not mean fine
- only explicitly set values are compared
- snapshot, not a watch
basics
~20 sCloudFormation drift coverage is partial: support exists per resource type and per property, unsupported resources come back NOT_CHECKED, and only values the template explicitly sets are compared. A run is also a point-in-time snapshot, so IN_SYNC is weak evidence, not proof.
solid answer
~50 sA stack-level `IN_SYNC` only says that nothing CloudFormation actually compared came back different — and it compares far less than people assume. Drift support is implemented per resource type, so any unsupported resource in the stack is reported `NOT_CHECKED` and contributes nothing; within a supported type, coverage is per property, so a changed field may simply not be one of the ones examined. It also only compares values the template or its parameters explicitly set, so a setting left to the service default is not tracked. On top of that, detection is a snapshot: it describes the moment the job ran, and a change made after it — or made and reverted before it — is invisible. So before believing a clean report, list the per-resource drifts and count the `NOT_CHECKED` entries. If drift matters, schedule detection and, better, remove the console write path.
go deeper
Know that drift detection does not cover everything: some resources come back NOT_CHECKED, and that status means nothing was compared rather than everything matched.
Be able to list the coverage limits concretely — per type, per property, only explicitly set values — and to pull the NOT_CHECKED entries out of the per-resource listing to show how much a report actually covered.
Demonstrate that you distrust the signal appropriately: schedule detection, alarm on drifted counts, run it before deployments, and argue for removing console write access rather than sampling for its consequences afterwards.
Own the framing that repeated drift is evidence of a change path people are working around. Decide the estate-wide policy — what is detected, what is prevented, who is accountable for a drifted stack, and what the acceptable window between runs is.
## What IN_SYNC actually asserts Read the claim narrowly. A stack-level `IN_SYNC` means: of the resources CloudFormation checked, and of the properties within those resources it supports, and at the instant the detection job read them, none differed from the expected values. Every clause in that sentence is a hole, and a senior answer names them. ## Coverage is per resource type Drift detection is not a generic mechanism CloudFormation applies to everything. Each resource type has to be supported explicitly, and the supported set — while large — is not everything. Any resource of an unsupported type comes back `NOT_CHECKED`, which is not a status people read carefully. Crucially, `NOT_CHECKED` resources do not push the stack roll-up to `DRIFTED`; a stack can be `IN_SYNC` overall while containing several resources nobody ever looked at. The practical move is to make that visible: ```bash aws cloudformation describe-stack-resource-drifts \ --stack-name payments-prod \ --stack-resource-drift-status-filters NOT_CHECKED ``` If that returns a long list, your clean report covers a fraction of the stack. Treat the count as a coverage metric, not as noise. ## Coverage is per property Even inside a supported type, drift comparison operates on properties, and not every property of a supported type is compared. A change to a field outside the compared set leaves the resource `IN_SYNC`. This is the sharpest edge of the whole feature, because the resource *was* checked — so nothing in the report hints that anything was skipped. ## Only what the template explicitly set CloudFormation compares against the *expected* configuration, which is built from what the template and its parameters actually declare. A property the template never mentioned, left at the service's default, is not part of the expectation — so a person setting it by hand is not obviously contradicting anything the stack said. The more sparse your template, the less surface drift detection has to compare, which is a real argument for declaring the properties you care about rather than relying on defaults. ## It is a snapshot, not a watch Detection runs when it is started and reports the state it read. Two consequences follow. First, a change made after the run is invisible until the next run, so the freshness of the answer is whatever your schedule is. Second, a change made *and reverted* between runs leaves no trace at all — the drift report is not an audit trail, and it never tells you who did anything. If the question is "who changed prod", the answer is an account-level audit log, not a drift report. ## Things it structurally never sees - **Resources next to the stack.** Detection walks the stack's own resources. A rule added to a security group the stack does not manage is out of scope entirely. - **Data, not configuration.** Bucket contents, table items, parameter values inside another service — none of those are resource properties. - **State that is not modelled.** Runtime conditions, in-place software on an instance, and anything a resource does after creation are outside the template's vocabulary. ## What to do instead of trusting the report A senior answer does not stop at the limitations; it says what changes. **Make detection routine and measured.** Because it is on demand, wire a scheduled job — an EventBridge scheduled rule invoking a small function that calls `DetectStackDrift` per stack, polls, and publishes `DriftedStackResourceCount` — and alarm on it. Publish the `NOT_CHECKED` count alongside so the coverage gap is visible rather than assumed. **Detect at the moment it matters.** Running detection immediately before a deployment is worth more than a nightly scan, because that is the moment an unnoticed hand edit is about to be silently overwritten or silently preserved. **Prefer prevention to detection.** The strongest control is that nobody can make the change by hand in the first place: production write access through the console is the thing that produced the drift, and permission design removes the whole class of problem that drift detection can only sample for afterwards. **Close the loop when drift is found.** Every `MODIFIED` result is a decision — revert the resource, or codify the change in the template and update. Leaving it means the template has stopped describing production, and the next person to redeploy the stack finds that out the hard way. ## The one-line version Silence is not proof. `IN_SYNC` means "nothing I looked at differed", and the interesting engineering question is always how much of the stack that sentence actually covers.
- How would you turn on-demand detection into ongoing assurance across a few hundred stacks?Schedule it: a periodic trigger invoking a function that lists stacks, starts `DetectStackDrift` on each, polls to completion and publishes `DriftedStackResourceCount` as a metric with an alarm. Publish the `NOT_CHECKED` count too, so coverage is visible. Respect the one-run-per-stack-at-a-time limit and API throttling by fanning out with concurrency control rather than firing everything at once.
- A drift report tells you a security group changed. How do you find out who did it?Not from the report — drift detection reads current configuration and records no actor, no timestamp and no history. It tells you the end state only. Attribution comes from the account's API audit trail, correlating the resource identifier with the write call that changed it. In an interview, say plainly that drift detection is a state comparison, not an audit log.
- Why is a sparse template worse for drift detection than an explicit one?Because the expected configuration is built from what the template and parameters actually declare. Properties left to service defaults are not part of the expectation, so someone setting them by hand is not contradicting anything recorded. Declaring the properties you care about — even when the default is already correct — widens the surface drift detection can compare and turns a silent change into a MODIFIED result.
saying these in an interview costs you the question
- Treats a NOT_CHECKED resource as equivalent to in sync
- Assumes every resource type and property supports drift detection
- Believes an IN_SYNC report proves nothing has changed since
- Expects the drift report to name who made the change
- Thinks properties left at service defaults are still compared