skip to content

You start drift detection on an AWS CloudFormation stack and the call returns immediately with only an ID. Which calls do you make to get the per-resource result, and what do the resource statuses IN_SYNC, MODIFIED, DELETED and NOT_CHECKED tell you?

level: middleimportance: should knowfreq 55%

answer

  1. start call returns only a job id
  2. poll the detection status
  3. roll-up first, per-resource detail second
  4. MODIFIED carries expected versus actual
  5. single resource answers synchronously

basics

~10 s

DetectStackDrift is asynchronous: it returns a StackDriftDetectionId that you poll with DescribeStackDriftDetectionStatus, then read DescribeStackResourceDrifts for per-resource statuses and the expected-versus-actual property differences.

solid answer

~40 s

`DetectStackDrift` only kicks the run off — it hands back a `StackDriftDetectionId`. You poll `DescribeStackDriftDetectionStatus` with that id until `DetectionStatus` becomes `DETECTION_COMPLETE` (or `DETECTION_FAILED`); that response also carries the roll-up `StackDriftStatus` and a `DriftedStackResourceCount`. The detail lives in `DescribeStackResourceDrifts` for the stack, which returns one entry per resource with a `StackResourceDriftStatus`. `IN_SYNC` means the compared properties matched; `MODIFIED` means at least one differed, and the entry carries `ExpectedProperties`, `ActualProperties` and a list of `PropertyDifferences` with the path, both values and a difference type; `DELETED` means the resource no longer exists; `NOT_CHECKED` means CloudFormation made no comparison at all. If you only care about one resource, `DetectStackResourceDrift` takes a stack name plus a logical id and answers synchronously.

code

bash · 15 lines
bash
ID=$(aws cloudformation detect-stack-drift \
  --stack-name payments-prod \
  --query StackDriftDetectionId --output text)

while true; do
  STATUS=$(aws cloudformation describe-stack-drift-detection-status \
    --stack-drift-detection-id "$ID" \
    --query DetectionStatus --output text)
  [ "$STATUS" = "DETECTION_IN_PROGRESS" ] || break
  sleep 5
done

aws cloudformation describe-stack-resource-drifts \
  --stack-name payments-prod \
  --stack-resource-drift-status-filters MODIFIED DELETED

go deeper

for a junior

Remember the shape: one call starts the job, another reports whether it finished, a third lists what differs. Know the four per-resource statuses by name.

for a middle

Be ready to write the poll loop and to read a PropertyDifferences entry aloud — path, expected value, actual value, difference type — and to say when you would check a single resource synchronously instead.

for a senior

Show how you turn this into automation: scheduled runs, DriftedStackResourceCount as a metric, filtering to MODIFIED and DELETED for an actionable alert, and handling DETECTION_FAILED as a permissions signal rather than a flake.

for a principal

Decide what the organisation does with the signal. Detection across hundreds of stacks costs API calls and produces noise, so define who owns a drifted stack, what response time is expected, and when drift becomes a delivery-process problem instead of a ticket.

## Why the first call gives you nothing useful Checking a stack for drift means calling the read API of every service that owns a resource in it. For a large stack that is slow and rate-limited, so CloudFormation models it as a job rather than a request. `DetectStackDrift` starts the job and immediately returns a `StackDriftDetectionId`. Nothing about the result is in that response. ```bash ID=$(aws cloudformation detect-stack-drift \ --stack-name payments-prod \ --query StackDriftDetectionId --output text) ``` One detection run per stack at a time — starting a second while one is in flight is rejected. ## Polling the job `DescribeStackDriftDetectionStatus` takes that id and reports on the job itself: - `DetectionStatus` — `DETECTION_IN_PROGRESS`, `DETECTION_COMPLETE`, or `DETECTION_FAILED`. Failure is usually a permissions problem: the caller could not read one of the resources. - `StackDriftStatus` — the roll-up: `DRIFTED`, `IN_SYNC`, `NOT_CHECKED`, or `UNKNOWN`. - `DriftedStackResourceCount` — how many resources came back drifted, which is the number worth turning into an alarm. A scripted run is a poll loop on `DetectionStatus`, not a single call. ## Reading the per-resource detail The roll-up tells you *whether*, never *what*. For that, `DescribeStackResourceDrifts --stack-name` returns one entry per resource, each carrying a `StackResourceDriftStatus`: **`IN_SYNC`** — every property CloudFormation compared matched the expected value. Note the qualifier: properties it does not support are simply not part of the verdict. **`MODIFIED`** — at least one compared property differs. This entry is the useful one. It carries `ExpectedProperties` and `ActualProperties` as resolved JSON, plus `PropertyDifferences`: a list where each element has a `PropertyPath` (a JSON pointer into the resource's properties), the `ExpectedValue`, the `ActualValue`, and a `DifferenceType`: - `NOT_EQUAL` — the property exists on both sides with different values. The everyday case. - `ADD` — a value is present in the live configuration that the expected configuration does not have, typically an extra entry in a list-shaped property such as a rule or a tag. - `REMOVE` — the property is gone from the live configuration. So an ingress rule added by hand at 3am shows up as an `ADD` under the security group's ingress property path, with the rule itself as the actual value. That is exactly the forensic detail a reviewer wants. **`DELETED`** — the resource is gone. Something deleted it outside the stack, and CloudFormation now holds a logical id pointing at nothing. A subsequent stack update will typically try to act on a resource that no longer exists. **`NOT_CHECKED`** — no comparison happened, almost always because drift detection is not supported for that resource type. It is not a clean bill of health, and in most real stacks there are several of these. You can narrow the listing with `--stack-resource-drift-status-filters` — filtering to `MODIFIED DELETED` is the usual way to get an actionable list, and filtering to `NOT_CHECKED` is how you find out how much of the stack the report never looked at. ## One resource at a time If you already know which resource you care about — say you are verifying a fix — `DetectStackResourceDrift --stack-name X --logical-resource-id Y` checks a single resource and answers **synchronously**, returning that one `StackResourceDrift` structure directly. No job id, no polling. It is also the cheap way to re-check after correcting something, instead of re-scanning a hundred-resource stack. ## Practical shape of an automation Because detection is on demand, teams wrap the whole sequence in a scheduled job: start detection on each stack, poll to completion, list the `MODIFIED` and `DELETED` entries, and publish `DriftedStackResourceCount` as a metric with an alarm on it. Two details bite people writing this the first time: the poll loop is mandatory because the start call is asynchronous, and the roll-up alone is not enough to page a human — the property differences are what make the alert actionable rather than merely alarming. A last subtlety worth carrying into the interview: the results are a snapshot taken while the job ran. If a deployment is in flight during detection you can get differences that are simply mid-update states, which is why scheduled detection is usually pointed at quiet windows.

  • What is the difference between DifferenceType NOT_EQUAL and ADD in a drift report?
    `NOT_EQUAL` means the property exists in both the expected and the actual configuration but the values differ — a changed instance class, a flipped boolean. `ADD` means the live configuration carries something the template never asked for, typically an extra element in a list-shaped property such as an ingress rule or a tag. Both appear in `PropertyDifferences` with the path and both values.
  • Detection comes back DETECTION_FAILED rather than complete. What do you check first?
    Permissions. Detection reads each resource through its owning service's API, so a role that can deploy the stack but cannot describe one of its resource types will fail the run. Check which principal started detection — the caller's own identity, or the stack role if one is in play — and confirm it holds read access to every resource type in the stack, then re-run.
  • Can you run drift detection on two stacks at once, or twice on the same stack?
    Across stacks, yes — they are independent jobs. On a single stack, no: one detection operation runs at a time and starting another while one is in progress is rejected, so an automation that fans out over an estate must track ids per stack rather than firing blindly and retrying.

saying these in an interview costs you the question

  • Expects DetectStackDrift to return the drift results directly
  • Reads only the stack-level status and never lists per-resource drifts
  • Treats NOT_CHECKED as equivalent to IN_SYNC
  • Thinks DELETED means CloudFormation deleted the resource
  • Assumes single-resource detection also needs polling

context