skip to content

CloudFormation

AWS's native templating service: stacks, change sets, and drift detection, with the service itself owning rollback. Interviewers ask about change sets and stuck rollback states, because that is what you fight during a real incident.

on this pageshow

questions

16

In AWS CloudFormation, what exactly does stack drift detection compare, and what does CloudFormation do about the drift it finds?

level: juniorimportance: must knowfreq 50%

answer

  1. expected template values versus live configuration
  2. a report, not a repair
  3. runs only when you ask
  4. per-resource status plus property diff
  5. re-applying an unchanged template fixes nothing

basics

~20 s

Drift detection compares a stack's live resource configuration against the values its template and parameters expect, then reports the differences. It is read-only and on demand: CloudFormation never reverts drift for you and never watches resources continuously.

solid answer

~50 s

CloudFormation keeps the template it last applied to a stack plus the parameter values used with it; together those are the *expected* configuration. When you start a drift detection run, it reads the current configuration of each stack resource from the underlying service and compares the two, resource by resource and property by property. The output is a report: each resource comes back `IN_SYNC`, `MODIFIED`, `DELETED` or `NOT_CHECKED`, and modified resources carry the expected and actual values of the properties that differ. That is all it does — it does not roll anything back, does not block the next stack update, and does not update the template to match reality. Fixing drift is a separate, deliberate act: either change the resource back, or change the template to describe what is now true and update the stack.

go deeper

for a junior

Know the one-line definition: it compares live resources against what the template and parameters said, and reports differences without fixing them. Say plainly that it runs only when you start it.

for a middle

Be ready to name the per-resource statuses and describe the two ways drift is resolved — change the resource back, or change the template and update. Explain why an unchanged template produces no update at all.

for a senior

Show judgment about what a clean report is worth: it is a point-in-time sample of a partial set of properties. Talk about scheduling detection and about preventing console writes rather than detecting them afterwards.

for a principal

Own the position that drift is a symptom of a broken change path, not just a report to act on. Decide where detection sits in the delivery pipeline and what an organisation does when the on-call fix and the code inevitably diverge.

## What CloudFormation means by "expected" Every CloudFormation stack has a stored template — the exact one that was last successfully applied — plus the parameter values it was applied with. Resolve the parameters, conditions and intrinsic functions in that template and you get a concrete statement of what every resource in the stack is supposed to look like: this bucket has versioning enabled, this security group has exactly these two ingress rules, this table has this billing mode. That resolved picture is the **expected configuration**. Drift is what happens when the real world stops matching it. Somebody opens the console at 3am to unblock an incident and adds an ingress rule. A support engineer flips a setting through a service API. An automated tool with write access to the account edits a tag. None of that goes through CloudFormation, so the stored template still describes the old shape while the live resource has moved on. ## What a detection run actually does Drift detection is a read operation. When you start it, CloudFormation walks the stack's resources, calls the owning service's read APIs to fetch each resource's current configuration, and compares that against the expected values. Each resource lands in one of four states: - `IN_SYNC` — the properties CloudFormation compared all match. - `MODIFIED` — at least one compared property differs; the report carries the expected value, the actual value and the path of each differing property. - `DELETED` — the resource no longer exists; something removed it outside the stack. - `NOT_CHECKED` — CloudFormation made no comparison, usually because drift detection is not supported for that resource type. Those per-resource results roll up to a stack-level status. If anything is `MODIFIED` or `DELETED`, the stack is `DRIFTED`. ## It reports; it does not repair This is the point candidates most often get wrong. A drift report changes nothing. The drifted resource keeps its hand-edited configuration, the stack keeps its stored template, and the next deployment proceeds as if nothing happened. There is no auto-revert switch and no "enforce" mode. So what do you do with a `MODIFIED` result? You make a decision, and there are only two honest options. **Make reality match the code.** Somebody made a change that should not stand, so you put the resource back. Beware of the obvious-looking move here: re-applying the same template does not do it. A stack update diffs the *new* template against the *stored* one, not against live resources — if the template has not changed, CloudFormation answers `No updates are to be performed` and the drift survives. You need an update that genuinely touches the drifted property, or you fix the resource directly. **Make the code match reality.** The change was legitimate and should have gone through the pipeline. Edit the template to describe the new value and update the stack, so the recorded expectation and the live resource agree again. What you must not do is leave it. A drifted resource means the template no longer describes production, which quietly breaks the promise the whole tool rests on: that you can redeploy this stack somewhere else and get the same thing. ## On demand, not continuous CloudFormation does not sit and watch. Drift is evaluated only when a detection run is started — from the console, from the API, or from something you schedule yourself. A stack that says `IN_SYNC` is telling you about the moment detection ran, not about now. Somebody can drift a resource thirty seconds after a clean report and nothing will notice until the next run. That is why teams that care about drift schedule detection rather than clicking it: a periodic job that starts detection across the estate and alerts on the drifted count. ## What it does not look at Drift detection is scoped to the resources the stack manages. It says nothing about resources someone created *next to* the stack — a hand-made security group in the same VPC is invisible to it, because the stack never claimed to own it. It also only looks at configuration, never at data: bucket contents, table rows and log entries are not properties, so they are not drift. And its coverage inside the stack is partial — drift support exists per resource type and per property, which is why `NOT_CHECKED` shows up constantly and why a clean report is weaker evidence than it looks.

  • A resource comes back MODIFIED and you want the template's value to win. Why is re-running the same stack update usually not enough?
    Because a CloudFormation update diffs the submitted template against the stack's stored template, not against live resources. If the template is unchanged there is nothing to do and the API answers `No updates are to be performed`, leaving the drift in place. You need an update that actually changes the drifted property — or you correct the resource directly and re-run detection to confirm.
  • Does drift detection tell you about a resource somebody created by hand alongside the stack?
    No. Detection only walks resources the stack manages, comparing each against its template entry. A security group created by hand in the same VPC was never claimed by the stack, so it does not appear in the report at any status. Finding unmanaged resources is a separate inventory or account-compliance problem, not something a stack drift report answers.
  • Is a stack update blocked or warned about if the stack is currently drifted?
    No. CloudFormation does not consult drift status before an update, and an update is not a drift check. It computes changes from the template diff and applies them, which can silently overwrite a hand edit — or leave it untouched if the update never touches that property. If drift matters to your pipeline, run detection explicitly before deploying.

saying these in an interview costs you the question

  • Says drift detection automatically reverts the resource to the template
  • Believes CloudFormation monitors stacks for drift continuously
  • Assumes IN_SYNC means every property of every resource was checked
  • Confuses a drift report with a change set preview of a pending update
  • Thinks re-applying the same template always corrects a drifted property

context

open as a page

In an AWS CloudFormation template, what is the difference between Ref and Fn::GetAtt when applied to a resource, and how do you know which one gives you that resource's ARN?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Ref returns one default value chosen by the resource type — usually its name or physical ID — while Fn::GetAtt returns a specific named attribute of that resource. Only the resource type's documentation says which value each one produces.

open as a page

In AWS CloudFormation, what is a change set, and which part of the described change set tells you a resource will be destroyed and recreated rather than updated in place?

level: middleimportance: must knowfreq 68%

basics

~20 s

A change set is a stored, named preview of what an AWS CloudFormation stack update would do, computed without applying anything. Its Replacement field is the thing to read: True means the resource is deleted and recreated with a new physical ID.

open as a page

An AWS CloudFormation stack update fails and the stack ends up in UPDATE_ROLLBACK_FAILED. What does that status actually mean, and how do you get the stack back to a state where you can deploy again?

level: seniorimportance: must knowfreq 52%

basics

~20 s

UPDATE_ROLLBACK_FAILED means CloudFormation could not restore the stack's previous state, so it refuses further updates. Fix whatever blocked the rollback, then call ContinueUpdateRollback — skipping unrecoverable resources only as a last resort, since that leaves records inaccurate.

open as a page

Your first `aws cloudformation create-stack` call fails and the stack now shows the status ROLLBACK_COMPLETE. Why can you neither retry the create nor update it, and what do you do next?

level: juniorimportance: should knowfreq 58%

basics

~10 s

ROLLBACK_COMPLETE means the creation failed and CloudFormation deleted everything it had made, leaving an empty stack record that only accepts deletion. Delete the stack, fix what the failure event reported, and create it again.

open as a page

You start drift detection on an AWS CloudFormation stack and the call returns immediately with only an ID. Which calls do you make to get the per-resource result, and what do the resource statuses IN_SYNC, MODIFIED, DELETED and NOT_CHECKED tell you?

level: middleimportance: should knowfreq 55%

basics

~10 s

DetectStackDrift is asynchronous: it returns a StackDriftDetectionId that you poll with DescribeStackDriftDetectionStatus, then read DescribeStackResourceDrifts for per-resource statuses and the expected-versus-actual property differences.

open as a page

A DynamoDB table was created by hand and now has to be managed by an existing AWS CloudFormation stack, with no downtime and no data loss. How do you bring it under the stack's control, and what does that operation require?

level: middleimportance: should knowfreq 45%

basics

~20 s

Use a CloudFormation import: create a change set with type IMPORT, supply a template entry for the resource carrying a DeletionPolicy, and map its logical id to the live resource's physical identifier. Executing it adopts the table without recreating it.

open as a page

In an AWS CloudFormation template, how do you make a whole resource and a single property exist only in production, and what happens to that resource when the condition later evaluates to false?

level: middleimportance: should knowfreq 42%

basics

~20 s

Declare a named condition in the Conditions section, attach it to a resource with the Condition attribute, and use Fn::If with AWS::NoValue to include or omit a single property. If an existing resource's condition turns false on update, CloudFormation deletes that resource.

open as a page

Given that Ref and Fn::GetAtt already order resource creation in an AWS CloudFormation template, when do you still need an explicit DependsOn attribute?

level: middleimportance: should knowfreq 58%

basics

~20 s

You need DependsOn when a resource must exist before another for a reason CloudFormation cannot see because no property references it. The classic cases are an internet gateway attachment and an IAM policy the workload needs at boot but never references.

open as a page

In an AWS CloudFormation template, why would you use Fn::Sub rather than Fn::Join, and how do you include a literal ${...} in the string that CloudFormation must not substitute?

level: middleimportance: should knowfreq 55%

basics

~10 s

Fn::Sub interpolates ${...} placeholders directly inside a string, which is far more readable than assembling fragments with Fn::Join. A literal dollar-brace is escaped by writing ${!Literal}, which renders as ${Literal} untouched.

open as a page

An AWS CloudFormation drift detection run reports a stack IN_SYNC, but you have good reason to believe someone changed one of its resources in the console. Give the reasons CloudFormation can miss a real change.

level: seniorimportance: should knowfreq 42%

basics

~20 s

CloudFormation drift coverage is partial: support exists per resource type and per property, unsupported resources come back NOT_CHECKED, and only values the template explicitly sets are compared. A run is also a point-in-time snapshot, so IN_SYNC is weak evidence, not proof.

open as a page

A production database was defined in an AWS CloudFormation template with DeletionPolicy: Retain, and an update destroyed it anyway. What is the difference between DeletionPolicy and UpdateReplacePolicy, and what should have been set?

level: seniorimportance: should knowfreq 45%

basics

~20 s

DeletionPolicy only covers a resource leaving the stack or the stack being deleted. An update that replaces a resource deletes the old one under UpdateReplacePolicy, which defaults to Delete. Stateful resources need both attributes set to Retain or Snapshot.

open as a page

A CloudFormation networking stack exports a subnet ID that three application stacks consume with Fn::ImportValue. What does that coupling stop you from doing later, and how would you loosen it?

level: seniorimportance: should knowfreq 46%

basics

~20 s

An imported export becomes immovable: CloudFormation refuses to delete the exporting stack or to change that output's value while any import exists. Loosen it by publishing the value to SSM Parameter Store and reading it as a template parameter instead.

open as a page

A production RDS database currently sits in one AWS CloudFormation stack and needs to move to a different stack, without being deleted or recreated. Walk through the steps, and say what happens if you get the order wrong.

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Set DeletionPolicy Retain on the resource and update the source stack, then remove it from that template so the stack releases it while the database survives, then adopt it into the target stack with an import change set and verify with drift detection.

open as a page

What does the line `Transform: AWS::Serverless-2016-10-31` at the top of a CloudFormation template do, and how does it change the way that template is deployed?

level: seniorimportance: nice to knowfreq 36%

basics

~10 s

It declares the SAM macro, so CloudFormation expands the template's AWS::Serverless::* resources into plain resources on the service side before deploying. Deployments go through a change set and must pass the CAPABILITY_AUTO_EXPAND capability.

open as a page

In AWS CloudFormation, what problem does a nested stack solve and what problem does a StackSet solve, and how would you choose between them when designing deployments for a large estate?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Nested stacks decompose one large deployment into child stacks that deploy and roll back as a single unit inside one account and region. StackSets push one template out as stack instances across many accounts and regions. They answer different questions and often combine.

open as a page