Your first `aws cloudformation create-stack` call fails and the stack now shows the status ROLLBACK_COMPLETE. Why can you neither retry the create nor update it, and what do you do next?
answer
- failed create, everything undone
- a record owning nothing
- only one operation permitted
- earliest failure event is the cause
- on-failure keeps the evidence
basics
~10 sROLLBACK_COMPLETE means the creation failed and CloudFormation deleted everything it had made, leaving an empty stack record that only accepts deletion. Delete the stack, fix what the failure event reported, and create it again.
solid answer
~50 sOn a failed `create-stack`, CloudFormation's default behaviour is to roll back — it deletes the resources it managed to create and parks the stack in `ROLLBACK_COMPLETE`. What is left is a record with no resources behind it, and the only operation CloudFormation accepts on a stack in that status is `delete-stack`; an update is rejected, and a `create-stack` with the same name collides with the existing record. So the loop is: read `describe-stack-events`, find the earliest `CREATE_FAILED` event and its status reason — that is the root cause, and the later events are just cascade — fix the template or the underlying problem, delete the stack, and create it again. If I expect to iterate, I create with `--on-failure DO_NOTHING` so the failed resources stay in place and I can look at the one that actually broke instead of only reading its error string.
code
bash · 7 linesaws cloudformation create-stack --stack-name payments-dev \
--template-body file://template.yaml \
--on-failure DO_NOTHING
aws cloudformation describe-stack-events --stack-name payments-dev \
--query 'reverse(StackEvents[?ResourceStatus==`CREATE_FAILED`])[0].[LogicalResourceId,ResourceStatusReason]' \
--output textgo deeper
Remember that a stack in ROLLBACK_COMPLETE owns nothing and only accepts deletion — delete it, fix the template, create again. Know to look in describe-stack-events for the reason.
Explain why the create path has no previous state to return to, so it differs from a failed update. Describe the --on-failure options and when you would keep the failed resources rather than roll back.
Show how you diagnose without guessing: earliest CREATE_FAILED event, read the status reason literally, distinguish root cause from cascade. Call out tombstone stacks blocking a pipeline's next run and how you stop them accumulating.
Set the conventions: stack naming and ownership so tombstones are attributable, a pipeline failure path that always cleans up, and a rule about which environments may use DO_NOTHING given that abandoned resources keep billing.
## What the status means CloudFormation treats a stack creation as all-or-nothing by default. If any resource fails to create, it enters `ROLLBACK_IN_PROGRESS` and deletes the resources it had already made, ending in `ROLLBACK_COMPLETE`. That status is a tombstone: the stack record exists, holds the template and the event history, and owns nothing at all. Because there is nothing to update, CloudFormation permits exactly one operation on a stack in `ROLLBACK_COMPLETE` — deleting it. `update-stack` returns an error. Creating a stack with the same name fails too, because the name is still taken by the tombstone. This is why the first CloudFormation experience of many engineers is a confusing loop of "my fix will not apply": the fix is fine, the stack is simply not eligible to receive it. So the recovery is mechanical: ```bash aws cloudformation delete-stack --stack-name payments-dev aws cloudformation wait stack-delete-complete --stack-name payments-dev aws cloudformation create-stack --stack-name payments-dev --template-body file://template.yaml ``` Note the contrast with a failed *update*: an update failure rolls back to the previous template and lands in `UPDATE_ROLLBACK_COMPLETE`, where the stack is still alive and updatable. Only a failed *create* produces a stack you must throw away, because there is no previous state to return to. ## Finding out why it failed The status alone tells you nothing about the cause, and by the time you look the resources are gone. The evidence is in the events: ```bash aws cloudformation describe-stack-events --stack-name payments-dev \ --query 'reverse(StackEvents[?ResourceStatus==`CREATE_FAILED`])[0].[LogicalResourceId,ResourceStatusReason]' \ --output text ``` Two habits matter here. First, events come back newest first, so the *earliest* `CREATE_FAILED` is the root cause; everything after it is either cascade failures or the rollback deleting things. Second, read `ResourceStatusReason` literally — it is usually an exact service error: an AccessDenied naming the API call, a quota message, an invalid parameter value, a name already in use, or `WaitCondition timed out` when something was supposed to signal back and never did. Common culprits on a first create: the deploying principal lacks a permission the template needs; `CAPABILITY_IAM` or `CAPABILITY_NAMED_IAM` was not acknowledged for a template that creates roles; a hard-coded resource name already exists in the account; an AMI, instance type or engine version is not available in the target region. ## Keeping the wreckage for inspection The error string is not always enough — sometimes you need to look at the half-built resource itself. `create-stack` accepts `--on-failure` with `ROLLBACK` (the default), `DELETE`, or `DO_NOTHING`: - `ROLLBACK` — delete created resources, stack ends in `ROLLBACK_COMPLETE`. - `DELETE` — remove the stack record entirely, leaving nothing to clean up. Handy in CI where an abandoned tombstone would block the next run. - `DO_NOTHING` — leave everything that was created in place; the stack ends in `CREATE_FAILED` so you can inspect the broken resource, its logs, and its configuration. Delete the stack when you are done, because those resources are still running and still billing. That trade is worth stating out loud in an interview: `ROLLBACK` gives you a clean account and a thin error message; `DO_NOTHING` gives you the crime scene and a cleanup obligation. ## The neighbouring failure states Two more statuses belong to this family and are worth recognising: - `ROLLBACK_FAILED` — the create rolled back and the rollback itself hit a problem. As with a stuck update rollback, the stack needs the blocker cleared before it can be finished or deleted. - `DELETE_FAILED` — a delete could not remove some resource, usually because it is not empty or is still referenced. `delete-stack` accepts `--retain-resources` naming the logical IDs to give up on, which lets the stack record be removed while leaving those resources orphaned in the account for you to deal with by hand. ## Why this is worth knowing early A stack in `ROLLBACK_COMPLETE` costs nothing and breaks nothing, but it blocks the name and confuses anyone who assumes a stack record implies running infrastructure. In a shared development account they accumulate. In a CI pipeline they are worse: a job that creates a stack and fails leaves a tombstone that makes the *next* run fail for an entirely different and misleading reason. Either delete the stack in the pipeline's failure path or create with `--on-failure DELETE` so no tombstone ever remains.
- How does a failed update differ from a failed create in what it leaves behind?A failed update rolls back to the stack's previous template and lands in UPDATE_ROLLBACK_COMPLETE — the stack is alive, owns its old resources, and accepts the next update. A failed create has no previous state to return to, so it deletes everything and leaves a tombstone in ROLLBACK_COMPLETE that can only be deleted.
- Why is a leftover ROLLBACK_COMPLETE stack a particular nuisance in a CI pipeline?The name is taken, so the next run's create-stack fails for a reason unrelated to the actual change, sending whoever is on call chasing a phantom. Either delete the stack in the failure path, or create with `--on-failure DELETE` so a failed run leaves no record behind.
- What does delete-stack do if a resource refuses to be deleted?The stack ends in DELETE_FAILED — commonly a non-empty S3 bucket or a resource still referenced elsewhere. Clear the blocker and retry, or rerun delete-stack with `--retain-resources` naming those logical IDs, which removes the stack record and leaves those resources orphaned in the account for you to handle manually.
saying these in an interview costs you the question
- Tries update-stack to apply the fix to a rolled-back stack
- Assumes ROLLBACK_COMPLETE means resources are still running
- Reads the last event instead of the first failure
- Deletes and recreates repeatedly without reading the status reason
- Thinks a failed update also has to be deleted and recreated