An AWS CloudFormation stack update fails and the stack ends up in UPDATE_ROLLBACK_FAILED. What does that status actually mean, and how do you get the stack back to a state where you can deploy again?
answer
- the safety net itself failed
- stack refuses further updates
- two operations still accepted
- repair by hand, then continue
- skipping costs record accuracy
basics
~20 sUPDATE_ROLLBACK_FAILED means CloudFormation could not restore the stack's previous state, so it refuses further updates. Fix whatever blocked the rollback, then call ContinueUpdateRollback — skipping unrecoverable resources only as a last resort, since that leaves records inaccurate.
solid answer
~50 sWhen an update fails, CloudFormation tries to put every resource back the way it was. If one of those restore steps itself fails — usually because someone deleted or altered the resource by hand, a bucket is no longer empty, or the stack's role lost a permission mid-flight — the rollback cannot finish and the stack lands in `UPDATE_ROLLBACK_FAILED`. In that status CloudFormation accepts no `update-stack`; the only ways forward are `ContinueUpdateRollback` or deleting the stack. So I read `describe-stack-events` to find the first failed resource and its status reason, repair that resource by hand so the rollback step can succeed, and call `aws cloudformation continue-update-rollback`. If a resource genuinely cannot be restored, I rerun it with `--resources-to-skip` naming that logical ID — which unblocks the stack but leaves CloudFormation's record wrong for it, so reconciling that resource afterwards is mandatory, not optional.
code
bash · 8 linesaws cloudformation describe-stack-events --stack-name payments-prod \
--query 'StackEvents[?ResourceStatus==`UPDATE_FAILED` || ResourceStatus==`UPDATE_ROLLBACK_FAILED`].[LogicalResourceId,ResourceStatusReason]' \
--output text
aws cloudformation continue-update-rollback --stack-name payments-prod
aws cloudformation continue-update-rollback --stack-name payments-prod \
--resources-to-skip AppDatabase LoggingStack.LogGroupgo deeper
Recognise that a stack can get stuck in a state where CloudFormation refuses new updates, and that the fix is a specific continue-rollback operation rather than retrying the deploy.
Explain the two-phase update and why the rollback can fail — a deleted resource, a non-empty bucket, a lost permission. Name continue-update-rollback and the resources-to-skip option and say what each does.
Show the diagnosis first: read the earliest failed event's status reason, repair the blocker by hand, then continue. Be explicit that skipping is a last resort that leaves CloudFormation's record wrong and creates work you must finish.
Own the prevention side: mandatory change-set review, a dedicated stack role so rollback never depends on a human's credentials, a ban on console edits to managed resources, and a documented runbook so an on-call engineer does not improvise on a wedged production stack.
## How a stack gets there A CloudFormation stack update is a two-phase operation with an automatic safety net. Phase one applies the new template. If any resource fails, CloudFormation stops and enters `UPDATE_ROLLBACK_IN_PROGRESS`, walking back every change it already made so the stack returns to the template it had before. The happy ending of that path is `UPDATE_ROLLBACK_COMPLETE` — a stack that is exactly where it started, updatable again. `UPDATE_ROLLBACK_FAILED` is what you get when the safety net itself tears. A rollback step is just another API call against a real resource, and it can fail for ordinary reasons: - Somebody deleted or renamed a stack-managed resource in the console, so there is nothing to restore. - A resource cannot be returned to its old configuration — an S3 bucket that must be emptied first, a security group still referenced by something outside the stack, a subnet with an ENI left in it. - The IAM identity or stack role executing the update lost a permission partway through, or a policy changed under it. - A resource that waits for a signal or a stabilisation timeout never reports back, and the rollback times out too. - A nested child stack failed its own rollback, which fails the parent's. The important operational point: the workload is still running. `UPDATE_ROLLBACK_FAILED` is not an outage by itself. It is a stack that is half-way between two templates and that CloudFormation refuses to touch further, which means every subsequent deploy is blocked until it is cleared. ## What CloudFormation will accept in that state Exactly two things: `ContinueUpdateRollback`, and `DeleteStack`. `update-stack` is rejected. `cancel-update-stack` is not applicable — that only works while an update is still in progress. Creating a change set against the stack does not help either, because you cannot execute it. ## The recovery, in order **1. Find the actual failure.** `describe-stack-events` returns newest first; scroll back to the first `UPDATE_FAILED` or `UPDATE_ROLLBACK_FAILED` event and read `ResourceStatusReason`. Later events are usually cascades from that one. This is the whole diagnosis — the reason string almost always names the problem ("The bucket you tried to delete is not empty", "resource does not exist", an AccessDenied). **2. Repair the blocker by hand.** This is the part that feels wrong to engineers who believe infrastructure should never be touched manually, and it is nonetheless the correct move. If a resource was deleted, recreate something CloudFormation can find at that identifier. If a bucket is not empty, empty it. If a permission was removed, put it back. You are not fixing the infrastructure; you are removing the obstacle to the rollback step. **3. Continue the rollback.** ```bash aws cloudformation continue-update-rollback --stack-name payments-prod ``` CloudFormation retries the remaining rollback steps from where it stopped. If it succeeds you reach `UPDATE_ROLLBACK_COMPLETE` and the stack is deployable again. **4. Skip only what cannot be recovered.** When a resource genuinely cannot be restored, name its logical ID in `--resources-to-skip` (nested-stack resources are addressed as `NestedStackName.LogicalId`): ```bash aws cloudformation continue-update-rollback --stack-name payments-prod \ --resources-to-skip AppDatabase LoggingStack.LogGroup ``` CloudFormation marks those resources as if their rollback had succeeded and finishes. That unblocks the stack, and it also means CloudFormation's recorded state for those resources no longer matches what exists. You have deliberately traded accuracy for a working stack. The debt is real: the next update computes its diff from a record you know is wrong. Pay it immediately, either by making the live resource match what the stack believes or by updating the template to describe what is actually there. **5. If none of that works,** the remaining option is deleting the stack and recreating it — viable in a lower environment, rarely acceptable in production unless the stateful resources are protected by retaining policies and can be brought back into a fresh stack. ## Reducing how often this happens - Review a change set before executing, so fewer updates fail in the first place. - Stop people editing stack-managed resources by hand; that single habit causes most unrecoverable rollbacks. - Give the stack its own role with `--role-arn` so the rollback does not depend on whichever human or pipeline credential started the update still having permissions. - Since 2021 you can pass `--disable-rollback` on `update-stack`. The failed update then stops in `UPDATE_FAILED` with the partially updated resources left alone, so you can inspect them; from there you either continue rolling back or call `rollback-stack` when you are ready. It converts a panicked automatic rollback into a decision you make with evidence. - Attaching CloudWatch alarms as rollback triggers via `--rollback-configuration`, with `MonitoringTimeInMinutes`, makes CloudFormation watch for a set period after the update and roll back if an alarm fires — worth knowing, and worth remembering that it adds another way a rollback can start.
- What exactly do you owe the stack after using --resources-to-skip?Reconciliation. CloudFormation marked those resources as rolled back without touching them, so its recorded template no longer describes reality for them. The next diff is computed from that wrong record. Either change the live resource to match what the stack believes, or amend the template to describe what actually exists — before the next deploy, not eventually.
- Why does giving the stack its own IAM role with --role-arn make rollbacks more reliable?Without one, CloudFormation acts with the credentials of whoever started the update. If that pipeline credential's permissions change, or a different principal is involved later, a rollback step can fail with AccessDenied. A dedicated stack role gives every phase — update and rollback alike — one stable, sufficient permission set.
- How does --disable-rollback on update-stack change what you are dealing with?Added in 2021, it stops a failed update in UPDATE_FAILED with the partially applied resources left in place instead of automatically unwinding. You inspect the actual broken resource, then choose: continue rolling back, or call rollback-stack when ready. It trades automatic tidiness for evidence, which is what you want on a hard-to-reproduce failure.
saying these in an interview costs you the question
- Says just run update-stack again with a corrected template
- Thinks the stack is down because the status says failed
- Reaches for --resources-to-skip before diagnosing anything
- Treats skipped resources as fully resolved afterwards
- Claims deleting and recreating the stack is the only option