skip to content

A production database was defined in an AWS CloudFormation template with DeletionPolicy: Retain, and an update destroyed it anyway. What is the difference between DeletionPolicy and UpdateReplacePolicy, and what should have been set?

level: seniorimportance: should knowfreq 45%

answer

  1. two attributes, two moments
  2. one covers stack deletion only
  3. replacement deletes the old one
  4. the other default is Delete
  5. policies salvage, stack policy refuses

basics

~20 s

DeletionPolicy only covers a resource leaving the stack or the stack being deleted. An update that replaces a resource deletes the old one under UpdateReplacePolicy, which defaults to Delete. Stateful resources need both attributes set to Retain or Snapshot.

solid answer

~50 s

They guard two different events and people assume one covers both. `DeletionPolicy` applies when the resource is removed from the template or the whole stack is deleted. `UpdateReplacePolicy` applies when an update forces a replacement: CloudFormation creates the new resource and then deletes the old one — and that deletion follows `UpdateReplacePolicy`, which defaults to `Delete` even if `DeletionPolicy` says `Retain`. So editing an immutable property such as `DBInstanceIdentifier` triggers a replacement and the original instance is destroyed exactly as the account owner did not intend. Both attributes take `Delete`, `Retain` or `Snapshot`, and on anything stateful I set both — `Snapshot` on databases and volumes that support it, `Retain` otherwise. I also read the change set first, because `Replacement: True` is the warning that these policies are about to matter, and I back them with a stack policy that denies `Update:Replace` on the resource so the replacement is refused rather than merely survived.

code

yaml · 13 lines
yaml
AWSTemplateFormatVersion: '2010-09-09'
Resources:
  AppDatabase:
    Type: AWS::RDS::DBInstance
    DeletionPolicy: Snapshot
    UpdateReplacePolicy: Snapshot
    Properties:
      DBInstanceIdentifier: app-prod
      Engine: postgres
      DBInstanceClass: db.t3.medium
      AllocatedStorage: '100'
      MasterUsername: appadmin
      ManageMasterUserPassword: true

go deeper

for a junior

Know that DeletionPolicy and UpdateReplacePolicy are resource attributes, not properties, and that a database in a CloudFormation template should have both set rather than relying on defaults.

for a middle

Explain the two distinct events — resource leaving the stack versus an update replacing it — and that UpdateReplacePolicy defaults to Delete. Name what forces a replacement, such as changing an immutable identifier.

for a senior

Demonstrate the operational judgment: Snapshot where supported, Retain otherwise, Delete for stateless resources, and a change set reviewed for Replacement: True before executing. Be honest that a snapshot means an outage and a manual restore.

for a principal

Own the guardrail layering: a stack policy that refuses replacement on protected resources, retaining policies as the net under it, real backups underneath both, and an automated check that no stateful resource type merges without these attributes set.

## Two attributes, two different moments Both are resource attributes rather than properties — they sit alongside `Type` and `Properties` in the resource definition, not inside it, which is itself a common source of silently-ignored configuration. **`DeletionPolicy`** answers: what happens to the real resource when CloudFormation is finished with it? That covers the stack being deleted, and the resource being removed from the template. **`UpdateReplacePolicy`** answers a narrower question: an update requires this resource to be replaced — CloudFormation creates the new one, cuts over, and then disposes of the old one. What happens to the old one? They are independent, and neither implies the other. `UpdateReplacePolicy` defaults to `Delete`. That default is the entire failure in this question: `DeletionPolicy: Retain` was set, everyone believed the database was protected, an update forced a replacement, and the replacement path — which `DeletionPolicy` does not govern — deleted the original. ## What forces a replacement Every resource type has properties that can be changed in place and properties that cannot. Renaming is the classic one: change `DBInstanceIdentifier` on an `AWS::RDS::DBInstance` and there is no API to rename it, so CloudFormation builds a new instance and deletes the old. The same applies broadly — an EC2 instance's AMI or subnet, a resource's name property, anything the service treats as immutable at creation. This is why the preview matters. A change set entry reports `Replacement: True` (or `Conditional`) for exactly these cases, and each detail carries a `RequiresRecreation` value of `Never`, `Conditionally` or `Always`. Reading that before executing is what turns "a database was destroyed" into "we caught it in review". ## The values - **`Delete`** — remove the resource. The default for almost every type. There is one notable exception: on `AWS::RDS::DBCluster` and on `AWS::RDS::DBInstance` resources that are members of a cluster, CloudFormation's default behaviour is to snapshot, which is why relying on defaults per type is a bad habit rather than a shortcut. - **`Retain`** — leave the resource alone. It survives, and it is no longer managed by any stack. That has costs of its own: an orphan nobody's template describes, still billing, and a name that will collide when someone recreates the stack. - **`Snapshot`** — take a final snapshot, then delete the resource. Only valid on types that support snapshots — RDS instances and clusters, EBS volumes, ElastiCache clusters and replication groups, Redshift clusters, Neptune and DocumentDB clusters among them. On any other type it is a template error, not a silent fallback. And a snapshot is not a running database: recovery is a manual restore that produces a new endpoint, so it protects the data while accepting the outage. - **`RetainExceptOnCreate`** — added in 2023, valid on `DeletionPolicy`. It behaves like `Retain` normally but allows the rollback of a *failed initial creation* to delete the resource. Without it, a create that fails and rolls back leaves retained resources orphaned, and the retry then collides with them. ```yaml AppDatabase: Type: AWS::RDS::DBInstance DeletionPolicy: Snapshot UpdateReplacePolicy: Snapshot Properties: Engine: postgres ``` ## Salvage versus refusal These policies are salvage: the destructive action still happens, and you are left with a retained resource or a snapshot. If what you want is for the destructive action not to happen at all, that is a stack policy — a JSON document set with `set-stack-policy` that denies `Update:Replace` or `Update:Delete` on specific logical IDs. CloudFormation then refuses the update outright, and an operator must temporarily override the policy to proceed. The mature setup uses both, because they fail differently. A stack policy is a gate that a human can lift under pressure at 3am; a retaining policy is what saves you when they lift it. Neither is a backup strategy — the database's own automated backups are, and a template attribute is not a substitute for them. ## Choosing in practice - Stateful resources — databases, volumes, buckets holding real data, log destinations: set both attributes explicitly. `Snapshot` where supported, otherwise `Retain`. - Stateless resources — load balancers, security groups, roles, functions: leave both at `Delete`. Retaining them just litters the account and blocks recreation with name collisions. - Environments matter: retaining everything in an ephemeral test account produces a slow accumulation of unowned, billing resources that nobody will admit to. `RetainExceptOnCreate` exists precisely because the naive setting bites hardest on repeated failed creates. - Whatever you choose, verify it the same way you verify anything else: a change set that reports `Replacement: True` on the resource you care about, reviewed before execution.

  • What is the downside of Retain that people discover afterwards?
    The resource survives but nothing manages it. It keeps billing, no template describes it, and its name is still taken — so recreating the stack fails on a collision, or worse, someone points the new stack at a leftover holding old data. Retain buys time to make a decision; it is not itself a decision.
  • How does a stack policy differ from these attributes as a guardrail?
    A stack policy denies the operation. Set with set-stack-policy, a statement denying Update:Replace on a logical ID makes CloudFormation refuse the update instead of performing it and salvaging the remains. DeletionPolicy and UpdateReplacePolicy are the net beneath that gate — you want both, because a gate can be lifted under pressure.
  • Why does Snapshot fail on some resource types rather than falling back to Retain?
    Snapshot is only defined for types with a snapshot API — RDS instances and clusters, EBS volumes, ElastiCache, Redshift, Neptune, DocumentDB. On anything else it is an invalid attribute value and the template is rejected. That is deliberate: a silent fallback would let you believe data was protected when nothing was capturing it.

saying these in an interview costs you the question

  • Assumes DeletionPolicy also covers replacement during updates
  • Sets Retain everywhere and calls it protection
  • Treats a snapshot as equivalent to a running database
  • Puts the attributes inside Properties, where they are ignored
  • Says these attributes replace real database backups

context