skip to content

A production RDS database currently sits in one AWS CloudFormation stack and needs to move to a different stack, without being deleted or recreated. Walk through the steps, and say what happens if you get the order wrong.

level: seniorimportance: nice to knowfreq 28%

answer

  1. one resource, one stack
  2. retain first, remove second
  3. separate deployments, not one edit
  4. import adopts, it does not create
  5. detect drift before the next update

basics

~20 s

Set DeletionPolicy Retain on the resource and update the source stack, then remove it from that template so the stack releases it while the database survives, then adopt it into the target stack with an import change set and verify with drift detection.

solid answer

~50 s

CloudFormation will not let two stacks own the same physical resource, so the source stack has to let go before the target can adopt. The sequence is: **first**, add `DeletionPolicy: Retain` to the database in the source template and update the stack — this changes only stack metadata and touches nothing live. **Second**, remove the resource from the source template and update again; because of the retain policy, CloudFormation stops managing it and leaves the database running. **Third**, add it to the target stack's template (also with a `DeletionPolicy`) and adopt it with a change set of type `IMPORT`, mapping the logical id to the instance's identifier. **Fourth**, run drift detection on the target stack, because import never checks that your template matches the live database. Reverse the first two steps and the outcome is brutal: removing the resource from the template without the retain policy in place deletes the production database.

code

yaml · 10 lines
yaml
Resources:
  OrdersDb:
    Type: AWS::RDS::DBInstance
    DeletionPolicy: Retain
    UpdateReplacePolicy: Retain
    Properties:
      DBInstanceIdentifier: orders-prod
      Engine: postgres
      DBInstanceClass: db.m6g.large
      AllocatedStorage: 200

go deeper

for a junior

Know that a resource belongs to exactly one stack and that removing it from a template normally deletes it — so moving one is a deliberate procedure, not an edit.

for a middle

Be able to name the four steps in order and explain why the retain policy has to be applied and deployed before the resource is removed from the source template.

for a senior

Show the operational care: two separate deployments, an explicit check that the first landed, a short unmanaged window under change control, and drift detection before any ordinary deployment touches the imported resource.

for a principal

Decide whether the move is worth it at all. Restructuring stacks around stateful resources carries irreversible risk, so weigh it against leaving the database where it is, and define what change control such a procedure requires in your organisation.

## The constraint that shapes everything A physical resource can belong to exactly one CloudFormation stack. There is no move operation, and you cannot import a resource that another stack still manages. So the move is really two halves: the source stack must **release** the database while leaving it alive, and the target stack must **adopt** it. The order of those halves is the whole question, because the naive release deletes production. ## Step 1 — make the retain policy real first Add `DeletionPolicy: Retain` to the resource in the source template and update the stack, as its own deployment, with no other changes: ```yaml OrdersDb: Type: AWS::RDS::DBInstance DeletionPolicy: Retain UpdateReplacePolicy: Retain Properties: DBInstanceIdentifier: orders-prod # ...unchanged ``` This is metadata about how CloudFormation should behave, so the update makes no change to the running database. Do it as a separate deployment and confirm it succeeded — the entire safety of the next step rests on this one having landed. `UpdateReplacePolicy: Retain` is worth adding at the same time for the same reason. ## Step 2 — release from the source stack Now delete the resource block from the source template and update the stack again. Normally, removing a resource from a template tells CloudFormation to delete it. With `Retain` recorded, CloudFormation instead stops managing it and leaves it in place: the stack events show the resource leaving the stack, and the database keeps serving traffic. Anything else in the source stack that referenced it — an output, a `Ref` from a security group rule — has to be untangled in the same change, or the update fails. At the end of this step the database is running and **unmanaged**. That window is the risky part of the procedure: no stack describes it, so nothing will recreate it and nobody's template review protects it. Keep the window short and do not let anything else deploy into that account in the middle of it. ## Step 3 — adopt into the target stack Add the resource to the target template, carrying a `DeletionPolicy` because import requires one, then create a change set with `--change-set-type IMPORT` and a `--resources-to-import` entry mapping the logical id to the live instance's identifier: ```json [{"ResourceType": "AWS::RDS::DBInstance", "LogicalResourceId": "OrdersDb", "ResourceIdentifier": {"DBInstanceIdentifier": "orders-prod"}}] ``` Execute it. The stack goes to `IMPORT_COMPLETE` with nothing created, replaced or deleted. ## Step 4 — verify with drift detection Import records your template as the expected configuration without comparing it against the live database. If the template you wrote in the target stack differs from reality — a different instance class, storage size, backup retention, parameter group — the stack now holds a wrong expectation, and the first ordinary update that touches the resource can apply it for real, which for a database means a modification or even a replacement. So run detection on the imported resource and reconcile every `MODIFIED` property in the template before letting any normal deployment near the stack: ```bash aws cloudformation detect-stack-resource-drift \ --stack-name orders-prod-v2 \ --logical-resource-id OrdersDb ``` The cheapest way to get the template right in the first place is to read the live configuration from the service and write the template from that, rather than copying the source stack's version and hoping nothing drifted while it lived there. ## Failure modes, in order of severity **Removing before retaining.** The update deletes the database. This is the reason the two source-stack updates are separate deployments rather than one clever template edit — combining them is exactly where people get it wrong, and there is no undo. **Importing before releasing.** Rejected: the resource still belongs to the source stack. Harmless, but it tells you the order is fixed, not a preference. **A wrong template at import time.** Silent. Nothing fails, and the damage appears later during an unrelated deployment. Drift detection is the only thing standing between you and that. **Deleting the source stack to release the resource.** Works only if every retained resource genuinely has a retain policy, and takes everything else in the stack with it. Prefer removing the single resource from the template. ## Say this in the interview Retain, release, import, detect. Two source-stack updates, deliberately separate, and a drift check as the last step rather than the optional one.

  • Why must the retain policy and the removal be two separate stack updates?
    Because the retain policy has to be recorded in the stack before CloudFormation processes a removal. Submitting a template that both drops the resource and would have added the policy leaves nothing recorded to protect it, and the update deletes the database. Splitting them costs one extra deployment and turns an irreversible mistake into an ordinary metadata change you can verify before proceeding.
  • What is the risk in the window between the source stack releasing the resource and the target stack importing it?
    The database is alive but unmanaged: no stack declares it, so no template review, no drift detection and no recreation path covers it. A concurrent cleanup, a mistaken manual delete, or simply forgetting to finish the procedure leaves production depending on an orphan. Keep the window to minutes, do it under change control, and confirm the import before walking away.
  • Could you instead delete the source stack to release the database?
    Only if every resource you need to survive carries a retain policy, and only if losing the rest of the stack is acceptable — which for a shared stack it rarely is. Removing the single resource from the template is the narrower, reversible-in-practice move. Deleting a stack to release one resource is how people discover which of its other resources had no retain policy.

saying these in an interview costs you the question

  • Removes the resource from the template before setting Retain
  • Assumes CloudFormation has a move-resource-between-stacks operation
  • Tries to import a resource another stack still manages
  • Skips drift detection after importing into the target stack
  • Believes import applies the target template to the live database

context