In AWS CloudFormation, what problem does a nested stack solve and what problem does a StackSet solve, and how would you choose between them when designing deployments for a large estate?
answer
- different questions, not alternatives
- one deployment's internal shape
- how many places it lands
- children share a blast radius
- targets follow the org structure
basics
~20 sNested stacks decompose one large deployment into child stacks that deploy and roll back as a single unit inside one account and region. StackSets push one template out as stack instances across many accounts and regions. They answer different questions and often combine.
solid answer
~50 sA nested stack is a resource of type `AWS::CloudFormation::Stack` inside a parent template, pointing at a child template in S3. It solves size and reuse: one deployment that is too big or too repetitive for a single template, split into components while remaining one unit of deploy and rollback — which is also its risk, since a child failure rolls the whole tree back. A StackSet solves multiplicity: the same template delivered as stack instances into many accounts and many regions from one operation, with per-instance parameter overrides and rollout controls like `FailureToleranceCount`, `MaxConcurrentCount` and `RegionConcurrencyType`. With the service-managed permission model it targets Organizations OUs and, with automatic deployment enabled, covers new accounts as they join. So I use nested stacks for the shape of one application and a StackSet for org-wide baselines — and often both, the StackSet delivering a template that is itself nested.
code
bash · 12 linesaws cloudformation create-stack-set \
--stack-set-name org-baseline \
--template-body file://baseline.yaml \
--permission-model SERVICE_MANAGED \
--auto-deployment Enabled=true,RetainStacksOnAccountRemoval=false \
--capabilities CAPABILITY_NAMED_IAM
aws cloudformation create-stack-instances \
--stack-set-name org-baseline \
--deployment-targets OrganizationalUnitIds=ou-abcd-11111111 \
--regions eu-west-1 us-east-1 \
--operation-preferences FailureToleranceCount=0,MaxConcurrentCount=1,RegionConcurrencyType=SEQUENTIALgo deeper
Know the plain distinction: a nested stack is a child stack created by a parent template, while a StackSet deploys one template into many accounts and regions at once.
Explain the mechanics — AWS::CloudFormation::Stack with a TemplateURL, parameters down and outputs up, versus stack instances created per account and region — and that a nested child failure rolls back the whole parent.
Show rollout judgment: failure tolerance, concurrency, sequential regions, a canary OU, and per-instance parameter overrides. Say what a StackSet operation does not do, which is undo instances that already succeeded.
Own the estate decision. Weigh the self-managed versus service-managed permission models with delegated administration, and argue honestly whether a native StackSet or a per-account deployment pipeline better fits the organisation's existing review and rollback discipline.
## They are not alternatives The question is asked as a comparison because the names sound similar, and the useful answer starts by refusing the framing. Nested stacks address the internal structure of one deployment. StackSets address how many places a deployment lands. You can want either, both, or neither. ## Nested stacks: decomposition with one blast radius A nested stack is an ordinary resource, `AWS::CloudFormation::Stack`, whose `TemplateURL` points at a child template stored in S3. CloudFormation creates a real, separate stack for each child and manages it as part of the parent's operation. Parameters flow down into children, and children hand values back up through their outputs. What it buys you: - **Size.** CloudFormation caps how many resources one template may contain, and a large system hits it. Nesting is the supported way past it. - **Reuse.** A network layer, a standard service scaffold, a monitoring bundle — written once, instantiated repeatedly with different parameters within the same deployment. - **Readability.** A 4,000-line template nobody reviews properly becomes a parent that reads like a table of contents. What it costs you: - **One unit of failure.** The tree deploys and rolls back together. A failure in one child rolls back the parent, which rolls back the siblings that had already succeeded. That atomicity is a feature when the pieces are genuinely one system and a liability when they are not. - **Operational coupling.** You update the parent, never a child directly; updating a child stack out of band puts the parent's record out of step with reality. - **Weaker default visibility.** A change set on the parent does not expand the children unless you ask it to with `IncludeNestedStacks`, so a replacement hiding inside a child can pass review unseen. - **Harder failure diagnosis.** The parent reports that a child stack failed; the actual reason is in the child's own event stream. The alternative to nesting is separate top-level stacks wired together by cross-stack references. That inverts the trade: independent deploys and independent blast radius, at the price of losing atomic rollback and gaining rigid coupling, because an exported value cannot change while another stack imports it. ## StackSets: one template, many targets A StackSet holds a template and a target set. From one operation, CloudFormation creates a *stack instance* — a real stack — in each target account and region. Update the StackSet and it propagates. Two permission models: - **Self-managed** — you create an `AWSCloudFormationStackSetAdministrationRole` in the admin account and an `AWSCloudFormationStackSetExecutionRole` in each target account that trusts it. Works without AWS Organizations, and every new target account is manual setup. - **Service-managed** — integrated with AWS Organizations. You target organizational units rather than listing account IDs, CloudFormation handles the roles, and a delegated administrator account can run the StackSets instead of the management account. Enabling automatic deployment means an account that joins a target OU receives the stacks automatically, and one that leaves has them removed or retained by your choice. That last property is the real reason StackSets exist for platform teams: the baseline is attached to the organizational structure, not to a list somebody must remember to update. Rollout is controlled rather than simultaneous. Operation preferences carry `FailureToleranceCount` or `FailureTolerancePercentage` — how many instances may fail before the whole operation stops — plus `MaxConcurrentCount` or `MaxConcurrentPercentage`, `RegionOrder`, and `RegionConcurrencyType` set to `SEQUENTIAL` or `PARALLEL`. A sensible org-wide change goes out with a failure tolerance of zero, one region at a time, into a canary OU first. Per-instance parameter overrides let one template adapt to accounts that legitimately differ. The limits are worth naming too: a StackSet operation is not transactional across accounts — it stops on failure, it does not undo the instances that already succeeded. Drift between instances is real, because someone with admin in a member account can edit what the StackSet delivered. ## Choosing Ask what varies: - **The deployment is large or repetitive within one account and region** → nested stacks (or separate stacks, if you would rather have independent blast radius than atomic rollback). - **The same thing must exist in many accounts or regions** → a StackSet, service-managed if you run Organizations. - **Both** → common and correct: a StackSet distributing a security or logging baseline whose template is itself composed of nested stacks. And the honest counterweight for a platform decision: a StackSet couples the whole estate to one team's operation and to CloudFormation specifically. Many organisations instead run a pipeline that deploys ordinary stacks per account from version-controlled code, trading the built-in fan-out for the review, testing and rollback discipline they already have around a pipeline. That is the tradeoff a principal is expected to weigh out loud rather than defaulting to the AWS-native answer.
- What does enabling automatic deployment on a service-managed StackSet give you?The target set follows the organizational structure instead of a maintained list. An account that joins a targeted OU receives the stack instances without anyone acting; one that leaves has them removed, or retained if you set RetainStacksOnAccountRemoval. That is what makes a baseline a property of the organization rather than a spreadsheet somebody updates.
- Why is a StackSet operation not a transactional rollout across accounts?Each stack instance is a separate stack operation. Failure tolerance decides when CloudFormation stops attempting further instances, but instances that already succeeded stay updated — you end with a partially rolled-out estate. That is why sensible rollouts use a zero failure tolerance, low concurrency, sequential regions and a canary OU.
- When would you prefer separate top-level stacks over nesting?When the pieces have different change rates, owners or risk profiles. Separate stacks deploy independently and one failure does not roll back the others. You give up atomic rollback across the system and take on coupling through cross-stack references, where an exported value cannot change while another stack imports it.
saying these in an interview costs you the question
- Describes StackSets as a stack containing other stacks
- Assumes a StackSet rollout rolls back everywhere on failure
- Thinks a child nested stack should be updated directly
- Says nested stacks give independent blast radius per component
- Presents StackSets as strictly better than a deployment pipeline