skip to content

Auto Scaling Groups and Launch Templates

An Auto Scaling group turns a launch template into a self-healing, elastic fleet by holding a desired capacity between min and max. I learn the policy types, the health-check source and the lifecycle hooks, because "how would you scale this tier?" is the standard EC2 design question.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

6

In an EC2 Auto Scaling group, what do the minimum, maximum and desired capacity settings each control, and what happens if you manually terminate one of the group's instances?

level: juniorimportance: must knowfreq 68%

answer

  1. three numbers, only one is live
  2. floor, ceiling, and current setpoint
  3. policies get clamped, never exceed
  4. kill one and it comes back
  5. desired is not lowered by manual termination

basics

~20 s

Desired capacity is the instance count the group tries to run right now; minimum and maximum are the floor and ceiling that scaling can never cross. Manually terminating an instance leaves the group below desired, so it launches a replacement.

solid answer

~50 s

An Auto Scaling group is a control loop, not a list of servers. `DesiredCapacity` is the live setpoint — the number of healthy instances the group is currently trying to run. `MinSize` and `MaxSize` are hard bounds: scaling policies, scheduled actions and manual updates all get clamped into that range, so a policy asking for 12 in a group with `MaxSize` 10 gets 10. If you terminate an instance yourself, desired capacity does not change — the group simply notices it is one short and launches a replacement from the launch template, which is exactly the self-healing behaviour you are buying. To genuinely remove an instance you lower desired capacity, or detach it with the flag that decrements desired at the same time. Setting min, max and desired to the same value gives you a fixed-size fleet that still replaces failures but never scales.

go deeper

for a junior

Be able to say plainly which of the three numbers the group is actively chasing, and that terminating an instance yourself gets you a replacement rather than a smaller fleet.

for a middle

Explain the clamping rules — what happens when a policy asks for more than the maximum, or when you lower the maximum below the current desired capacity — and how to remove an instance without triggering a backfill.

for a senior

Show that you size minimum capacity for survivability, not steady state: enough instances left after losing an Availability Zone, plus enough headroom that a spike does not have to wait on a cold boot.

for a principal

Own the guardrail argument. Maximum capacity is a blast-radius and spend control that has to be set high enough to absorb a real incident but low enough that a runaway scaling loop cannot bankrupt a team; explain how you pick and review that number.

## The group is a control loop, not a list of servers The most common beginner mistake is picturing an Auto Scaling group (ASG) as a folder containing servers you created. It is the opposite. The group holds a *target number* plus a launch template describing how to build one instance. A background loop continuously compares how many instances it currently considers healthy against that target, and issues launches or terminations until the two match. Every other Auto Scaling feature — scaling policies, health checks, lifecycle hooks — is either a way of moving that number or a way of deciding what counts as healthy. ## The three numbers - **`MinSize`** — the floor. The group will never voluntarily go below it, no matter what a scale-in policy asks for. - **`MaxSize`** — the ceiling. The group will never voluntarily go above it, no matter what a scale-out policy asks for. This is your blast-radius and budget guardrail. - **`DesiredCapacity`** — the live setpoint, the only one of the three that is actively acted on. It must sit between min and max. A useful mental model: min and max are the ends of a slider's track; desired is where the handle currently sits. ## Who moves desired capacity Several things write to `DesiredCapacity`, and they all get clamped: - **You**, via `update-auto-scaling-group` or the console. - **Dynamic scaling policies** (target tracking, step, simple) reacting to CloudWatch alarms. - **Scheduled actions**, which can set min, max and desired at a point in time. - **Predictive scaling**, which raises the effective minimum ahead of a forecast load. If a policy computes 12 and `MaxSize` is 10, the group goes to 10 and CloudWatch shows the policy as capped rather than failing. Likewise, if you lower `MaxSize` below the current desired capacity, desired is pulled down with it and instances are terminated. ## The self-healing loop in practice Suppose desired capacity is 4 and you terminate one instance directly in EC2: 1. The group's next check counts 3 healthy instances against a desired of 4. 2. It launches one replacement from the configured launch template version. 3. The new instance boots, registers with any attached target groups, and passes health checks. Desired capacity was never reduced, because a manual termination is not a signal of intent — from the group's point of view it is indistinguishable from a failure. The same loop is what replaces an instance whose health check fails, and what refills the group after an Availability Zone outage. If you actually want one fewer instance, you have two clean options: ```bash # Option 1: move the setpoint aws autoscaling set-desired-capacity \ --auto-scaling-group-name web-asg --desired-capacity 3 # Option 2: take an instance out of the group and shrink the setpoint with it aws autoscaling detach-instances \ --auto-scaling-group-name web-asg \ --instance-ids i-0123456789abcdef0 \ --should-decrement-desired-capacity ``` The second is how you pull a suspect instance aside for forensics without the group immediately backfilling it and without over-provisioning. ## Availability Zones and distribution Desired capacity is a total across the group, not a per-AZ figure. The group spreads instances across the subnets you gave it and tries to keep the zones balanced, which is why a desired capacity of 3 across three zones is a very different resilience story from 3 in one zone. Sizing a fleet so that it still serves peak load after losing an entire zone is the reason many teams run min capacity higher than steady-state need. ## Common shapes - **`min = max = desired`** — a fixed-size fleet. No elasticity, but you still get automatic replacement of dead instances. This is a perfectly legitimate use of an ASG and often the first one a team adopts. - **`min = 0`** — allows scaling to zero, useful for batch or non-production tiers that should cost nothing when idle. Be aware that the first request after zero has to wait for a full boot. - **min above steady state** — headroom bought deliberately, so a spike or a zone loss does not have to wait for launches. ## What these settings do not do They do not reserve capacity: `MaxSize` 50 is permission to launch 50, not a promise that AWS has 50 of that instance type free in your zones. They are also not a spend cap — they cap instance *count*, and cost depends on what each instance is. And they say nothing about whether an instance is actually serving traffic correctly; that is the health-check configuration's job.

  • You want to pull a misbehaving instance out of the group to investigate it, without the group launching a replacement. How?
    Detach it with `--should-decrement-desired-capacity`, which removes it from the group and lowers desired by one in a single step, so no replacement is launched. If you want the instance to keep serving while you look, you can instead enable scale-in protection on it, but that only stops scale-in — it does not stop health-check replacement.
  • What happens if you set MaxSize lower than the current desired capacity?
    Desired capacity is pulled down to the new maximum and the group terminates the excess instances using its termination policy. It is a real scale-in event, not a validation error, so lowering the ceiling on a live production group takes traffic-serving instances away immediately.
  • Is it sensible to run an Auto Scaling group where min, max and desired are all the same number?
    Yes, and it is common. You give up elasticity but keep automatic replacement of failed instances, even spread across Availability Zones, and a single place that defines how an instance is built. Many teams start there and add scaling policies later once they trust the launch path.

saying these in an interview costs you the question

  • Thinks desired capacity is the maximum the group can reach
  • Believes a manual termination permanently shrinks the group
  • Says a scaling policy can push capacity above MaxSize
  • Assumes MaxSize reserves capacity with AWS in advance
  • Treats desired capacity as a per-Availability-Zone number

context

open as a page

EC2 Auto Scaling supports target tracking, step, simple, scheduled and predictive scaling policies. How would you choose between them for a web tier, and can they be combined?

level: middleimportance: must knowfreq 70%

basics

~20 s

Target tracking is the default choice: name a metric and the value to hold, and Auto Scaling works out the capacity. Step scaling handles graded reactions to severity, scheduled actions cover known clock-driven load, and predictive scaling pre-scales repeating daily patterns. They compose.

open as a page

An EC2 Auto Scaling group references a launch template by version. What is the practical difference between pinning it to $Latest, $Default, or a specific version number?

level: middleimportance: should knowfreq 46%

basics

~20 s

Launch template versions are immutable, so changes create new versions. $Latest means every new instance uses the newest version the moment it exists; $Default uses whichever version you explicitly promoted; a pinned number is fully deterministic. Running instances are never altered.

open as a page

Instances in an EC2 Auto Scaling group pass their EC2 status checks, but the application on them returns errors and the load balancer has marked them unhealthy. Why has Auto Scaling not replaced them, and what would you change?

level: seniorimportance: should knowfreq 56%

basics

~20 s

The group's health check type is EC2, which only reflects hypervisor and instance status checks — a running instance with a dead application passes. Setting the health check type to ELB makes the group consume the load balancer's verdict and replace failing instances.

open as a page

A scale-in event on an EC2 Auto Scaling group terminates workers that are still processing long-running jobs. How do Auto Scaling lifecycle hooks let you drain that work before the instance goes away?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A terminating lifecycle hook pauses the instance in a Terminating:Wait state and emits an event, giving your drain logic a window to finish or hand off work. You end the wait by calling CompleteLifecycleAction, or extend it with a heartbeat.

open as a page

A tier behind an EC2 Auto Scaling group takes about eight minutes from launch to serving traffic, but its load can double in under two. How would you decide between warm pools, predictive scaling, and simply carrying more headroom?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Reactive scaling cannot beat an eight-minute boot, so the choice is about predictability and cost. Predictable load gets predictive scaling or scheduled actions; unpredictable spikes need pre-paid capacity — headroom or a warm pool — and the real fix is shortening boot time.

open as a page