skip to content

EC2 Auto Scaling supports target tracking, step, simple, scheduled and predictive scaling policies. How would you choose between them for a web tier, and can they be combined?

level: middleimportance: must knowfreq 70%

answer

  1. reactive versus clock versus forecast
  2. declare the value you want held
  3. severity-graded response needs steps
  4. cooldown blocks, warmup only excludes metrics
  5. forecast raises the floor, never scales in

basics

~20 s

Target tracking is the default choice: name a metric and the value to hold, and Auto Scaling works out the capacity. Step scaling handles graded reactions to severity, scheduled actions cover known clock-driven load, and predictive scaling pre-scales repeating daily patterns. They compose.

solid answer

~60 s

For a normal web tier I start with a **target-tracking** policy — you declare a metric and a target value, such as `ALBRequestCountPerTarget` at some number per instance or `ASGAverageCPUUtilization` at 50 percent, and Auto Scaling manages the alarms and the arithmetic. It is self-correcting and hard to misconfigure. I reach for **step scaling** only when I want a graded response — add one instance on a small breach, five on a large one — or when the metric is something the group cannot use as a per-instance average. **Simple scaling** is the legacy form and I would not choose it today; its cooldown blocks all further scaling while it waits. **Scheduled actions** handle load that is a function of the clock rather than of a metric: a batch window, a market open, a marketing send. **Predictive scaling** learns a repeating daily or weekly shape and raises capacity ahead of it, which matters when instances take minutes to become useful. They combine: predictive or scheduled sets the baseline, and target tracking stays on as the reactive safety net.

code

bash · 11 lines
bash
aws autoscaling put-scaling-policy \
  --auto-scaling-group-name web-asg \
  --policy-name requests-per-target \
  --policy-type TargetTrackingScaling \
  --target-tracking-configuration '{
    "PredefinedMetricSpecification": {
      "PredefinedMetricType": "ALBRequestCountPerTarget",
      "ResourceLabel": "app/web-alb/50dc6c495c0c9188/targetgroup/web-tg/73e2d6bc24d8a067"
    },
    "TargetValue": 800.0
  }'

go deeper

for a junior

Know that target tracking is the sane default and that you configure it by naming a metric and the value you want held, rather than by writing alarms and step tables yourself.

for a middle

Explain when a graded step policy beats target tracking, why the tracked metric must fall as instances are added, and the difference between a cooldown and an instance warmup.

for a senior

Show production judgment about metric choice and layering: a forecast or schedule for the known shape, target tracking underneath as the net, and a minimum capacity chosen for zone loss rather than average day.

for a principal

Own the argument about where elasticity belongs at all. Argue when a queue or a serverless tier absorbs bursts better than any policy can, and how you decide the cost of carrying headroom against the cost of arriving late.

## The distinction that organises all of them Every Auto Scaling policy answers one question — *what should desired capacity be?* — but they differ in what they use as input. **Dynamic** policies (target tracking, step, simple) react to a CloudWatch metric that has already moved. **Scheduled** actions react to the clock. **Predictive** scaling reacts to a forecast built from history. Reactive policies are always late by the length of a boot; the other two are attempts to stop being late. ## Target tracking — the default You give it a metric and the value you want held, and Auto Scaling creates and manages the underlying CloudWatch alarms itself. The predefined metric types include `ASGAverageCPUUtilization`, `ALBRequestCountPerTarget`, `ASGAverageNetworkIn` and `ASGAverageNetworkOut`, and you can supply a customised metric specification for anything else. ```bash aws autoscaling put-scaling-policy \ --auto-scaling-group-name web-asg --policy-name cpu-target \ --policy-type TargetTrackingScaling \ --target-tracking-configuration '{"PredefinedMetricSpecification":{"PredefinedMetricType":"ASGAverageCPUUtilization"},"TargetValue":50.0}' ``` The key property is that the metric must be one that **falls as you add instances**. Average CPU and requests-per-target both do. A queue depth does not — adding consumers drains it faster, but the raw depth is not proportional to fleet size, which is why a backlog-per-instance metric is the usual fix there. Target tracking scales out aggressively and scales in conservatively by design, because being briefly over-provisioned is much cheaper than being under. You can also attach `DisableScaleIn` and let something else own shrinking. If you attach several target-tracking policies, the group scales out to satisfy the most demanding one and only scales in when all of them agree. ## Step scaling — graded reactions Step scaling binds a CloudWatch alarm to a set of step adjustments keyed by how far the metric has breached, using `MetricIntervalLowerBound` and `MetricIntervalUpperBound`. Each step chooses an adjustment type: `ChangeInCapacity`, `PercentChangeInCapacity` or `ExactCapacity`. So CPU 10 points over target might add one instance while 40 points over adds six. Use it when the severity of the breach should change the response, or when you already own a well-tuned alarm. The cost is that you are now maintaining the thresholds and the arithmetic yourself, which is exactly what target tracking exists to avoid. ## Simple scaling — legacy Simple scaling applies one adjustment per alarm and then blocks *all* scaling activity for the duration of the cooldown before it will re-evaluate. During a genuine ramp that is disastrous: you add two instances, then sit still for the cooldown while load keeps climbing. Step and target tracking supersede it, and they do not use the group's default cooldown at all — they use an **instance warmup** instead, configured group-wide as `DefaultInstanceWarmup`. Warmup means "do not count this new instance's metrics until it has settled", which prevents the classic over-scaling loop where booting instances report low CPU and drag the average down. Being able to say *cooldown blocks scaling, warmup only excludes an instance's metrics* is a strong signal in an interview. ## Scheduled actions — load that is a function of the clock A scheduled action sets min, max and/or desired at a time, optionally on a cron recurrence with a time zone (so daylight-saving shifts do not silently move your business hours). Use it when you *know* the shape: office-hours tiers, a nightly batch window, a televised event. It is not a forecast — it is an instruction, and it will happily provision for a peak that does not arrive. ## Predictive scaling — forecast, then pre-scale Predictive scaling analyses the group's history and produces a forward forecast, then raises capacity ahead of the predicted rise. It needs a repeating pattern and at least a day of history to say anything useful, and it runs in one of two modes: `ForecastOnly`, where you compare the forecast against reality before trusting it, and `ForecastAndScale`, where it acts. Critically, **it only scales out** — it raises the effective minimum capacity ahead of predicted load and leaves scale-in to your dynamic policy. That design is what makes it safe to run alongside target tracking. You can also configure a `MaxCapacityBuffer` so the forecast is allowed to exceed the current maximum by a percentage. Predictive scaling pays for itself when boot time is long relative to the ramp, and does nothing for load that is genuinely unpredictable. ## How to combine them The production-grade arrangement for a web tier is layered: 1. **Predictive or scheduled** sets the baseline ahead of known or forecast demand. 2. **Target tracking** stays on permanently as the reactive net for whatever the forecast missed. 3. **Minimum capacity** is set for survivability — enough to serve after losing an Availability Zone — not for the average day. When policies disagree, the group takes the largest scale-out request, so layering is additive on the way up and cautious on the way down.

  • Why is average CPU a poor scaling metric for a tier whose work is dominated by waiting on a downstream call?
    Because CPU barely moves while threads block on I/O, so the metric never breaches even as latency climbs and queues build. Scale on something that reflects the actual constraint — requests per target, in-flight requests, or a custom metric such as backlog per instance — otherwise the fleet stays flat through the incident.
  • What is the difference between a cooldown and an instance warmup in EC2 Auto Scaling?
    A cooldown blocks the group from taking any further scaling action for a period, and belongs to simple scaling. A warmup instead says a newly launched instance's metrics should not be counted until it has settled, so target tracking and step scaling can keep reacting while a boot is in progress without being dragged around by the booting instance's own numbers.
  • Can you run predictive scaling and target tracking on the same group at once?
    Yes, and it is the recommended pairing. Predictive scaling only raises capacity ahead of a forecast, so it cannot fight a reactive policy for control of scale-in. Target tracking then handles everything the forecast did not anticipate. Start predictive in ForecastOnly mode and compare the forecast against actual load before letting it act.

saying these in an interview costs you the question

  • Picks CPU as the scaling metric for every workload
  • Thinks target tracking requires you to write CloudWatch alarms
  • Believes predictive scaling terminates instances when demand falls
  • Confuses cooldown with instance warmup
  • Says only one scaling policy may be attached to a group

context