skip to content

A batch fleet running entirely on EC2 Spot keeps losing most of its instances within the same minute. What makes a whole Spot fleet vanish at once, and how would you configure it so a single reclamation cannot take everything?

level: seniorimportance: should knowfreq 54%

answer

  1. one pool means correlated failure
  2. type plus size plus zone plus platform
  3. cheapest pool is often the shallowest
  4. many interchangeable types, many zones
  5. capacity-aware allocation, not price-only

basics

~20 s

The fleet is concentrated in one Spot capacity pool — one instance type and size in one Availability Zone — so a single reclamation event hits every instance. The fix is to spread across many pools and let EC2 pick the deepest ones instead of the cheapest.

solid answer

~50 s

Spot capacity is allocated per **pool**: a specific instance type and size, in a specific Availability Zone, for a specific platform. If the whole fleet launched into one pool, AWS reclaiming that pool takes the whole fleet in one go — that is correlated failure, not bad luck. The fix has three parts. First, give the fleet many pools to choose from: a list of interchangeable instance types across several Availability Zones, sized by vCPU and memory rather than pinned to one type. Second, change the allocation strategy away from `lowest-price`, which deliberately concentrates you into the cheapest and usually shallowest pools, to `price-capacity-optimized`, which weighs available capacity as well as price. Third, enable capacity rebalancing so the group replaces instances that receive a rebalance recommendation before they are taken. On top of that, keep an On-Demand floor if the fleet has throughput it must never drop below.

code

json · 20 lines
json
{
  "LaunchTemplate": {
    "LaunchTemplateSpecification": {
      "LaunchTemplateName": "batch-worker",
      "Version": "$Latest"
    },
    "Overrides": [
      { "InstanceType": "m6i.large" },
      { "InstanceType": "m6a.large" },
      { "InstanceType": "m5.large" },
      { "InstanceType": "m5a.large" },
      { "InstanceType": "m5n.large" }
    ]
  },
  "InstancesDistribution": {
    "OnDemandBaseCapacity": 4,
    "OnDemandPercentageAboveBaseCapacity": 0,
    "SpotAllocationStrategy": "price-capacity-optimized"
  }
}

go deeper

for a junior

Know that Spot capacity is grouped into pools by instance type, size and Availability Zone, and that using only one pool means everything can go at once.

for a middle

Explain the allocation strategies and what each optimises for, and describe how widening the instance-type list multiplies the pools a fleet can draw on.

for a senior

Diagnose this as correlated failure, sequence the fixes by leverage, note that pool concentration also blocks scale-up, and say how you would measure interruptions per pool to verify the change.

for a principal

Set the fleet defaults across the organisation — capacity-aware allocation, a minimum pool count, an On-Demand floor tied to committed throughput — so individual teams do not rediscover this during an incident.

## The unit of failure is the pool A **Spot capacity pool** is the set of unused instances of one instance type and size, in one Availability Zone, for one platform. AWS reclaims capacity pool by pool, because that is how capacity is physically organised and how On-Demand demand arrives. That single fact explains the symptom entirely. If a fleet of 200 instances is all `m6i.large` in `us-east-1a`, it is one pool, and its instances are not 200 independent bets — they are one bet, made 200 times. A demand spike in that pool reclaims them together. Interviewers like this scenario because it is a correlated-failure problem wearing a cost-optimisation costume. ## How teams end up concentrated without meaning to Two causes account for most cases. **A single instance type in the launch configuration.** The workload was sized on `c6i.4xlarge`, so the fleet asks for `c6i.4xlarge`, and there is nowhere else for it to go. Any pool the fleet cannot use is a pool that cannot rescue it. **The `lowest-price` allocation strategy.** This one is counter-intuitive and is the real trap. `lowest-price` fills from the cheapest pools first, and cheap pools are cheap *because* demand there is low relative to supply — but they are also often small, and the strategy actively pushes you to concentrate into a few of them. You save a little on the rate and buy a much higher interruption rate. ## The fix, in the order it matters **1. Widen the pool list.** Express the requirement as a set of interchangeable options, not one SKU: several instance types and sizes with comparable vCPU and memory, across every Availability Zone the workload can use. Ten viable types across three zones is thirty pools instead of one. If the workload runs on both x86 and Arm builds, both architectures widen it further. This is the highest-leverage change by a wide margin — everything else is choosing well *among* pools, and it only helps if there are pools to choose among. **2. Choose the allocation strategy deliberately.** In an Auto Scaling group's mixed-instances policy, `SpotAllocationStrategy` accepts `lowest-price`, `capacity-optimized`, `capacity-optimized-prioritized` and `price-capacity-optimized`. `capacity-optimized` picks the pools with the deepest spare capacity and ignores price. `price-capacity-optimized` weighs both, and is the sensible default for most fleets — the marginal price difference between pools is small compared with the cost of being interrupted repeatedly. **3. Turn on capacity rebalancing.** With capacity rebalancing enabled, the Auto Scaling group reacts to a rebalance recommendation by launching a replacement *while the at-risk instance is still running*, rather than noticing a shortfall after the fact. It turns some interruptions into proactive replacements. **4. Leave the price ceiling alone.** Setting `SpotMaxPrice` below the On-Demand price does not reduce what you pay — you always pay the current pool price — but it does add a second reason to be evicted, and it silently removes pools from consideration. Defaulting the ceiling keeps the full pool set available. **5. Put a floor under it if the workload needs one.** A mixed-instances policy also carries `OnDemandBaseCapacity` (a fixed number of instances always bought On-Demand) and `OnDemandPercentageAboveBaseCapacity` (the split for everything above that floor). A fleet with a base of On-Demand instances degrades to reduced throughput when Spot capacity evaporates instead of going to zero. ## The second, quieter failure Diversification protects against more than interruption. When a fleet is pinned to one pool, it also cannot **scale up** if that pool is exhausted — the Auto Scaling group tries to launch, gets an insufficient-capacity error, and sits below desired capacity while the queue backs up. Nobody gets paged for an interruption they never saw; they get paged for a backlog. Widening the pool list fixes both symptoms with one change. ## Verifying rather than assuming Record the interruption events for the fleet — instance type, Availability Zone, timestamp — and chart interruptions per pool over time. That turns the discussion from folklore into evidence: you can see which pools are genuinely deep for your shape of instance, drop the ones that churn, and measure whether the allocation-strategy change actually moved the interruption rate. AWS also publishes interruption-frequency guidance per pool, which is a useful starting point before you have your own data. ## What good sounds like in the interview "The fleet is one pool, so the failures are correlated. I would make the launch template accept a dozen comparable instance types across all usable Availability Zones, switch the allocation strategy from lowest-price to price-capacity-optimized, enable capacity rebalancing, and set an On-Demand base equal to the throughput we cannot drop below. Then I would measure interruptions per pool and prune the bad ones." That answer is mechanism, tradeoff and verification in three sentences.

  • Why can lowest-price allocation end up more expensive than price-capacity-optimized?
    Because the rate is only part of the cost. `lowest-price` concentrates the fleet into the cheapest and typically shallowest pools, so interruptions rise, work is retried, and instance-hours are spent on runs that never finish. `price-capacity-optimized` accepts a marginally higher rate for far deeper pools, and usually delivers more completed work per dollar.
  • The fleet is spread across three Availability Zones but still lost everything at once. What is likely missing?
    Instance-type diversity. Three zones of a single instance type is three pools, and correlated demand for a popular type — a new generation everyone is adopting, for instance — can squeeze all three together. Adding a set of comparable types across those zones is what actually multiplies the pool count.
  • How would you size the On-Demand base capacity for a Spot-heavy fleet?
    Set it to the throughput the business genuinely cannot drop below — the level at which the queue stops growing, or the requests you must serve. That floor is bought at On-Demand rates deliberately as the cost of a guarantee; everything above it rides Spot and is allowed to disappear.

saying these in an interview costs you the question

  • Blaming random bad luck rather than a single shared capacity pool
  • Assuming lowest-price allocation is always the cheapest overall
  • Diversifying across Availability Zones but keeping one instance type
  • Setting a low max price believing it reduces the hourly cost
  • Ignoring that a concentrated fleet also cannot scale up when the pool is dry

context