skip to content

How would you use an ECS capacity provider strategy that mixes FARGATE and FARGATE_SPOT for a production service, and what must the workload be able to tolerate for that to be a responsible choice?

level: principalimportance: nice to knowfreq 32%

answer

  1. a floor and a proportional remainder
  2. only one entry may fix an absolute count
  3. reclaim is not the only failure mode
  4. two minutes is only useful if you use it
  5. split by service, not by exception

basics

~20 s

Set a base of on-demand FARGATE tasks to guarantee a floor, then weight FARGATE_SPOT to carry the elastic remainder. Spot tasks can be reclaimed with a two-minute SIGTERM warning, so the workload must shut down gracefully.

solid answer

~50 s

A capacity provider strategy is a list of `{capacityProvider, weight, base}` entries. `base` is a fixed number of tasks placed on one provider before anything else is considered — only one entry may set it — and `weight` splits every task beyond that base proportionally. So `FARGATE` with `base: 6, weight: 1` plus `FARGATE_SPOT` with `weight: 4` gives you six guaranteed on-demand tasks and then four-fifths of the growth on Spot. The judgment is in choosing the base: it should be the capacity you are unwilling to lose if Spot capacity is reclaimed or simply unavailable, which for a customer-facing service usually means enough to serve your committed traffic floor. The workload must tolerate a task disappearing at any time: Fargate Spot signals a reclaim by sending SIGTERM with roughly two minutes' warning before SIGKILL, so the process must drain in-flight requests and checkpoint or requeue any in-progress work. Stateful, non-idempotent, or long-single-unit jobs are poor fits.

go deeper

for a junior

Know that Fargate Spot is cheaper capacity that AWS can take back, and that a capacity provider strategy decides how tasks are split between it and on-demand Fargate.

for a middle

Explain base and weight precisely — an absolute floor on one provider, ratios for everything above it — and describe the SIGTERM warning a reclaimed task receives.

for a senior

Design against both failure modes: reclaim and unavailability. Show graceful shutdown, correct alarming on running rather than desired capacity, and which workloads you would refuse to put on Spot.

for a principal

Own it as policy: what the guaranteed floor is across services, whether Spot is the cluster default, how the savings are measured against the operational surface, and where the exception boundary lives.

## The mechanism ECS clusters expose capacity providers; on Fargate the two built-in ones are `FARGATE` and `FARGATE_SPOT`. A service (or a `RunTask` call) can specify a `capacityProviderStrategy` instead of a `launchType`: ```json { "capacityProviderStrategy": [ {"capacityProvider": "FARGATE", "base": 6, "weight": 1}, {"capacityProvider": "FARGATE_SPOT", "weight": 4} ] } ``` Two fields do the work: - **`base`** — an absolute count of tasks placed on that provider before any weighting applies. At most one entry in the strategy may set a base. - **`weight`** — the relative share of the *remaining* tasks. Weights are ratios, not percentages: 1 and 4 mean one-fifth and four-fifths of everything above the base. With the example above, a service at `desiredCount: 6` runs entirely on-demand. At 16, it runs 6 on-demand plus 2 more on-demand and 8 Spot. Scaling changes the split continuously, which is exactly the intent: the base is your floor, the weights govern the elastic part. A strategy can also be set as the cluster's default so new services inherit it, which is how a platform team makes Spot the norm rather than an opt-in. ## What Spot actually costs you Fargate Spot is discounted capacity that AWS reclaims when it needs it. Two distinct failure modes follow, and candidates usually name only the first. **Interruption.** A running task is reclaimed. Fargate sends **SIGTERM** to the task's containers and follows with SIGKILL roughly two minutes later, and emits a task state change event you can route through EventBridge. Two minutes is generous compared with most interruption budgets, but it is only useful if the process actually handles the signal — the same graceful-shutdown discipline a rolling deployment demands. **Unavailability.** This is the one people forget. Spot capacity may simply not be available, so a scale-out that expects Spot tasks may place fewer than requested, or none. ECS will not silently promote them to on-demand. If your scaling event coincides with a regional capacity crunch, the Spot portion of your fleet does not arrive. Your base must therefore be sized so that *base alone* keeps the service correct, if degraded — not merely so that base plus Spot is enough. ## Designing the mix Ask three questions in order. 1. **What is the floor you must always serve?** That is the base — typically the capacity for your committed traffic or your error-budget-safe minimum, not your average load. 2. **Is a unit of work cheap to lose and redo?** Stateless HTTP request handling and idempotent queue consumers are ideal: an interrupted task's work is retried by the client or redelivered by the queue. A 40-minute non-restartable job, a task holding a lease or a sticky session, or anything writing non-idempotently is a bad fit. 3. **Does the discount matter at this scale?** Spot complexity buys a meaningful percentage off compute. For a service costing tens of dollars a month, that is not worth the operational surface; for a large asynchronous fleet, it is one of the largest levers available. ## Failure modes to design against - **Correlated interruption.** Reclaims are not spread evenly by design; a chunk of your Spot tasks may go at once. Keep the base able to absorb the gap while replacements start, and make sure your scaling policy reacts on a timescale shorter than your tolerance. - **Interruption during deployment.** A deployment already runs at reduced or doubled capacity; a simultaneous reclaim compounds it. `minimumHealthyPercent` at 100 protects you here. - **Retry storms.** If interrupted work is retried without backoff and jitter, a reclaim event becomes a self-inflicted load spike. - **Metrics that hide it.** Alarm on the *served* capacity, not the desired count. A service whose Spot tasks never placed still reports the desired count it wanted. ## Where it does not belong Singleton controllers, leader processes, schema migrations, anything holding an exclusive lock, and long single-unit batch jobs without checkpointing should stay on on-demand — put them in a separate service with a pure `FARGATE` strategy rather than trying to express the exception inside one mixed strategy. Splitting by service, not by weight, is the cleaner boundary: the strategy is a property of the service, so "which workloads may run on Spot" becomes a legible platform policy rather than a per-team guess. The honest summary for an interview: Spot is a capacity-elasticity decision dressed as a cost decision. You are trading a guaranteed floor for a cheaper ceiling, and the base is where you write down how much guarantee you are buying.

  • What happens if Fargate Spot capacity is unavailable when your service scales out?
    The Spot portion of the requested tasks simply does not launch; ECS does not fall back to on-demand for you. The service runs below desired count until capacity returns. That is why the on-demand base must be sized as a standalone floor, and why alarms should watch running task count and served traffic rather than the desired count you asked for.
  • How does a workload find out a Fargate Spot task is being reclaimed?
    The platform sends SIGTERM to the task's containers roughly two minutes before termination, and emits a task state change event that can be routed through EventBridge. A process that traps SIGTERM can stop accepting work, finish or checkpoint what it holds, and exit cleanly. A process that ignores it loses whatever was in flight when SIGKILL arrives.
  • Would you use a mixed strategy for a nightly batch job that takes 90 minutes and cannot resume?
    No. A single non-resumable unit of work longer than the interruption window means any reclaim throws away the whole run, and reruns cost more than the discount saved. Either make the job checkpoint and resume — at which point Spot becomes attractive — or run it on on-demand capacity as a separate task with a pure FARGATE strategy.
  • How does this differ from an EC2 Auto Scaling group capacity provider on the EC2 launch type?
    There you attach a capacity provider backed by an Auto Scaling group, and ECS managed scaling adjusts the group toward a target capacity utilization while managed termination protection keeps it from terminating instances that still run tasks. You are then managing instance lifecycle and Spot interruption at the host level as well as the task level — more control, more to operate.

saying these in an interview costs you the question

  • Treats Spot as a pure discount with no capacity risk
  • Sets a base equal to zero for a customer-facing service
  • Assumes ECS falls back to on-demand when Spot is unavailable
  • Ignores SIGTERM in the container and loses in-flight work
  • Runs singleton or non-idempotent workloads on Spot

context