skip to content

As the platform lead for an engineering organisation, how do you decide which workloads may run on EC2 Spot and which must never, and how do you bound the damage when Spot capacity disappears across the fleet?

level: principalimportance: should knowfreq 44%

answer

  1. classify the workload, do not negotiate
  2. what does a two-minute eviction cost
  3. only-copy-of-state is the hard no
  4. floor of On-Demand, elastic tier on Spot
  5. savings must not become an SLA

basics

~20 s

Judge each workload by what an unannounced two-minute eviction costs it: restart cost, statelessness, and customer-visible impact. Spot suits retryable, replaceable work; singletons holding state must not use it. Bound the damage with an On-Demand floor sized to the throughput the business cannot lose.

solid answer

~50 s

I would make it a property of the workload, not a preference of the team. The test is what an eviction with two minutes' notice actually costs: can the work be retried without a human, is any state on the node the only copy, and does a customer see it? Retryable, replaceable things — CI agents, batch and ETL with checkpointing, stateless queue consumers, elastic serving capacity above a baseline — are Spot candidates. Singletons holding the only copy of state, licence servers, self-managed database primaries, brokers with undelivered messages, and long non-resumable jobs must not be. For everything in between, the answer is a mix rather than a yes or no: an On-Demand base capacity sized to the throughput we cannot drop below, with Spot carrying the elastic tier above it. I would also insist that Spot fleets are diversified across many capacity pools and that interruption rates are measured per pool, so that the savings are auditable against the retries and incidents they cost.

code

json · 7 lines
json
{
  "InstancesDistribution": {
    "OnDemandBaseCapacity": 6,
    "OnDemandPercentageAboveBaseCapacity": 25,
    "SpotAllocationStrategy": "price-capacity-optimized"
  }
}

go deeper

for a junior

Know the simple test: if the instance vanishing in two minutes would lose data or break a user-facing request, it is not a Spot workload.

for a middle

Be able to place common workloads into the categories and explain why a mixed fleet with an On-Demand base is usually the answer for serving traffic rather than all-or-nothing.

for a senior

Show that you would size the On-Demand floor to committed throughput, mandate pool diversification and a shared interruption handler, and measure interruption rates rather than trusting last quarter's luck.

for a principal

Own the policy and its economics: publish the never-Spot list, decide the fallback behaviour in advance, watch for organisation-wide pool correlation, and stop Spot savings from becoming an unfunded availability promise.

## Turn the question into a property of the workload The wrong version of this conversation is a negotiation between a finance-minded platform team pushing Spot and product teams resisting it. The right version is a **classification** that every workload gets, applied consistently, with the tradeoff written down. The classifying question is narrow and answerable: *if this instance disappears in two minutes, with no human involved, what is the cost?* Break it into four sub-questions: 1. **Is the work retryable without human intervention?** If a killed unit of work simply reappears and completes elsewhere, the eviction cost is near zero. 2. **Is any state on that node the only copy?** In-memory session state, an unreplicated database volume, an accumulated cache that takes an hour to rebuild — each converts an eviction into a real loss. 3. **Is the effect customer-visible?** Latency, a failed request, a missed deadline. Elastic capacity absorbing a traffic peak fails differently from the last instance serving a checkout page. 4. **What is the restart cost relative to the discount?** A job with a twenty-minute warm-up that is interrupted every hour has negative savings, no matter what the hourly rate says. ## The categories that fall out **Spot by default.** CI and CD build agents, test environments, batch and ETL with checkpointing, media transcoding, model batch inference, load generators, big-data worker nodes, stateless queue consumers. These already treat a node as disposable, and the discount is close to free money. **Spot for the elastic tier only.** Web and API fleets, stream processors, container clusters. Here the honest position is not "Spot or not" but *how much*: a floor of On-Demand capacity sized to the traffic that must always be served, with Spot carrying the peaks. In an Auto Scaling group's mixed-instances policy that is `OnDemandBaseCapacity` for the floor and `OnDemandPercentageAboveBaseCapacity` for the split above it. **Never Spot.** Self-managed database primaries and any singleton holding the only copy of state; message brokers with undelivered messages on local storage; licence and identity servers everything depends on; cluster control-plane or coordinator nodes where quorum loss is catastrophic; long non-resumable jobs; and anything on a contractual deadline whose miss costs more than a year of the savings. Publish this list — it is far more useful as a short, explicit prohibition than as case-by-case judgment. ## Bounding the blast radius Suitability decides *whether*. These decide *how badly it can go*. - **An On-Demand floor per fleet**, sized to committed throughput rather than to a comfortable-sounding percentage. That floor is the guarantee, bought deliberately at full rate. - **Pool diversification as a platform default.** A fleet allowed to use only one instance type in one zone has correlated failure regardless of how suitable the workload is. Many interchangeable types across all usable zones, with a capacity-aware allocation strategy, is a default the platform sets — not a per-team decision. - **Graceful interruption handling in the shared runtime.** A base image or sidecar that watches for the interruption notice and drains, checkpoints and exits cleanly, so every team inherits correct behaviour instead of writing it badly. - **Correlation awareness across teams.** If every fleet in the organisation diversifies into the same three popular instance types, the organisation is still concentrated. Spread the *portfolio* across families and generations, not just each fleet. - **A fallback rule.** Decide in advance what happens when Spot capacity is unavailable for an extended period: automatically top up with On-Demand and eat the cost, or let the queue grow. Both are defensible; discovering which one you chose during an incident is not. ## Measure the savings honestly The discount is gross. The net is the discount minus retried work, minus engineering time spent on interruption handling, minus incidents. Instrument it: record every interruption with instance type and zone, chart interruption rate per pool, and track wasted instance-hours from runs that never completed. That data ends most arguments — it identifies pools to stop using, and it shows when a fleet's Spot ratio has been pushed past the point where it still pays. ## The organisational trap to name The subtle failure is **Spot savings silently becoming an availability promise**. A team runs at 90% Spot, the interruption rate happens to be low for a quarter, the savings get booked into a budget, and the fleet is now implicitly committed to capacity nobody guaranteed. When the pool tightens, the choice is a visible outage or an unbudgeted cost — and that decision arrives during an incident, at the worst possible time. The defence is to keep the guarantee and the discount separate in the way the organisation talks about them: reliability comes from the On-Demand floor and from diversification, and Spot is an *upside* on top. If a service's stated availability depends on Spot capacity being there, the stated availability is wrong, and that is a conversation to have on a calm afternoon rather than at 3 a.m.

  • A team argues their service is fine on 100% Spot because last quarter's interruption rate was under one percent.
    Last quarter's rate is an observation, not a guarantee — Spot capacity is a residual, and the residual disappears exactly when Region-wide demand spikes, which correlates with the events you most need to survive. If the service has a stated availability target, part of its capacity must be guaranteed. Keep Spot as upside above an On-Demand floor.
  • How do you size the On-Demand base capacity rather than picking a round percentage?
    Work from the load the business genuinely cannot fail to serve: the request rate below which errors become customer-visible, or the queue drain rate below which the backlog grows without bound. That number, converted to instances, is the floor. A percentage is arbitrary and drifts as traffic changes; a throughput-derived floor stays meaningful.
  • What organisation-level Spot risk survives even when every individual fleet is well diversified?
    Correlation across fleets. If every team diversifies into the same handful of popular current-generation types, the organisation is concentrated even though each fleet looks spread. Treat instance-family and generation spread as a portfolio property, and steer different fleets toward different families — including Arm-based options where the software supports them.
  • When is a low-priority workload still worth running on On-Demand instead of Spot?
    When its restart cost swamps the discount — a long non-resumable job, or one with an expensive warm-up that is interrupted more often than it completes. Priority is not the criterion; interruption economics are. A low-priority job that must run to completion by a deadline can be a worse Spot candidate than a high-priority stateless worker.

saying these in an interview costs you the question

  • Setting a blanket Spot percentage target across all workloads
  • Running a self-managed database primary on Spot because it has replicas
  • Treating a low observed interruption rate as a guarantee for next quarter
  • Counting the gross discount without the cost of retried and lost work
  • Diversifying each fleet while the whole organisation uses the same few types

context