skip to content

Purchase Options and Spot Economics

The same instance can cost wildly different amounts depending on how I buy it, so cost questions are really commitment questions. I learn On-Demand versus Reserved versus Savings Plans versus Spot, and how to build a workload that survives a two-minute Spot interruption notice.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

6

In EC2, what do you get and what do you give up by running an instance as Spot instead of On-Demand, and how is the Spot price actually determined?

level: juniorimportance: must knowfreq 78%

answer

  1. cheap because interruptible
  2. spare capacity, reclaimable
  3. two minutes of warning
  4. no auction since 2017
  5. MaxPrice is a ceiling, not a bid

basics

~20 s

Spot Instances run on EC2 capacity AWS currently has idle, at a steep discount off On-Demand, but AWS can reclaim them at any moment with a two-minute warning. Spot prices move gradually with supply and demand; there is no bidding auction.

solid answer

~50 s

Spot lets you run EC2 instances on capacity AWS has spare right now, typically at a large discount off the On-Demand rate — often quoted as up to 90%, though the real discount varies a lot by instance type and Availability Zone. What you give up is tenure. AWS can take the instance back whenever it wants the capacity, giving you a two-minute interruption notice first, and it can simply refuse to launch new Spot instances when a pool is tight. Since the 2017 pricing change there is no bidding war: each pool's price moves gradually with supply and demand, and the optional `MaxPrice` you can set is only a ceiling above which you stop being launched or get interrupted — not a bid you win. So Spot fits work that can be killed and retried, and is wrong for anything whose interruption costs more than the discount saves.

code

bash · 5 lines
bash
aws ec2 run-instances \
  --image-id ami-0abcdef1234567890 \
  --instance-type m6i.large \
  --instance-market-options 'MarketType=spot,SpotOptions={SpotInstanceType=one-time}' \
  --count 1

go deeper

for a junior

Know that Spot runs on AWS's spare capacity at a big discount and can be reclaimed with a two-minute notice, and be able to name two workloads that tolerate that (CI builds, batch jobs).

for a middle

Explain that a capacity pool is instance type plus size plus Availability Zone plus platform, that pricing moves gradually rather than by auction, and what MaxPrice actually controls today.

for a senior

Show you weigh the restart cost of a workload against the discount, and mention that failing to launch Spot capacity is as real a risk as being interrupted.

for a principal

Frame the purchase option as a reliability contract the organisation is choosing, and be ready to say where Spot savings must not be allowed to become an implicit availability promise.

## The purchase option is a promise about tenure, not just a price The *same* EC2 instance type, in the same Availability Zone, running the same AMI, can cost wildly different amounts depending on how you buy it. The reason is not that AWS gives volume discounts for fun — it is that each purchase option is a different bargain about **who bears the risk of capacity**. - **On-Demand** — you pay the published per-second rate and the instance is yours until *you* stop or terminate it. AWS never takes it away. You make no commitment and get no discount. - **Spot** — you run on capacity that is currently unsold. It is deeply discounted precisely because AWS reserves the right to take it back the moment an On-Demand or reserved customer wants it. That is the whole trade in one line: **Spot is cheaper because it is interruptible.** ## What "interruptible" concretely means When AWS decides to reclaim your Spot Instance, three things happen that you must design for: 1. **You get a two-minute interruption notice.** It appears in the instance metadata service and as an EventBridge event. Two minutes is short but not nothing — it is enough to stop accepting new work, flush a checkpoint, and drain. 2. **The instance then goes away** with the *interruption behavior* configured on the request. The default is terminate; stop and hibernate are available for persistent requests. Anything on an instance store volume is gone. 3. **You may not be able to get another one.** This is the failure people forget. Interruption is not the only risk — a tight pool also means a *launch* request for Spot capacity simply fails. Spot gives you no capacity guarantee going forward either. ## How the price is set — and the myth of the auction Spot began life as a genuine auction: you submitted a bid, and if your bid exceeded the current market price you ran, and your instance died the moment the price crossed your bid. That model was retired in late 2017 and **candidates still describe the old one in interviews**, which is a reliable tell. Today, AWS sets the price for each **capacity pool** — a given instance type and size, in a given Availability Zone, for a given platform — and moves it *gradually* based on long-run supply and demand for that pool. It no longer spikes to the On-Demand ceiling on a moment's notice, which is what made the old model so frightening to build on. You pay the price in effect for each second you run. You may still set a `MaxPrice` on the request. Its meaning is now narrow: if the pool's price rises above your ceiling you will not be launched, and a running instance can be interrupted for that reason. The default `MaxPrice` is the On-Demand price for that type, and leaving it at the default is usually right — setting it lower does not save you money (you pay the market price either way), it only adds a second reason to be interrupted. ```bash # The current price of one pool, most recent entries first aws ec2 describe-spot-price-history \ --instance-types m6i.large \ --product-descriptions "Linux/UNIX" \ --max-results 5 ``` ## Where Spot belongs The usable rule is: **can this workload be killed mid-flight and restarted somewhere else without a human noticing?** Good fits — CI/CD build agents, batch and ETL jobs with checkpointing, media transcoding, stateless queue consumers, big-data worker nodes, load generators, ephemeral test environments. All of them already treat a node as disposable. Bad fits — the single primary of a self-managed database, a message broker holding undelivered messages, a licence server, a long non-resumable job, anything whose restart cost or customer-visible latency exceeds what you saved. Between those poles sits the common production answer: run the fleet as a *mix*, keeping a floor of On-Demand instances for the traffic you must always serve, and letting Spot carry the elastic part above it. ## The mental model to carry into the interview Don't describe Spot as "cheap EC2". Describe it as **a different reliability contract**. You are not buying a discount; you are selling AWS an option to evict you, and the discount is the premium you collect for it. Every design decision that follows — checkpointing, diversification, an On-Demand base, graceful shutdown on the notice — is just you making sure that option is cheap for you to honour.

  • If you leave MaxPrice unset, what does AWS use, and why is that usually the right choice?
    The default ceiling is the On-Demand price for that instance type. Leaving it there is usually right because you always pay the current Spot price, not your ceiling — a lower ceiling saves nothing and only adds a second cause of interruption when the pool price drifts up.
  • Besides being interrupted, what other Spot failure mode should a design account for?
    Not getting capacity in the first place. When a pool is tight, a Spot launch request simply fails, so an Auto Scaling group can sit below its desired capacity indefinitely. That is why production designs diversify across pools and keep an On-Demand floor rather than relying on Spot to scale up on demand.
  • Does the discount differ between instance types, and how would you check before committing to a design?
    Yes — the discount is per capacity pool and varies substantially by type, size, Region and Availability Zone. Check `describe-spot-price-history` for the pools you intend to use and the published Spot placement/interruption-frequency guidance, and design against the pools that are both cheap and deep, not just cheap.

Spot is standby airline seating: far cheaper because the airline can bump you the moment a full-fare passenger shows up, and the deal only works if being bumped is genuinely survivable for you.

saying these in an interview costs you the question

  • Saying you bid against other customers and the highest bid wins
  • Claiming Spot prices spike unpredictably to the On-Demand rate
  • Believing a low MaxPrice makes the instance cheaper to run
  • Assuming AWS terminates Spot instances with no warning at all
  • Treating Spot as a way to guarantee cheap capacity is always available

context

open as a page

An EC2 Spot Instance is about to be reclaimed by AWS. What signals do you get beforehand, where does a process running on the instance read them, and what should that process do in response?

level: middleimportance: must knowfreq 68%

basics

~20 s

AWS publishes two signals: an optional earlier rebalance recommendation, and a two-minute Spot interruption notice. Both appear in instance metadata and as EventBridge events. The worker should stop taking new work, checkpoint or requeue in-flight work, and drain from its load balancer.

open as a page

A team needs certainty that twenty EC2 instances of a specific type will be available in one Availability Zone for a launch next month. Which EC2 purchase mechanism actually holds that capacity, and which ones only change the bill?

level: middleimportance: should knowfreq 42%

basics

~20 s

Only an On-Demand Capacity Reservation actually holds EC2 capacity in a specific Availability Zone. Savings Plans and regional Reserved Instances are billing constructs with no capacity guarantee; a zonal Reserved Instance is the one commitment that also reserves capacity.

open as a page

A batch fleet running entirely on EC2 Spot keeps losing most of its instances within the same minute. What makes a whole Spot fleet vanish at once, and how would you configure it so a single reclamation cannot take everything?

level: seniorimportance: should knowfreq 54%

basics

~20 s

The fleet is concentrated in one Spot capacity pool — one instance type and size in one Availability Zone — so a single reclamation event hits every instance. The fix is to spread across many pools and let EC2 pick the deepest ones instead of the cheapest.

open as a page

As the platform lead for an engineering organisation, how do you decide which workloads may run on EC2 Spot and which must never, and how do you bound the damage when Spot capacity disappears across the fleet?

level: principalimportance: should knowfreq 44%

basics

~20 s

Judge each workload by what an unannounced two-minute eviction costs it: restart cost, statelessness, and customer-visible impact. Spot suits retryable, replaceable work; singletons holding state must not use it. Bound the damage with an On-Demand floor sized to the throughput the business cannot lose.

open as a page

In EC2, what is the difference between Dedicated Instances and a Dedicated Host, and in what situation does that difference actually matter?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

Both isolate you from other AWS customers' instances. Dedicated Instances give isolation only — AWS still chooses the hardware. A Dedicated Host allocates you a whole physical server, exposing its socket and core counts and letting you pin instances to it, which is what per-socket software licensing needs.

open as a page