skip to content

An EC2 Spot Instance is about to be reclaimed by AWS. What signals do you get beforehand, where does a process running on the instance read them, and what should that process do in response?

level: middleimportance: must knowfreq 68%

answer

  1. two signals, not one
  2. advisory first, then the hard notice
  3. poll metadata or subscribe to events
  4. nothing sends your process a signal
  5. stop intake, drain, checkpoint, exit

basics

~20 s

AWS publishes two signals: an optional earlier rebalance recommendation, and a two-minute Spot interruption notice. Both appear in instance metadata and as EventBridge events. The worker should stop taking new work, checkpoint or requeue in-flight work, and drain from its load balancer.

solid answer

~50 s

There are two distinct signals. The **rebalance recommendation** arrives when EC2 judges the instance at elevated risk — it can come well before an interruption and carries no guarantee that one follows. The **interruption notice** is the hard one: roughly two minutes before EC2 takes the instance. A process reads either from the instance metadata service — `/latest/meta-data/events/recommendations/rebalance` and `/latest/meta-data/spot/instance-action`, which returns 404 until a notice exists and then returns JSON with the action and time — or subscribes to the matching EventBridge events at the account level. AWS does not signal your process; you must poll or subscribe. On the notice, a well-behaved worker stops pulling new work, deregisters from its target group so the load balancer drains it, finishes or checkpoints what it holds, and returns any in-flight queue message so another consumer picks it up promptly rather than waiting out a timeout.

code

json · 8 lines
json
{
  "source": "aws.ec2",
  "detail-type": "EC2 Spot Instance Interruption Warning",
  "detail": {
    "instance-id": "i-0abcdef1234567890",
    "instance-action": "terminate"
  }
}

go deeper

for a junior

Know that a Spot instance gets roughly two minutes of warning and that the warning shows up in instance metadata — nothing is pushed to your program, so something has to look for it.

for a middle

Explain both signals and both delivery paths, name the metadata path for the interruption notice, and describe the shutdown order: stop intake, drain, checkpoint, exit.

for a senior

Show that you treat the notice as best-effort — the work must already be idempotent and restart-safe, because instances also disappear with no notice, and telemetry must be flushed or interrupted nodes go unmeasured.

for a principal

Own the standard: a fleet-wide interruption handler and capacity rebalancing as defaults, plus measurement of interruption rates per pool so the organisation knows what its Spot discount is actually costing in retries.

## Two signals, not one People conflate these constantly, and interviewers probe the difference because it changes what you can build. **Rebalance recommendation.** EC2 emits this when it judges an instance to be at *elevated risk* of interruption. It can arrive noticeably earlier than the two-minute notice — sometimes long before — but it is advisory. An instance that got a rebalance recommendation may run for hours afterwards and may never be interrupted at all. Its value is that it gives you time to replace capacity *proactively*, while you still have the old instance. **Interruption notice.** This is the commitment: EC2 has decided to reclaim the instance, and it will do so about two minutes later. There is no negotiating and no extension. Design consequence: use the rebalance recommendation to start bringing up a replacement, and the interruption notice to shut down cleanly. If you only handle the notice you have two minutes to do everything, including waiting for a new instance to boot and warm up — which is usually not enough. ## Where you read them There are two delivery paths, and mature designs use both. **From inside the instance, via the instance metadata service.** The Spot notice lives at `/latest/meta-data/spot/instance-action`. It returns HTTP 404 until a notice exists, and then a small JSON document naming the action (`terminate`, `stop` or `hibernate`) and the time it takes effect. The rebalance signal lives at `/latest/meta-data/events/recommendations/rebalance` and carries a `noticeTime`. A daemon polls these every few seconds — polling metadata is free and local. ```bash TOKEN=$(curl -sX PUT "http://169.254.169.254/latest/api/token" \ -H "X-aws-ec2-metadata-token-ttl-seconds: 60") curl -s -o /dev/null -w '%{http_code}' \ -H "X-aws-ec2-metadata-token: $TOKEN" \ http://169.254.169.254/latest/meta-data/spot/instance-action # 404 = no notice yet; 200 = you have ~2 minutes ``` **From outside, via EventBridge.** EC2 also publishes account-level events — an EC2 Spot Instance Interruption Warning and an EC2 Instance Rebalance Recommendation — which you can route to a Lambda function, a Step Functions workflow, or an SNS topic. This path is right for fleet-level reactions: draining a node, updating an inventory, recording interruption rates per pool. **Critically, AWS does not send your process a signal.** On a bare EC2 instance nothing arrives as SIGTERM from the platform; if your application does nothing, it simply vanishes mid-request. (Container orchestrators layer their own behaviour on top — that is the orchestrator reacting to the same notice, not EC2 signalling you.) ## What a good worker actually does in those two minutes Order matters, because the point is to make the interruption invisible rather than merely survivable. 1. **Stop accepting new work first.** Close the intake — stop long-polling the queue, stop claiming new batch shards. Everything else is wasted if new work keeps arriving. 2. **Leave the traffic path.** For a serving instance, deregister from its target group so the load balancer stops sending requests and drains open connections. Doing this early buys most of the two minutes for in-flight requests to finish. 3. **Checkpoint or hand back in-flight work.** For a batch job, flush a checkpoint to durable storage — S3 or EBS, never the instance store, which disappears with the instance — so a replacement resumes rather than restarts. For a queue consumer, explicitly return the message you are holding so another consumer receives it immediately instead of waiting for a timeout to expire. 4. **Flush what is only in memory.** Buffered logs, metrics and traces are lost otherwise, and losing exactly the telemetry from interrupted instances is how teams end up blind to how often Spot bites them. 5. **Exit.** Do not try to finish work that plainly cannot finish in the window; hand it back so it is retried, not half-applied. ## Idempotency is the real requirement The honest framing is that the notice is a *courtesy*, not a contract you can lean on. Two minutes may not be enough, the process might be mid-write, and an instance can also be lost with no notice at all for ordinary hardware reasons. So the design must be idempotent and restart-safe on its own merits: work is retried at least once, checkpoints are written atomically, and partial output is either invisible or overwritten. The interruption handler then turns a correct-but-ugly outcome into a clean one — it should never be what makes the system correct. ## In an Auto Scaling group If the Spot instances sit in an Auto Scaling group, turn on capacity rebalancing so the group reacts to the rebalance recommendation by launching a replacement *before* the old instance is taken, rather than discovering the shortfall afterwards. Pair that with a graceful in-instance handler, and the two mechanisms cover the two halves of the problem: the group restores capacity, and the process protects the work it was holding.

  • Why is it not enough to handle the interruption notice alone, if the code is already graceful?
    Because two minutes is only enough to shut down, not to replace capacity. The rebalance recommendation arrives earlier and lets an Auto Scaling group launch and warm a replacement while the doomed instance is still serving, so throughput never dips. Handling only the notice means every interruption is a capacity gap.
  • Your consumer stops after each interruption but downstream sees duplicate side effects. What went wrong?
    The handler returned in-flight work for retry without the processing being idempotent. Anything a worker can be killed halfway through must be safe to repeat — write with a deterministic key, make the final commit atomic, or record a processed-message marker. The interruption exposed a correctness gap it did not cause.
  • How would you find out how often your fleet is actually being interrupted?
    Route the EC2 Spot Instance Interruption Warning events from EventBridge to a target that records instance type, Availability Zone and time, and chart interruptions per pool. That per-pool interruption rate is what tells you whether a pool is worth using, and it is also the input for choosing diversification.

saying these in an interview costs you the question

  • Believing AWS sends the application SIGTERM when Spot is reclaimed
  • Treating the rebalance recommendation as a guaranteed interruption
  • Checkpointing to instance store, which dies with the instance
  • Assuming two minutes is always enough, so retries are unnecessary
  • Holding an in-flight queue message rather than returning it before exit

context