skip to content

A provider can reclaim a batch job's machine with minutes of warning - what must fit inside that notice, and what does it not promise?

level: middleimportance: must knowfreq 62%

answer

  1. a warning, not an extension
  2. drain, do not finish
  3. state must outlive the machine
  4. best-effort delivery
  5. unit smaller than the window

basics

~20 s

The reclaim notice buys just enough time to stop taking new work, flush a checkpoint to durable shared storage and hand the unfinished unit back. It promises no time to finish, no extension, and not even that it will arrive.

solid answer

~50 s

Reclaimable capacity is rented on the condition that the provider may take the machine back when it wants that capacity, and the reclaim notice is a short warning delivered to the machine before it goes - minutes rather than hours, and providers publish different lengths. What has to fit inside it is a drain: stop claiming new work units, write a checkpoint of the unit in progress to storage that outlives the machine, release or return that unit so another worker can pick it up, and deregister from whatever was sending work here. What it does not promise is time to finish the current unit, an extension if you ask, a veto, or delivery at all - capacity is also lost abruptly with no warning. The notice is an optimisation on top of a design that already tolerates sudden loss, not a substitute for one.

code

pseudocode · 9 lines
pseudocode
on reclaimNoticeReceived:
    acceptNewWork = false                  // 1. stop claiming
    writeCheckpoint(currentUnit, to = durableSharedStore)   // 2. make it resumable
    workSource.release(currentUnit)        // 3. hand it back now, do not wait for lease expiry
    workRouter.deregister(thisMachine)     // 4. stop anything new arriving
    exit()                                 // 5. before the deadline

// note: the unit is NOT finished here - whatever is left is redone
// by whichever worker claims it next, starting from the checkpoint

go deeper

for a junior

Recall the shape: capacity that is cheaper because the provider can take it back, and a short warning sent to the machine first. Know that work must be saved somewhere other than that machine.

for a middle

Explain what the window is for - stop claiming, checkpoint, hand the unit back, deregister - and state plainly that it promises neither completion nor delivery. Naming the ordering earns as much as naming the steps.

for a senior

Show the design that survives no notice at all: expiring leases, idempotent restart, durable progress. Then position the notice as the thing that makes recovery cheaper, and size work units against the published window.

for a principal

Frame it as a risk posture: an option the provider holds, paid for in engineering. Decide which fleets may hold that option against them and what the estate's standard drain contract is.

## What reclaimable capacity actually is A provider sells the same machines under several postures. **Metered on-demand capacity** is yours until you release it. A **term commitment** is capacity you promised to keep paying for, so the risk sits with you. **Reclaimable capacity** is the third posture: the machine is yours until the provider wants it back, and it is priced lower because the *provider* kept an option, not because you signed anything. That difference in who holds the risk is the whole design consequence - a commitment hurts when your demand shrinks, reclaimable capacity hurts when someone else's demand grows. Because the provider holds that right, it needs a way to exercise it, and the reclaim notice is that mechanism. ## The notice, precisely - It is delivered **to the machine**, as a signal software running there can observe - not to a human, and not to your work queue or your scheduler. - It is **short and fixed**: minutes rather than hours. Providers publish a length; the lengths differ, and some capacity classes carry a longer one than others. - It starts a countdown you **cannot extend, defer or decline**. Nothing you send back changes the outcome. - It is **best-effort**. A machine can also disappear with no notice at all, because ordinary hardware and network failures do not announce themselves. ## What has to fit inside the window In order, because the order is what makes the work recoverable: 1. **Stop claiming new work.** Anything claimed after the notice is work you are certain not to finish. 2. **Checkpoint the unit in progress** to durable shared storage that outlives this machine. The machine's own local disk goes away with it, so a checkpoint written there is gone exactly when it is needed. 3. **Hand the unit back** - release the lease or return it to the queue so another worker can claim it immediately instead of waiting for a lease to expire. 4. **Deregister** from whatever was routing work or requests here, so nothing new arrives during the last seconds. 5. **Exit.** Whatever is unfinished at the deadline is cut off wherever it happens to be. | The notice gives you | The notice does not give you | |---|---| | A bounded window to make work resumable | Time to finish the unit in progress | | A defined point to stop accepting work | An extension, a deferral or a veto | | A chance to release leases and held state | Any guarantee that it will arrive | ## Size the work unit to the window The practical move is to make the **unit of work smaller than what you can wrap up inside the window**, and to checkpoint often enough that the redo left over is bounded and small. If one indivisible step runs longer than the notice, the notice buys that step nothing: the step is cut off mid-way whatever your handler does, and the only recovery is the last checkpoint. Teams discover this when a job that checkpoints every hour meets a warning measured in minutes and loses the hour every time. Two related constraints follow. First, the drain itself must be fast: if writing state takes longer than the window, shrink what a checkpoint contains or write it incrementally. Second, the handler must be **safe to run twice and safe to interrupt**, because a notice can arrive while a scheduled checkpoint is already being written. ## Why the notice is an optimisation, not the design If the notice were a guarantee, you could build around it. It is not, so the design has to survive its absence: - Work is handed out under a **lease that expires**, so a silent disappearance eventually returns the unit. - A restart is **idempotent**, so a unit that runs twice does no harm. - Progress lives in **durable shared storage**, never only in memory or on the machine's disk. Given those three, the notice makes recovery *faster and cheaper*: less redone work, and no waiting for a lease to expire before another worker picks the unit up. Without them, the notice only shrinks the loss on the machines that happened to get one. ## What interviewers listen for A strong answer separates three things that candidates routinely merge: what the notice *is* (a short, one-way, best-effort warning), what it is *for* (making the work resumable elsewhere), and what the design must do *anyway* (tolerate loss with no warning). A weak answer treats reclamation as a rare failure to be handled by a shutdown hook, which is exactly backwards: on reclaimable capacity, interruption is normal operation.

  • The handler releases the unit back to the queue. Why not just let the lease expire instead?
    Both end with the unit re-offered, but an explicit release happens in seconds while an expiry waits out the whole lease - dead time during which nobody is working the unit. Releasing is the optimisation the notice makes possible; the expiry is the safety net for the machines that got no notice. Keep both.
  • A single step of the job takes longer than the notice window. What can the handler usefully do?
    Only make the loss recoverable: release the unit and exit, so another worker restarts that step from the last checkpoint. It cannot finish the step and cannot buy more time. The real fix is upstream - split the step, or checkpoint inside it, so the redo is bounded rather than a full step.
  • Should the reclaim handler write a checkpoint even if one was written a minute earlier?
    Usually yes, if the write fits comfortably in the window: it is the cheapest minute of redone work you will ever buy back. The exception is state so large that the write would not complete, in which case an interrupted write must not overwrite the good checkpoint - publish atomically under a new name and keep the previous one.

saying these in an interview costs you the question

  • Assumes the notice always arrives, so silent loss is impossible
  • Believes the warning is long enough to finish the current unit
  • Thinks you can ask the provider to extend or defer the reclamation
  • Writes the reclaim-time checkpoint to the machine's own local disk
  • Treats reclamation as a rare failure rather than normal operation