skip to content

Designing for Reclaimable Capacity

Running work on capacity the provider can take back at short notice: checkpointing, draining on the reclaim signal, mixing it with guaranteed machines. Asked because the warning is short.

on this pageshow

questions

5

How do you choose how often an hours-long training job checkpoints when its machine can be reclaimed at any minute?

level: middleimportance: must knowfreq 50%

answer

  1. two costs, opposite directions
  2. half the interval is redone
  3. plus the restore, not just the redo
  4. balance write cost against interruption rate
  5. atomic publish, keep the previous

basics

~20 s

Balance two costs that pull in opposite directions: redone work, which is about half the interval every time capacity is reclaimed, against the runtime each checkpoint write itself consumes. Measure both, then pick the interval where they are roughly equal.

solid answer

~50 s

A reclamation costs you everything done since the last checkpoint, and reclamation lands anywhere inside the interval, so on average you redo **half the interval** each time - plus the time to restore and re-claim the unit. Checkpointing more often shrinks that, but every checkpoint stalls the job for as long as the write takes. So the interval is set by two measured numbers: the cost of writing one checkpoint, and the mean time between reclamations for this fleet. Redone work per hour grows with the interval and overhead per hour shrinks with it, and the sensible interval is near where the two are equal - roughly the square root of twice the checkpoint cost times the mean time between reclamations. Then make the checkpoint restorable: write it to durable shared storage, publish it atomically so an interrupted write never becomes the one you restore, and keep the previous one.

code

pseudocode · 9 lines
pseudocode
// two measured numbers, both in minutes
checkpointCost = 0.5              // runtime one checkpoint write consumes
meanBetweenReclamations = 240     // observed for this fleet, this shape

interval = squareRoot(2 * checkpointCost * meanBetweenReclamations)   // ~15.5

redoneFraction   = interval / (2 * meanBetweenReclamations)   // ~3.1% of runtime
overheadFraction = checkpointCost / interval                  // ~3.2% of runtime
// the interval is well chosen when these two land close together

go deeper

for a junior

Recall that work is saved periodically so an interruption resumes instead of restarting, and that the saved copy has to live somewhere other than the machine doing the work.

for a middle

Explain the trade-off with the mechanics: about half the interval is redone per interruption, each write costs runtime, and the two measured numbers that locate the balance are write cost and interruption rate.

for a senior

Demonstrate that you have restored from one: atomic publish, a retained previous checkpoint, resumable boundaries, and a write time re-measured as state grows over the life of a long job.

for a principal

Treat it as a fleet-wide default: what interval and what checkpoint contract teams get for free, and how you know the interruption rate the arithmetic depends on is measured rather than assumed.

## What one reclamation actually costs When the provider takes the machine back, the job loses everything computed since its last checkpoint. Reclamation is not synchronised with your checkpoint schedule, so over many interruptions it lands roughly uniformly inside the interval: **the expected redo is about half the interval**. On top of that sits a fixed recovery cost that candidates usually forget - the time for another worker to claim the unit, read the checkpoint back, and warm up before it is producing again. So the honest cost of an interruption is `interval / 2 + restore`, not `interval / 2` alone, and not the whole run. ## The two costs pull in opposite directions 1. **Redone work** rises as the interval grows. As a fraction of runtime it is `interval / (2 x meanTimeBetweenReclamations)`. 2. **Checkpoint overhead** rises as the interval shrinks. As a fraction of runtime it is `checkpointCost / interval`, where the cost is the runtime the write consumes. Add them and the total has a minimum. Differentiating gives a tidy rule of thumb: the interval is near `squareRoot(2 x checkpointCost x meanTimeBetweenReclamations)`, and at that point the two fractions are roughly equal - which is the version worth remembering, because it tells you what to measure rather than what to recite. ## A worked example Assume, purely as illustration, that writing one checkpoint costs about **half a minute** of runtime and that this fleet loses a machine roughly every **four hours** (240 minutes). Both numbers are invented; the point is the shape. | Checkpoint interval | Redone work | Write overhead | Total lost runtime | |---|---|---|---| | 60 minutes | ~12.5% | ~0.8% | ~13.3% | | 15 minutes | ~3.1% | ~3.3% | ~6.5% | | 5 minutes | ~1.0% | ~10.0% | ~11.0% | The square-root rule puts the balance near 15 minutes, which is where the table bottoms out. Notice the shape of the curve: it is **flat near the minimum and steep at both ends**. Being within a factor of two of the right interval costs almost nothing, while checkpointing once at the end, or every thirty seconds, costs a great deal. That is why this is a measurement question and not a tuning obsession. The same arithmetic explains why a fleet on capacity that is reclaimed often needs a *shorter* interval than one that is rarely interrupted, even with identical state: halving the mean time between reclamations pulls the balance down by about a third. ## Granularity: what a checkpoint has to contain An interval is meaningless if the checkpoint cannot be resumed from. A usable checkpoint carries **everything the next worker needs to continue and nothing it can recompute**: - the position in the work - which step, which pass, which input offset; - the mutable state the computation carries forward; - the identity of the work unit, so the restart is recognisably the same unit. And it must be taken at a **boundary the code can resume from**. A checkpoint captured halfway through an update that the restart cannot reproduce is not a shorter interval, it is a corrupt one. If natural boundaries are further apart than the interval you want, the boundaries are the thing to change. ## Writing one that survives the interruption - Write to **durable shared storage**. The machine's local disk is a cache that dies with the machine; it may hold a copy for speed, but it can never be the only copy. - **Publish atomically**: write under a temporary name, then move or register it under the canonical one. Otherwise a reclamation during the write leaves a truncated checkpoint in the place the restart looks. - **Keep the previous one or two.** A checkpoint that faithfully captures a state the job cannot continue from is worth exactly as much as no checkpoint, and the earlier one is the way out. - **Measure the write, do not estimate it.** Checkpoint cost is one of the two inputs to the interval, and it grows with state size over the life of a job. ## What interviewers listen for The answer they want is a trade-off stated with numbers the candidate could actually obtain - how long a write takes, how often this fleet is interrupted - rather than a habit ("every ten minutes"). The strongest answers add the second half: an interval is only as good as the restore, so they describe the atomic publish and the retained previous checkpoint without being asked.

  • The job's state grows as it runs, so each checkpoint takes longer to write. What follows for the interval?
    The balance moves: as the write cost rises, the interval that minimises lost runtime lengthens, roughly with its square root. Recompute it from the measured write time rather than fixing the interval at the start, and if the write starts approaching the reclaim window, shrink what a checkpoint contains - incremental or partial state - instead of simply checkpointing less.
  • Would you ever checkpoint on a work boundary instead of on a timer?
    Often, and it is usually better. A boundary checkpoint is guaranteed resumable, while a timer can fire mid-update and force the code to carry extra machinery to make that point restorable. Use the timer to decide how often boundaries should occur - if they are further apart than the interval the arithmetic wants, create more of them.
  • A reclamation interrupts the checkpoint write itself. What does the next worker read?
    The previous checkpoint, provided the new one is published atomically - written under a temporary name and only moved into place once complete. Without that, the restart reads a truncated file in the canonical location and either fails or, worse, resumes from nonsense. Keeping one earlier checkpoint is the cheap insurance.

It is the same judgment as saving a long document on a machine whose power is unreliable: save once an hour and an outage costs you an hour of retyping, save after every keystroke and you spend the day watching the save spinner instead of writing.

saying these in an interview costs you the question

  • Checkpoints only at the end, so an interruption costs the whole run
  • Assumes more frequent checkpoints are always better regardless of cost
  • Ignores the runtime each checkpoint write itself consumes
  • Counts redone work but never the restore and re-claim time
  • Restores from a checkpoint that was interrupted halfway through its write
open as a page

A provider can reclaim a batch job's machine with minutes of warning - what must fit inside that notice, and what does it not promise?

level: middleimportance: must knowfreq 62%

basics

~20 s

The reclaim notice buys just enough time to stop taking new work, flush a checkpoint to durable shared storage and hand the unfinished unit back. It promises no time to finish, no extension, and not even that it will arrive.

open as a page

Every machine in a reclaimable training fleet disappeared inside one minute - why does that happen, and what spreads the risk?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Reclamation is not an independent per-machine accident: it follows the provider's demand for one machine shape in one place, so identical machines from one capacity pool share a single fate. Spreading means several shapes and several zones, and a workload able to run on all of them.

open as a page

How would you set the split between guaranteed and reclaimable machines for a fleet, and which workloads would you bar from the reclaimable half?

level: principalimportance: should knowfreq 34%

basics

~20 s

Size the guaranteed baseline at the demand that must still be served when every reclaimable machine is taken back at once, and put resumable, deadline-tolerant work on the surge. Keep coordinators, sole copies of state and the recovery path itself off the reclaimable half.

open as a page

Some reclaimed machines vanish without ever delivering the warning - what must the work distribution do so a unit is not lost?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Hand work out under a lease that expires unless the worker keeps renewing it, so a silent disappearance returns the unit automatically. That makes delivery at-least-once, which in turn requires restarts and outputs to be idempotent.

open as a page