skip to content

A machine the provider may take back costs far less per hour. For which shapes of distributed job does that saving fail to arrive?

level: principalimportance: should knowfreq 35%

answer

  1. the rate is not the cost
  2. count the machine-hours redone
  3. exposure grows with the run
  4. a latency budget has no slack
  5. judge the tail, not the mean

basics

~20 s

For jobs that redo more than they save: long runs whose later rounds depend on output held only on those machines, and latency-bound continuous jobs, where each withdrawal buys a pause, a rewind and a backlog that a job with no spare throughput cannot clear.

solid answer

~50 s

Compare the hourly rate with the machine-hours actually billed. A job absorbs withdrawals when its pieces are independent, its rounds are short and its results land durably as it goes: the exposure is the work since the last durable write, so a withdrawal costs a slice. It does not absorb them when later rounds consume output that lives only on those machines, because the redone volume grows with the run, and a long run should expect more than one withdrawal, each dearer than the last. Latency-bound continuous jobs are the other exclusion: a withdrawal there means a pause, a rewind to a saved recovery point and a backlog, and a job sized for its input rate has no headroom to catch up. Deadline work is judged on the tail of its completion time, not the mean, and the tail is exactly what this capacity worsens.

go deeper

for a junior

Recall that cheap machines can be withdrawn and that the cheapness is not free: the job may pay it back in work it has to do a second time.

for a middle

Explain the two job shapes on either side of the line - independent pieces writing durable output, against later rounds consuming output that lives only on the machines.

for a senior

Quantify it: exposure per machine, expected withdrawals per run, redone hours added to billed hours, and the effect on elapsed time for whoever is waiting.

for a principal

Own the policy. Say which classes of work may sit on this capacity, make the criterion a property of the job rather than a preference, and decide on the tail of completion time rather than its mean.

## The rate is not the cost The advertised saving is per machine-hour. What an organisation pays is machine-hours billed, and a job on withdrawable capacity is billed for hours spent producing something a second time as readily as for hours of progress. The decision therefore is not "is this capacity cheaper" - it is - but "does this job's shape let the discount survive contact with a withdrawal". The property that decides it is the one a general treatment of reclaimable capacity cannot state: **in a job whose workers feed each other, losing a machine can destroy work that was already finished**, because intermediate output produced there was still being fetched. A job where that is untrue behaves like any other interruptible workload. A job where it is true has a loss that grows with how far it has got. ## The arithmetic to actually do 1. Estimate the **exposure**: at a typical moment, how much produced-but-unconsumed output sits on one machine, and how much finished work is discarded if it goes - one piece, or the wave that contained it. 2. Estimate how often a withdrawal lands during a run of this length. The frequency is a property of the capacity market, not of the job, and it varies by machine type and region; take it from your own history rather than from a published figure. 3. Multiply, and add the redone hours to the billed hours. Compare that total against the guaranteed-capacity total at full rate. 4. Then price the **elapsed time** separately. Redone work sits on the critical path, so it delays the answer as well as costing machine-hours, and whoever waits for the answer pays that. ## Shapes that absorb a withdrawal - Pieces that are independent: each reads its slice, computes, writes its own result out durably. Nothing another piece needs lives on the machine. - Short rounds, so little produced output is unconsumed at any instant. - Incremental durable output, so finished work stops being at risk as soon as it is written. - Slack in the deadline, so a longer tail in completion time costs nothing but patience. - Runs short enough that expecting zero or one withdrawal is reasonable. ## Shapes that should not sit here | shape | what the withdrawal actually costs | why the discount does not survive | |---|---|---| | a long finite run whose later rounds fetch earlier rounds' output | the product of several completed rounds, plus the wave held up waiting for it | exposure grows with elapsed time, so the expected waste grows faster than the run does | | a latency-bound continuous job | a pause, a rewind to a saved recovery point, and a backlog to work off | a job sized for its input rate has no surplus throughput, so the backlog persists and the latency commitment is already broken | | work with a hard delivery deadline | variance in completion time rather than mean cost | the commitment is met or missed by the tail, and this capacity widens the tail | | a job holding a large per-key retained set | the retained set of the lost workers, which must exist again before those keys are processed | the pause scales with how much is remembered, not with the piece in flight | ## What varies between engines, and why it changes the answer A claim that is true of one runtime here is routinely false of a rival, and the judgment turns on which one you are on. - Where each round's output is materialised for the next round to fetch, a withdrawal destroys finished product - the expensive case above. - Where records are pushed onward as they are produced, there are no files to lose, but the affected part of the job rewinds to a saved recovery point, which is the latency case above. - Where continuous work is run as a rapid succession of small finite jobs, exposure resets every time one of those small jobs completes, which makes the same workload markedly more tolerant. So the same business workload can be a good or a bad fit depending on the engine it runs on, and the honest answer says which property it depends on rather than naming a verdict. ## Reshape the job before rejecting the capacity Several of the disqualifying properties are choices, not facts. Cutting one long run into several shorter ones that each land durable output, shortening rounds, and reducing how much of one machine's product the next round depends on all shrink the exposure. That is a change to the job, and it is usually cheaper than paying full rate for capacity the job did not need to be so fragile about. ## What this question does not cover How the warning signal is handled, how work in progress is drained, how risk is spread across capacity pools and what proportion of a fleet should be guaranteed against discounted are all questions about buying reclaimable capacity in general, and they have the same answers whatever runs on it. The part that is specific here is the one above: what a withdrawal costs a job whose workers feed each other.

  • Which measurement tells you whether the discount is actually arriving?
    The share of billed machine-hours that went into producing something a second time, measured across many runs, together with the spread of wall-clock completion time. If redone hours stay small and the spread is tolerable, the saving is real; if either climbs as runs get longer, the capacity is being paid for twice.
  • Can a long dependent run be reshaped to sit here rather than moved off?
    Often. Cutting it into shorter runs that each land durable output, shortening rounds so less product is unconsumed, and reducing how much of one machine's output the next round depends on all shrink the exposure. That is a change to the job rather than to the purchase, and it usually costs less.
  • Why is a job holding a large per-key retained set a poor fit?
    Because what has to exist again after a withdrawal is not only the piece in flight but everything those workers were remembering, which must be recovered and redistributed before their keys can be processed again. The pause scales with how much is retained, so the loss is unrelated to how far along the piece was.

saying these in an interview costs you the question

  • Compares the two hourly rates and stops there.
  • Assumes redone work is a rounding error at any job length.
  • Puts a latency-bound continuous job on this capacity because the discount is large.
  • Judges a deadline-bound workload by its average completion time.
  • Thinks any workload becomes suitable once results are saved more often.
  • Assumes every engine loses the same thing when a machine is taken back.