You are budgeting the outage window for a planned stop and resume of a continuous job: what goes into it, and what must it fit inside?
answer
- the window is longer than the deploy
- five spans, not one
- the restore read is the forgotten one
- catch-up divided by spare capacity
- must fit inside the source's retention
basics
~20 sFive spans: draining, writing the restart point, the deploy and getting machines, reading the saved copy back, and catching up the input that piled up. It must fit inside the source's retention and whatever freshness the consumers were promised.
solid answer
~40 sThe outage window — the wall-clock time during which the job produces no output — is longer than the deploy. It is the drain, plus writing out what the job was holding, plus the build and submission and whatever wait there is for machines, plus reading the saved copy back onto the new workers, plus the tail nobody counts: the input that arrived while all that happened still has to be processed before output is current again. Two hard limits bound it. The source keeps its records only so long, and an outage longer than that retention means records are gone rather than delayed. And the catch-up tail is set by spare capacity: a job running at 1.05 times its arrival rate clears a ten-minute backlog in about two hundred minutes, not ten.
go deeper
Understand that a job taken down on purpose produces nothing while it is away, and that the records still arriving during that time do not disappear — they wait, and they have to be processed later.
Be able to list the spans that make up the window rather than just the deploy, and explain why the input that piled up means output is not current the moment the job restarts.
Do the catch-up arithmetic out loud: backlog is arrival rate times outage, draining at capacity minus arrival rate. Say which span you measured last time and what you would check against the source's retention.
Own the number as a commitment. Decide what freshness the organisation is actually paying for, how much headroom that implies as standing cost, and who is told before a stop that exceeds it.
## Why there is a window at all A stateless web service is replaced instance by instance and never stops answering. A job that has accumulated results cannot generally be replaced that way: the accumulated contents have exactly one owner, and handing them to a new set of processes means the old set lets go first. So there is a span with no output, and the engineering question is not how to make it zero but how long it honestly is and what it must fit inside. ## The five spans 1. **Drain.** Bounded by the slowest work already inside the job and by the destination's willingness to accept it. Usually seconds; occasionally the surprise, when one enormous group or a struggling destination holds everything up. 2. **Writing the restart point.** A whole copy of what the job holds, pushed to storage outside any one machine. It scales with the size of that accumulated set, and unlike the saves taken while running, none of it overlaps with useful work. 3. **The deploy itself.** Build, artifact distribution, submission, and the wait for machines. That wait is not always yours to control: where machines come from a shared pool, the cluster resource manager may queue you behind someone else, and where a cluster is raised per job, creation takes what it takes. 4. **Reading the saved copy back.** Every new worker pulls its share over the network before it can process a record. For a large accumulated set this frequently dominates everything above it, and it is the span most often left out of the estimate entirely. 5. **Catch-up.** Output resumes here, but it is not *current* here. Everything that arrived during spans one to four is still ahead of the job. Only when that backlog is cleared is the pipeline honestly back, and how a job behaves while running behind is its own subject; what matters for the budget is simply that this tail exists and is usually the longest of the five. ## The arithmetic that gets skipped Catch-up is not proportional to the outage; it is proportional to the outage divided by your spare capacity. If records arrive at rate **R** and the job can sustain **C**, then an outage of length **T** leaves a backlog of **R × T**, draining at **C − R**: | Sustained capacity | Spare | Backlog after a 10-minute stop | Time to clear | |---|---|---|---| | 2.0 × R | 1.0 × R | 10 minutes of input | ~10 minutes | | 1.5 × R | 0.5 × R | 10 minutes of input | ~20 minutes | | 1.05 × R | 0.05 × R | 10 minutes of input | ~200 minutes | | 1.0 × R | none | 10 minutes of input | never | A job sized to keep up exactly can never recover from any outage, and most teams discover their true headroom for the first time during a deploy. The corollary is that headroom is not waste — it is the budget for every planned stop you will ever make. ## What the window must fit inside - **The source's retention horizon.** **A retained, re-readable input** keeps its records for a while and lets a reader resume from a named position — *for a while*. Exceed it and the restart cannot resume; it can only skip forward to what still exists, which is data loss dressed as a successful deploy. This is the one limit where overrunning converts delay into permanent loss. - **Whatever freshness consumers were promised.** If a dashboard, an alert or a downstream job assumes numbers are never more than n minutes old, the outage window is a number the organisation already committed to without writing it down. - **Accumulation limits at the source itself.** Unread records occupy space somewhere. A long outage on a high-rate input can push the source into its own trouble before it pushes you into yours. - **The clock consumers actually read.** An outage that straddles an hourly or daily boundary produces a number someone reads as final while it is still filling. ## How a lead shrinks it - Have the machines before you stop, not after. The wait for capacity is often the largest controllable span. - Keep real headroom above the arrival rate, and know what it is, because it sets the catch-up tail. - Separate the parts of a change that require the job down from the parts that do not, and stop only for the former. - Publish the number. A deploy procedure whose cost is unstated gets run at nine in the morning by someone who assumed it was instant. ## Where engines differ A runtime serving an endless input as **repeated small finite jobs** already stops between pieces, so its drain is close to free — but its restore is not, and neither is its catch-up. A **record-at-a-time** runtime holding a large accumulated set has the reverse profile: a short drain to manufacture and a restore that can dominate. And a job carrying nothing between records skips spans one, two and four entirely: its window is the deploy plus catch-up, which is why the same procedure can cost thirty seconds for one job and half an afternoon for its neighbour.
- Which span surprises teams most often?Two, in practice. Reading the saved copy back, because it scales with the accumulated set and nothing overlaps it — the job is up and idle until it finishes. And the catch-up tail, because people declare the deploy done when output resumes rather than when output is current, which on a job with thin headroom can be hours apart.
- How would you shrink the window without touching the job's logic?Acquire the machines before stopping, so the deploy does not wait on capacity; keep genuine headroom above the arrival rate, which is what sets the catch-up tail; and split the change so that only the parts genuinely needing the job down cause a stop. Each is an operational choice, and together they usually beat any tuning of the save itself.
- What do you check afterwards to confirm the window was what you budgeted?The gap in the output series against the wall clock, the moment lag returned to its normal band rather than the moment the job started, and the oldest unread record at the source during the stop against the retention horizon. If the third came close, the budget was wrong regardless of how well it went this time.
saying these in an interview costs you the question
- Counts only the deploy and calls that the outage window.
- Declares the pipeline recovered when output resumes rather than when it is current.
- Ignores the time spent reading the saved copy back onto new workers.
- Assumes catch-up takes about as long as the outage.
- Never checks the outage against the source's retention horizon.
- Treats headroom above the arrival rate as waste.