skip to content

What do you weigh when choosing how often a long-running job writes a recovery point, and what makes each one expensive?

level: seniorimportance: should knowfreq 50%

answer

  1. two costs pulling opposite ways
  2. half an interval redone on average
  3. recovery time is more than replay
  4. capture duration approaching the interval
  5. start from the promise, not a number

basics

~20 s

Weigh what one capture costs — the slowdown while workers save and the bytes pushed to shared storage — against the input a crash makes you reprocess, which is about half an interval on average.

solid answer

~40 s

Two costs pull against each other. Each capture takes time and bandwidth: workers pause or slow while their part is written, every part travels to storage outside any one machine, and they all arrive at once, so that storage sees a burst rather than a trickle. Against that, a crash costs the input since the last completed capture — about half an interval on average, a whole one at worst — plus noticing the failure, obtaining machines, loading the parts and then catching up at better than the input rate. Start from the recovery time the service actually promises, work backwards to an interval, then check the per-capture cost fits comfortably inside it. The alarm is not a number but a ratio: once a capture's duration approaches the interval, the interval is already fiction.

go deeper

for a junior

Recall that captures are not free and that a crash costs the work done since the last one, so the interval is a trade rather than a setting to push as low as it goes.

for a middle

Explain both sides as mechanism: what a worker pays while its part is written, what a restart pays in re-reading, and why the average loss is about half an interval.

for a senior

Derive the interval from the recovery time the service promises, and treat a capture whose duration approaches the interval as evidence that the promise is already broken.

for a principal

Own the promise itself: what a tighter recovery time is worth to the business, and whether the money buys a shorter interval, faster storage for the parts, or a smaller retained set.

## The two costs that pull against each other On one side, each capture of a **recovery point** — a durable copy of everything the job would otherwise lose, written while the job keeps running — costs something while it happens. On the other, the gap between captures is input that a crash makes you process again. Choosing an interval is the arithmetic between those two, and neither is a number you can look up. ## What one capture costs - **A slowdown, sometimes a pause.** How much depends on the mechanism: a worker reconciling two inputs may hold one of them back, which is a real pause on that path, while a worker that saves at once instead writes more bytes. Either way the job's throughput dips for the duration. - **Bytes to storage outside any one machine.** Size is driven by how much the job has accumulated and by how many workers hold it. The picture cannot be written to the local disk of a machine that may be the one that dies. - **A burst, not a trickle.** All the parts are written at roughly the same time, so the storage behind them sees a spike in throughput and in request count, often while the job is reading and writing its own data through the same path. - **A coordination round.** The capture has to be started, tracked and recorded as complete. That is small, but it is per capture, so it grows as the interval shrinks. ## What a restart actually costs The replayed input is only part of it. In order: 1. The failure is noticed, which is not instantaneous and depends on how failure detection is set up. 2. Machines are obtained again, which is slow if capacity has to be acquired rather than reused. 3. Each worker loads its part of the last **completed** capture; one that was in progress is discarded, so the effective gap can be a whole interval plus part of another. 4. The input is re-read from the recorded read position in that capture and reprocessed. 5. Fresh input kept arriving throughout all of the above, so the job is now behind and has to run faster than its input rate to close the gap. The last step is the one people forget. If the job is five minutes behind by the time it is processing again and it can sustain one and a half times its input rate while catching up, the backlog drains in five divided by nought point five — ten minutes, not five. In general, a backlog of B minutes drained at r times the input rate takes B divided by (r minus one). On average a crash costs about half an interval of reprocessing, because it is roughly equally likely anywhere within the interval. Budget with the average; promise with the worst case. ## Full against incremental capture Some runtimes write the whole accumulated set every time, some write only what has changed since the last capture, some offer both, and some give no choice. | | Full capture | Incremental capture | |---|---|---| | Bytes per capture | The whole retained set | Only what changed | | Interval it makes affordable | Longer | Shorter | | Resuming | Read one picture | Read a chain, or a compacted base plus deltas | | The cost that hides | Bandwidth and storage every time | Background compaction, and a long chain slowing resume | Two traps follow from that table. The first is assuming incremental capture makes recovery cheaper too: it usually makes it dearer, because the read side has to assemble what the write side avoided writing. The second is assuming that doubling the interval halves the bytes: with incremental capture each delta simply covers twice as much change, and only repeated updates to the same key coalesce away. What a longer interval reliably saves is the fixed per-capture overhead and the number of bursts, not the volume of change. ## How to choose 1. Start from the recovery time the service actually promises — failure to caught-up, which is the only number anyone outside the team cares about. 2. Subtract the fixed parts: detection, obtaining machines, loading the parts. 3. What is left bounds how much input you can afford to reprocess and drain, and that gives an interval. 4. Measure a capture's duration and size at that interval. If duration is a large fraction of the interval, the interval is fiction and something has to change. ## Symptoms at each extreme - **Too frequent:** periodic dents in throughput, storage and request-rate cost out of proportion to the data, captures starting before the previous one finished, and a job that never settles into steady throughput. - **Too rare:** a long reprocessing window after every restart, and a backlog that takes longer to drain than the reprocessing itself took. When the cost of a capture is the problem, the interval is often the least effective dial. Reducing what is being saved, or putting the parts on faster and less contended storage, usually moves more.

  • Why is the work redone about half an interval on average rather than a whole one?
    Because a crash is roughly equally likely at any point within the interval, so the expected distance back to the last completed capture is about half of it. The worst case is a whole interval plus the duration of a capture that had begun and did not finish, since an incomplete capture is discarded. Budget against the average and promise against the worst case.
  • A capture now takes almost as long as the interval between captures. What does that mean?
    That the stated interval is no longer real: the next capture starts as the previous one finishes, or is skipped, so the actual distance a restart rewinds is longer than the setting suggests. The usual causes are a retained set that has grown, storage that is slower or busier than it was, or one worker taking far longer than the rest. Lengthening the interval is honest; leaving it as it is, is not.
  • What can you change other than the interval when captures cost too much?
    Reduce what is being saved, move to writing only what changed if the runtime offers it, give the parts faster or less contended storage, or spread the burst by letting workers write at slightly different times where that is supported. The interval is the dial everybody reaches for first and often the least effective, because the cost per capture is dominated by bytes and by the storage behind them.

Saving a long document while you are still writing it. Save on every keystroke and you spend the morning waiting for the disk; save once a day and a crash costs you the day. What you pick depends on how much retyping you can stand, not on how large the document is.

saying these in an interview costs you the question

  • Names a single interval that is right for every job
  • Counts only the replayed input as the recovery time
  • Assumes a longer interval cuts bytes written proportionally
  • Thinks writing only what changed makes restoring cheaper too
  • Ignores that every worker's part hits shared storage at once