skip to content

For a 60 GB keyspace under heavy sustained writes, how do you choose a copy interval, and when does a whole-keyspace copy stop being worth taking?

level: principalimportance: should knowfreq 38%

answer

  1. two curves, opposite directions
  2. price one copy, then the schedule
  3. duration approaching the interval
  4. a shorter interval never reaches zero

basics

~20 s

A shorter copy interval shrinks the writes sitting outside the newest copy, while every copy costs memory, tail latency and device bandwidth. Choose the longest interval whose exposure the business will sign, and stop when copies cannot finish cheaply.

solid answer

~50 s

Start from what the exposure means: at interval N, a crash can lose up to N minutes of acknowledged writes, and that sentence is what a business signs rather than a setting. Then price a single copy on this tier — measured duration, peak resident memory during the write, and the tail rise callers see — because the total cost is that figure multiplied by copies per hour. Pick the longest interval whose exposure is acceptable, not the shortest the machine survives. The signal that you have gone too far is measured duration approaching the interval: cuts then arrive while the previous write is still running and the tier is effectively copying continuously, paying full cost permanently for a schedule it is not meeting. And when no interval is acceptable — the state cannot tolerate any window at all — the answer is not a shorter interval but a different home for that state, because a copy always leaves a gap.

go deeper

for a junior

Recall that copies are taken on a schedule and that the schedule is a trade: more often means less lost, but taking a copy is not free.

for a middle

Explain both currencies concretely — writes outside the newest copy on one side, divergence memory and tail latency per copy on the other — and that the cost scales with copies per hour.

for a senior

Show that you price a single copy from measurements on this tier and watch for duration approaching the interval, which means the schedule is no longer being met.

for a principal

Own the sentence: state the maximum acknowledged-write loss in plain words, decide when copies should stop being taken at all, and rule the tier out for state that tolerates no window.

## Two curves, moving in opposite directions Choosing a copy interval is not tuning; it is pricing a trade that has a curve on each side. - **Shorten the interval** and the writes sitting outside the newest copy shrink roughly linearly. At a 15-minute interval you may lose up to 15 minutes of acknowledged writes; at five, up to five. - **Shorten the interval** and the cost of taking copies rises just as directly, because every effect of taking one is *per copy*: the divergence memory held while it is written, the tail latency callers feel, the device bandwidth consumed, and the risk that the spike lands during peak traffic. The decision is where on that curve the workload belongs, and the honest form of the answer names both currencies at once. ## Price one copy before you price the schedule On a 60 GB keyspace under heavy sustained writes, the inputs are measurable rather than theoretical: 1. **Measured copy duration.** How long does writing 60 GB actually take on this device, with this machine's other work running? This is the term everything else scales with. 2. **Peak resident memory during the write.** Divergence is roughly the distinct data changed during the copy — write rate times that duration — so a long write under heavy writes is the expensive corner of the space. 3. **The tail rise callers see** for the duration, and whether the machine has the headroom to keep that a tail rather than a collapse. 4. **The success rate.** A copy that fails partway leaves the previous one in place, so the exposure you can actually quote runs between *successful* copies. Multiply the per-copy cost by copies per hour, and the schedule prices itself. ## The signal that the interval is fiction The clearest operational tell is **measured duration approaching the configured interval**. When a 60 GB keyspace takes twelve minutes to write and the interval is fifteen, the tier is close to a state where the next cut arrives before the previous write has finished. Practically, that means: - The tier is paying the copy's full cost — memory, latency, bandwidth — essentially all of the time rather than in a bounded window. - The schedule on paper no longer describes what happens; cuts are skipped or queued, so the real gap is longer than the number you quoted. - Any growth in the keyspace or the write rate makes it worse without any change on your side. At that point shortening the interval further buys nothing at all — the tier cannot go faster than continuous. ## When a whole copy stops paying A periodic whole copy earns its cost in a narrow band: the keyspace is valuable enough that coming back empty hurts, and quiet or small enough that a copy completes without damage. It stops paying when: - **The write rate makes divergence unaffordable**, so each copy risks the machine it is protecting. - **The keyspace is so large that duration swallows the interval**, as above. - **What the tier holds is reconstructible cheaply** and the system of record behind it can absorb a repopulating tier — then the honest decision is to stop taking copies and invest in making an empty restart survivable instead. - **The state cannot tolerate any window.** This is the one case where tuning is the wrong instrument entirely: no interval reaches zero, so state that must not lose an acknowledged write needs a home built for that, and the tier holds a copy of it rather than the only instance. ## Write the window down as a sentence The deliverable of this decision is not a number in a configuration; it is a claim somebody owns: *"we may lose up to 15 minutes of acknowledged writes on this tier, and here is what that means for each thing we keep on it."* Written that way it can be argued with by people who do not operate the tier, which is the point. It also forces the inventory question — what is actually on here, and which of those things would be materially harmed by that sentence being true. ## Across a fleet At fleet scale two more considerations appear. Copies that all cut at the same clock minute concentrate their cost, so the same total work is felt as a synchronised spike rather than a smooth load; spacing the schedules across instances flattens it. And every instance's exposure is the *worst* of its recent gaps, not the configured one, so the number reported upward should come from measured successful copies rather than from the schedule anybody believes is in force.

  • A team responds to a data-loss incident by cutting the copy interval from fifteen minutes to one. What would you check before agreeing?
    Measured copy duration first: if a copy takes several minutes, a one-minute interval is unachievable and the tier will copy continuously, paying divergence memory and tail latency permanently. Then peak resident memory during a copy, since fifteen times more copies means fifteen times more chances for the spike to land at peak traffic.
  • What number would you report as this tier's exposure to someone outside the team?
    The largest recent gap between successful copies, expressed as a sentence: we may lose up to N minutes of acknowledged writes. Reporting the configured interval overstates the guarantee, because a failed copy silently widens the gap and a copy whose duration approaches the interval means cuts are not landing on schedule.

saying these in an interview costs you the question

  • Answers "take copies more often" to any tolerance requirement
  • Assumes the cost of copying is independent of how often you do it
  • Quotes the configured interval as the exposure actually achieved
  • Believes a short enough interval makes the tier a safe only home
  • Ignores that copy duration grows with the keyspace and swallows the interval