skip to content

Point-in-Time Copies

A whole copy of the keyspace written periodically while writes continue: what that divergence costs in memory and latency, and how much you lose between copies.

on this pageshow

questions

4

A store writes a whole copy of its keyspace to disk every 15 minutes while still accepting writes; which writes does that copy hold?

level: juniorimportance: must knowfreq 68%

answer

  1. one instant, not a range
  2. the cut, not the file close
  3. writes after the cut are absent
  4. gap runs between successful copies

basics

~20 s

A point-in-time copy holds the keyspace as of the instant it was cut, not when its file finished writing. Writes accepted after the cut are absent, so up to one whole copy interval of acknowledged writes can be missing.

solid answer

~50 s

The copy is anchored to one instant — the moment it was cut — and the store keeps serving throughout the minutes it takes to write the bytes out, so a write acknowledged while the file was still open is not in it. If the copy is cut at 12:00 and the file closes at 12:04, what you can restore is the 12:00 keyspace; the 12:02 write lives only in memory until the next cut at 12:15 catches it. That makes the writes at risk bounded by the copy interval: a process that dies just before the next cut loses almost a full interval of acknowledged writes. And the gap that matters runs between *successful* copies, not between scheduled ones — a copy that failed partway leaves the previous one standing and doubles the real gap.

go deeper

for a junior

Recall that a copy is a photograph of one instant, not a running record: the writes that arrive after the cut are simply not in it, and a crash loses them.

for a middle

Explain the two clocks — when the copy was cut and when its file closed — and show that the writes at risk are bounded by the copy interval rather than by how long the write took.

for a senior

Demonstrate that you monitor measured copy duration and copy success, and that you quote the exposure from the gap between successful copies rather than from the configured schedule.

for a principal

Frame it as a number the business signs: state the maximum acknowledged-write loss in plain words, and be explicit that no copy interval makes this tier a safe only home for state.

## The copy and the copy interval A **point-in-time copy** is a whole copy of an in-memory store's **keyspace** — every key and its value — written out to durable storage while the store carries on answering callers. The schedule on which it is taken is **the copy interval**. Its purpose is narrow: so that a process which dies and comes back does not come back holding nothing. The interview question underneath it is always the same one: *what is inside that copy, and what is not?* ## The cut, not the completion A copy represents the keyspace as it stood at a single instant — the moment the copy was **cut** — and not the moment its file finished being written. Those are two different clocks, and on a large keyspace they can be minutes apart, because pushing tens of gigabytes to a device takes time and the store does not stop serving while it happens. - Cut at 12:00, file closed at 12:04: the artefact holds the **12:00** state. - A write acknowledged at 12:02 is **absent**, even though the file was still open when it arrived. - The next cut, at 12:15, is the first copy that can contain it. An **acknowledged write** — one the caller was told had succeeded — is therefore not the same thing as a write that is inside a copy. The store answered the caller from memory; the copy is a separate, periodic act. ## Why one instant and not a smear If a copy were produced by walking the keyspace and recording whatever each entry happened to hold at the moment it was reached, restoring it would produce a state that **never existed**: one entry as of 12:00, another as of 12:03, a two-key change half applied. Stores avoid that in different ways — some hand the writing side an unchanging view of memory (a child process does the writing in some of them), some cut an incremental checkpoint and version whatever changes afterwards. What they share is the anchor to one instant. Engines that do permit a fuzzy copy make it restorable only by pairing it with a separate record of the writes that ran during it, which is a different mechanism entirely. ## What sits outside every copy 1. Take the instant the most recent copy was cut. 2. Every write acknowledged after that instant exists only in memory. 3. If the process dies now, those writes are gone. So the exposure is bounded by the copy interval, and in the worst case it *is* the interval: dying moments before the next cut loses almost all of it. Written the way a business can sign it: "we may lose up to 15 minutes of acknowledged writes." | Moment | What the live keyspace holds | What the newest usable copy holds | |---|---|---| | 11:59, next cut pending | everything written so far | the 11:45 state | | 12:00, copy cut | everything written so far | the 12:00 state | | 12:04, file closed | everything written so far | the 12:00 state | | 12:10, process dies | nothing | the 12:00 state | ## A whole copy, not a delta Each copy is self-contained: restoring it requires no other copy. Some stores assemble it internally from incremental checkpoints rather than by writing the whole keyspace out in one pass, but the artefact still stands for the entire keyspace at one instant. That is what makes "which writes are in it" answerable with a clock reading rather than a list of files. ## The interval you configured is not always the gap you have - A copy that fails partway normally leaves the previous one intact, because the new one is written aside and swapped in only on success — so the real gap is the one between **successful** copies. - If taking a copy comes to take longer than the interval, cuts start arriving while the previous write is still running, and the schedule you configured is fiction. - What tells you the truth is measured copy duration and copy success, not the number in the schedule. ## What the copy is actually worth The copy exists so a restart begins from a populated keyspace rather than an empty one — and that is the whole of what it confers. It does not make this tier a system of record, and it does not make data whose only home is this tier safe, because there is always a window of acknowledged writes on the wrong side of the last cut.

  • The copy takes four minutes to write and its file closes at 12:04. Which instant does it represent, and why does the distinction matter operationally?
    It represents 12:00, the instant it was cut. It matters because the exposure you quote is measured from the cut, not from the file timestamp: on a big keyspace the file's modification time can be minutes newer than the state inside it, and sizing the window from the file time understates what is lost.
  • A scheduled copy fails partway through and the next one succeeds an interval later. What is the practical effect?
    The previous copy is still there, so you are not left with nothing, but the real gap between usable copies is now two intervals instead of one — twice as many acknowledged writes sit outside every copy. This is why copy success and copy duration are monitored, rather than trusting the configured schedule.

saying these in an interview costs you the question

  • Says the copy holds everything up to the moment its file closed
  • Thinks writes accepted while the copy was written get folded into it
  • Assumes the configured interval is the real gap even after a copy fails
  • Treats an acknowledged write as already inside the copy
  • Calls data safe because a copy exists, with no other home for it
open as a page

While a whole copy of a live keyspace is being written, what can happen to resident memory, and what determines the size of that effect?

level: middleimportance: must knowfreq 58%

basics

~20 s

Resident memory can rise above the keyspace size while a copy is written: an unchanging image is held for the copy while callers change the live keyspace. The extra is sized by what changes during the write, not by keyspace size.

open as a page

Callers see tail write latency rise while a whole copy of the keyspace is being written; which parts of taking that copy explain it?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Taking a copy competes with serving: a brief pause to establish the unchanging image, a cost on the first change to each region after the cut, and device contention from pushing the whole keyspace out.

open as a page

For a 60 GB keyspace under heavy sustained writes, how do you choose a copy interval, and when does a whole-keyspace copy stop being worth taking?

level: principalimportance: should knowfreq 38%

basics

~20 s

A shorter copy interval shrinks the writes sitting outside the newest copy, while every copy costs memory, tail latency and device bandwidth. Choose the longest interval whose exposure the business will sign, and stop when copies cannot finish cheaply.

open as a page