skip to content

A team says their in-memory tier "has persistence turned on" — what must you still ask before you can say what a crash costs?

level: juniorimportance: must knowfreq 62%

answer

  1. a capability, not a number
  2. which posture is in force
  3. then the copy interval or flush policy
  4. seconds of acknowledged writes
  5. some stores offer no posture at all

basics

~20 s

"Persistence is on" is not a loss window. Ask which posture is in force and the flush policy or copy interval that governs it — those numbers, not the word, say how many seconds of acknowledged writes a crash costs.

solid answer

~50 s

The phrase names a capability, not a number. To say what a crash costs you need two things: the **posture** in force — keep nothing, keep a periodic whole copy, keep a replayed write log, or keep both — and the interval that governs it, which is the copy interval for a periodic copy and the flush policy for a write log. Only then can you write the sentence that actually answers the question: *we may lose up to N seconds of acknowledged writes*. An acknowledged write is one the caller was told had succeeded, and the window is the gap between that acknowledgement and the bytes reaching disk. It is also worth asking whether the store offers a posture at all — some stores in this class keep nothing across a restart by design and have no setting that changes it.

go deeper

for a junior

Recall that the answer depends on what the store keeps: nothing, a periodic whole copy, a replayed write log, or both. Then recall that a posture alone is not a number — you also need its interval.

for a middle

Explain the mechanics behind each posture's window: the copy interval bounds one, the flush policy bounds the other, and the flush policy's three granularities trade window size against cost on the write path.

for a senior

Show that you would go and read the actual configuration rather than repeat what a store is reputed to do, and that you would state the result as seconds and as a count of records at peak.

for a principal

Treat the unexamined setting as the defect. The interesting question is who signed off on this window, what they were told it meant, and whether any state on the tier deserves a smaller one.

## What "persistence is on" leaves out An in-memory store answers from memory, so everything it holds disappears when the process dies unless the store has been writing something to disk as it went. Saying persistence is "on" names a capability. The question anyone actually needs answered — a business owner, an incident review, an interviewer — is a different one: *when this process dies, how much of what we told callers had succeeded do we lose?* That is answered with one sentence carrying one number: **we may lose up to N seconds of acknowledged writes.** An **acknowledged write** is a write the caller was told had succeeded. The entire subject here is the gap between that acknowledgement and the moment the bytes are on disk — a write acknowledged to the caller but not yet flushed to disk is the exposure you are budgeting. ## The two things you must name to get a number 1. **The posture in force.** A store of this class is in one of four: keep nothing, keep a periodic whole copy, keep a replayed write log, or keep both. 2. **The interval that governs that posture.** For a periodic copy that is the copy interval; for a write log it is the flush policy — how often appended bytes are flushed to disk. Without both, "how much do we lose?" has no answer, because the honest answers range from *everything since the process started* to *the last fraction of a second*. | Posture | In memory after a restart | The window you can state | |---|---|---| | Keep nothing | an empty tier | every write since the process last started | | Keep a periodic whole copy | the keyspace as of the last copy | up to one copy interval of acknowledged writes | | Keep a replayed write log | the keyspace replayed from the log | up to one flush interval of acknowledged writes | | Keep both | whichever source the store restores from | that source's window — stores differ in which one they prefer and how they combine the two | ## Not every store has a posture to name Stores in this class differ here more than almost anywhere else, and the difference is not a setting. Some are built to keep nothing across a restart and offer nothing that changes that; on those the sentence is fixed — *a restart costs every write made since the process started* — and everything around the tier has to be designed for it. Others offer a periodic copy, a write log, or both, and some let a caller ask for stronger durability on an individual write. Asking "which posture?" and hearing "this store has none" is a complete answer, not a gap. ## The flush policy is the number that actually moves Where a write log exists, its flush policy has three granularities: - **flush on every write** — the smallest window, paid for on the write path by every single operation; - **flush on a timer (about a second)** — a window of roughly that timer; - **leave flushing to the operating system** — no bound you control, because the window is whatever the operating system had not yet written, which can be far larger than a second when the machine is under memory pressure. Notice the shape of that list: as the window shrinks, the cost moves onto the write path. That trade is what the sentence is really about. ## Writing the sentence 1. Name the posture and the interval that governs it. 2. State the window in seconds: *we may lose up to N seconds of acknowledged writes.* 3. Multiply by the write rate at the busiest moment, so seconds become records — "about three thousand orders" is something a non-engineer can judge; "one second" is not. 4. Say what those records are, and read the sentence to whoever owns them. ## The default nobody chose This is an interview question because most teams never made the choice. A store was installed, whatever it shipped with stayed, and the loss window is therefore an accident rather than a decision. That is what the interviewer is probing: not whether the window is one second or five minutes, but that nobody can say which, and nobody has put the sentence in front of the person who would have to live with the consequence. ## What the sentence does not promise Persistence here bounds a restart. It decides whether the tier comes back empty or nearly current, and how much replay at start costs in recovery time. It does not turn the tier into a durable system of record: under most flush policies the caller is told the write succeeded before the bytes are safe, the disk beneath it is a single failure domain, and the whole component is built around memory as its medium. The sentence is a budget for a bounded, survivable loss — which is exactly why this tier is never the only copy of anything that matters. Keep the scope honest, too. The sentence describes a crash or restart of this single process. An acknowledged write lost because another node was promoted without it has a different cause and a different remedy, and does not belong in this number.

  • If the answer is "this store keeps nothing across a restart and has no setting for it", have you failed to state a loss window?
    No — that is the loss window, fully stated: every write made since the process last started is gone. It is the sharpest version of the sentence, and it shapes the design around the tier, because nothing held there may be the only copy of anything and a restart must be survivable by the system of record behind it.
  • Why turn the seconds into a count of records before showing the sentence to anyone?
    Seconds are meaningless to whoever owns the data. Multiplying the window by the write rate at the busiest moment converts it into "about M records of this kind", which is the form someone can accept or refuse. It also exposes the fact that the same one-second window costs far more at peak than at night.
  • Does shortening the loss window also shorten the time the tier is unavailable after a crash?
    No, and conflating the two is common. The loss window is how much is gone; recovery time is how long replay at start keeps the process unavailable. Flushing more often shrinks the first and does nothing for the second — a longer accumulated write log usually makes recovery time worse, not better.

saying these in an interview costs you the question

  • Says persistence being on means no acknowledged write can be lost.
  • Answers with a single number without naming the posture in force.
  • Assumes every store of this class can be made to keep data across a restart.
  • Treats a write log as making the tier a system of record.
  • Assumes whatever the store shipped with was chosen for this workload.
  • Confuses how much is lost with how long the restart takes.