skip to content

Durability Options

What a volatile tier keeps across a restart and what keeping it costs while serving: nothing, a periodic whole copy, or a replayed log of writes. Asked because most teams never actually chose.

on this pageshow

questions

20

A team says their in-memory tier "has persistence turned on" — what must you still ask before you can say what a crash costs?

level: juniorimportance: must knowfreq 62%

answer

  1. a capability, not a number
  2. which posture is in force
  3. then the copy interval or flush policy
  4. seconds of acknowledged writes
  5. some stores offer no posture at all

basics

~20 s

"Persistence is on" is not a loss window. Ask which posture is in force and the flush policy or copy interval that governs it — those numbers, not the word, say how many seconds of acknowledged writes a crash costs.

solid answer

~50 s

The phrase names a capability, not a number. To say what a crash costs you need two things: the **posture** in force — keep nothing, keep a periodic whole copy, keep a replayed write log, or keep both — and the interval that governs it, which is the copy interval for a periodic copy and the flush policy for a write log. Only then can you write the sentence that actually answers the question: *we may lose up to N seconds of acknowledged writes*. An acknowledged write is one the caller was told had succeeded, and the window is the gap between that acknowledgement and the bytes reaching disk. It is also worth asking whether the store offers a posture at all — some stores in this class keep nothing across a restart by design and have no setting that changes it.

go deeper

for a junior

Recall that the answer depends on what the store keeps: nothing, a periodic whole copy, a replayed write log, or both. Then recall that a posture alone is not a number — you also need its interval.

for a middle

Explain the mechanics behind each posture's window: the copy interval bounds one, the flush policy bounds the other, and the flush policy's three granularities trade window size against cost on the write path.

for a senior

Show that you would go and read the actual configuration rather than repeat what a store is reputed to do, and that you would state the result as seconds and as a count of records at peak.

for a principal

Treat the unexamined setting as the defect. The interesting question is who signed off on this window, what they were told it meant, and whether any state on the tier deserves a smaller one.

## What "persistence is on" leaves out An in-memory store answers from memory, so everything it holds disappears when the process dies unless the store has been writing something to disk as it went. Saying persistence is "on" names a capability. The question anyone actually needs answered — a business owner, an incident review, an interviewer — is a different one: *when this process dies, how much of what we told callers had succeeded do we lose?* That is answered with one sentence carrying one number: **we may lose up to N seconds of acknowledged writes.** An **acknowledged write** is a write the caller was told had succeeded. The entire subject here is the gap between that acknowledgement and the moment the bytes are on disk — a write acknowledged to the caller but not yet flushed to disk is the exposure you are budgeting. ## The two things you must name to get a number 1. **The posture in force.** A store of this class is in one of four: keep nothing, keep a periodic whole copy, keep a replayed write log, or keep both. 2. **The interval that governs that posture.** For a periodic copy that is the copy interval; for a write log it is the flush policy — how often appended bytes are flushed to disk. Without both, "how much do we lose?" has no answer, because the honest answers range from *everything since the process started* to *the last fraction of a second*. | Posture | In memory after a restart | The window you can state | |---|---|---| | Keep nothing | an empty tier | every write since the process last started | | Keep a periodic whole copy | the keyspace as of the last copy | up to one copy interval of acknowledged writes | | Keep a replayed write log | the keyspace replayed from the log | up to one flush interval of acknowledged writes | | Keep both | whichever source the store restores from | that source's window — stores differ in which one they prefer and how they combine the two | ## Not every store has a posture to name Stores in this class differ here more than almost anywhere else, and the difference is not a setting. Some are built to keep nothing across a restart and offer nothing that changes that; on those the sentence is fixed — *a restart costs every write made since the process started* — and everything around the tier has to be designed for it. Others offer a periodic copy, a write log, or both, and some let a caller ask for stronger durability on an individual write. Asking "which posture?" and hearing "this store has none" is a complete answer, not a gap. ## The flush policy is the number that actually moves Where a write log exists, its flush policy has three granularities: - **flush on every write** — the smallest window, paid for on the write path by every single operation; - **flush on a timer (about a second)** — a window of roughly that timer; - **leave flushing to the operating system** — no bound you control, because the window is whatever the operating system had not yet written, which can be far larger than a second when the machine is under memory pressure. Notice the shape of that list: as the window shrinks, the cost moves onto the write path. That trade is what the sentence is really about. ## Writing the sentence 1. Name the posture and the interval that governs it. 2. State the window in seconds: *we may lose up to N seconds of acknowledged writes.* 3. Multiply by the write rate at the busiest moment, so seconds become records — "about three thousand orders" is something a non-engineer can judge; "one second" is not. 4. Say what those records are, and read the sentence to whoever owns them. ## The default nobody chose This is an interview question because most teams never made the choice. A store was installed, whatever it shipped with stayed, and the loss window is therefore an accident rather than a decision. That is what the interviewer is probing: not whether the window is one second or five minutes, but that nobody can say which, and nobody has put the sentence in front of the person who would have to live with the consequence. ## What the sentence does not promise Persistence here bounds a restart. It decides whether the tier comes back empty or nearly current, and how much replay at start costs in recovery time. It does not turn the tier into a durable system of record: under most flush policies the caller is told the write succeeded before the bytes are safe, the disk beneath it is a single failure domain, and the whole component is built around memory as its medium. The sentence is a budget for a bounded, survivable loss — which is exactly why this tier is never the only copy of anything that matters. Keep the scope honest, too. The sentence describes a crash or restart of this single process. An acknowledged write lost because another node was promoted without it has a different cause and a different remedy, and does not belong in this number.

  • If the answer is "this store keeps nothing across a restart and has no setting for it", have you failed to state a loss window?
    No — that is the loss window, fully stated: every write made since the process last started is gone. It is the sharpest version of the sentence, and it shapes the design around the tier, because nothing held there may be the only copy of anything and a restart must be survivable by the system of record behind it.
  • Why turn the seconds into a count of records before showing the sentence to anyone?
    Seconds are meaningless to whoever owns the data. Multiplying the window by the write rate at the busiest moment converts it into "about M records of this kind", which is the form someone can accept or refuse. It also exposes the fact that the same one-second window costs far more at peak than at night.
  • Does shortening the loss window also shorten the time the tier is unavailable after a crash?
    No, and conflating the two is common. The loss window is how much is gone; recovery time is how long replay at start keeps the process unavailable. Flushing more often shrinks the first and does nothing for the second — a longer accumulated write log usually makes recovery time worse, not better.

saying these in an interview costs you the question

  • Says persistence being on means no acknowledged write can be lost.
  • Answers with a single number without naming the posture in force.
  • Assumes every store of this class can be made to keep data across a restart.
  • Treats a write log as making the tier a system of record.
  • Assumes whatever the store shipped with was chosen for this workload.
  • Confuses how much is lost with how long the restart takes.
open as a page

An in-memory store restarts empty and takes traffic again. What does the system of record behind it experience, and how large is that effect?

level: juniorimportance: must knowfreq 64%

basics

~20 s

Until entries accumulate again, every read lands on the system of record behind the tier. The multiplier is the inverse of the miss rate: at a 95% hit ratio that system briefly sees twenty times its normal read load.

open as a page

A store writes a whole copy of its keyspace to disk every 15 minutes while still accepting writes; which writes does that copy hold?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A point-in-time copy holds the keyspace as of the instant it was cut, not when its file finished writing. Writes accepted after the cut are absent, so up to one whole copy interval of acknowledged writes can be missing.

open as a page

What does an in-memory store hold after a restart if it keeps nothing, a periodic copy, a write log, or both?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Keep nothing gives an empty store. A periodic whole copy gives the keyspace as of the last copy. A write log gives it replayed to the last flush to disk. Keeping both replays the log on top of a copy.

open as a page

When an in-memory store is configured to keep a write log, what is appended while it runs and what happens to that log at start?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A write log records every state-changing operation in the order it was accepted. At start the process replays that log from the beginning, applying each entry again, to rebuild the keyspace it held before it stopped.

open as a page

An in-memory tier keeps a write log flushed on a one-second timer, plus a copy every five minutes — what is the loss window?

level: middleimportance: must knowfreq 57%

basics

~20 s

Roughly a second, not five minutes: the write log holds everything acknowledged since the last copy, so the flush timer governs. The five-minute copy interval is the fallback window if the log cannot be replayed.

open as a page

While a whole copy of a live keyspace is being written, what can happen to resident memory, and what determines the size of that effect?

level: middleimportance: must knowfreq 58%

basics

~20 s

Resident memory can rise above the keyspace size while a copy is written: an unchanging image is held for the copy while callers change the live keyspace. The extra is sized by what changes during the write, not by keyspace size.

open as a page

For an in-memory store, how do a periodic whole copy and a replayed write log differ on writes lost, memory, latency and recovery time?

level: middleimportance: must knowfreq 61%

basics

~20 s

A copy loses a window of time and costs memory while it is written, but reloads fast. A log loses only what missed the disk and costs the write path, but replaying it is slower. Teams pair them to get both.

open as a page

In a store that appends every write to a write log, which flush granularities decide when those appends reach disk, and what does each cost?

level: middleimportance: must knowfreq 66%

basics

~20 s

A flush policy decides when appended writes reach disk: flush on every write, flush on a timer of about a second, or leave flushing to the operating system. Each trades write-path cost against writes at risk.

open as a page

A restarted in-memory store can be filled from a point-in-time copy or by pre-loading a chosen key set. What does each buy?

level: middleimportance: should knowfreq 46%

basics

~20 s

Restoring from a copy returns the whole keyspace as it was, with no read load behind the tier, but it costs time before the node serves and comes back stale. Pre-loading returns only what you chose, current.

open as a page

What does compacting a write log replace it with, what triggers it, and what does the compaction itself cost while the store serves?

level: middleimportance: should knowfreq 54%

basics

~20 s

Compacting replaces an accumulated write log with a shorter one that reconstructs the same state, dropping history later overwritten or removed. Growth thresholds, a schedule or an operator trigger it, and it costs disk, memory and a latency bump.

open as a page

A team wants to hold state with no other home on an in-memory tier whose loss window is one second — how do you test whether that is acceptable?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Convert the window into records at peak, then ask what one lost record costs and who finds out. For a reconstructible copy the window is a performance question; for state with no other home it is a correctness budget.

open as a page

You must restart every node of a six-node in-memory tier with no node promoted in its place. How do you choose the order and pacing so the system of record behind it survives?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Name what one node's restart removes - a share of the keyspace, or a share of capacity. Then one node at a time, cheapest share first, and no next node until the load behind the tier is back inside its headroom.

open as a page

Callers see tail write latency rise while a whole copy of the keyspace is being written; which parts of taking that copy explain it?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Taking a copy competes with serving: a brief pause to establish the unchanging image, a cost on the first change to each region after the cut, and device contention from pushing the whole keyspace out.

open as a page

If one candidate in-memory store keeps nothing across a restart by design and another offers a periodic copy and a write log, how should that change your design?

level: seniorimportance: should knowfreq 49%

basics

~20 s

It changes how big an empty restart is, not whether you need a system of record behind the tier. Either way, everything in the tier must be reconstructible or its loss must be acceptable; persistence only shortens and shrinks the restart.

open as a page

A store restarted with a large write log stays unavailable for minutes while replaying it, so what drives that recovery time and what shortens it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Recovery time follows the number of entries in the write log and the cost of applying each one, so it grows with everything appended since the last compaction, not with live data. Compaction shortens it.

open as a page

How do you write an in-memory tier's loss window as a sentence the business signs, and what if they refuse it?

level: principalimportance: should knowfreq 34%

basics

~20 s

One sentence with a number and a noun: if this process dies we may lose up to N seconds of acknowledged writes — about M orders. If refused, the levers are flushing every write, or another home for the state.

open as a page

For a 60 GB keyspace under heavy sustained writes, how do you choose a copy interval, and when does a whole-keyspace copy stop being worth taking?

level: principalimportance: should knowfreq 38%

basics

~20 s

A shorter copy interval shrinks the writes sitting outside the newest copy, while every copy costs memory, tail latency and device bandwidth. Choose the longest interval whose exposure the business will sign, and stop when copies cannot finish cheaply.

open as a page

For an in-memory tier, when is "keep nothing across a restart and make the restart cheap" the honest choice rather than an unmade decision?

level: principalimportance: should knowfreq 38%

basics

~20 s

It is honest when everything in the tier is reconstructible, the system of record behind it can absorb a full rebuild, and the team has weighed that against what a keeping posture would cost while serving. Otherwise it is a default nobody chose.

open as a page

When is the right answer to accept an empty restart and make restarting cheap, rather than engineering restores and pre-loading for a whole fleet?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Accept it when the tier holds only reconstructible entries and the system of record behind it has been measured absorbing the full read rate. Engineer around it when that capacity is a hard ceiling, or entries have no other home.

open as a page