skip to content

Three Honest Postures

Three postures a volatile tier can take on a restart - keep nothing, keep a periodic whole copy, keep a replayed write log - plus the common pairing of the last two, and what each costs.

on this pageshow

questions

4

What does an in-memory store hold after a restart if it keeps nothing, a periodic copy, a write log, or both?

level: juniorimportance: must knowfreq 72%

answer

  1. four answers, not three
  2. nothing, copy, log, or both
  3. ask what is in memory at start
  4. the copy's loss is a time window
  5. the log's loss is what missed disk

basics

~20 s

Keep nothing gives an empty store. A periodic whole copy gives the keyspace as of the last copy. A write log gives it replayed to the last flush to disk. Keeping both replays the log on top of a copy.

solid answer

~40 s

There are four answers in use. **Keep nothing**: the process comes back empty, and everything it held is gone. **Keep a periodic whole copy**: the store loads the most recent completed copy, so every write made after that copy is lost. **Keep a replayed write log**: each state-changing write is appended to a log on disk and replayed at start, so what is lost is whatever had not been flushed to disk yet. **Keep both**: the copy caps how much of the log has to be replayed, which shortens the restart. Not every store in this class offers all four — some keep nothing by design and have no setting that changes it, so "an in-memory store writes a copy in the background" is true of some products and false of others.

go deeper

for a junior

Memorise the four answers and what each leaves in memory: empty, the keyspace as of the last copy, the log replayed to its last flush, or a copy with the log's tail on top. Being able to list them cleanly is most of the value here.

for a middle

Be able to say why the two keeping postures lose different things — a copy loses a window of time, a log loses whatever missed the disk — and why combining them shortens the restart rather than duplicating effort.

for a senior

Say out loud that some stores in this class keep nothing by design, so the posture is a property of the store you picked as well as a setting. Interviewers listen for whether you describe the class or one product's behaviour.

for a principal

Treat the posture as a decision with a written owner. The value of naming all four is that it exposes the one your platform is running by default, which is almost never the one anybody argued for.

## The question every deployment has already answered An in-memory store answers from memory, and memory does not outlive the process. So every deployment of one has already answered a question, whether or not anyone in the room asked it out loud: **when this process comes back, what is in it?** There are four answers in general use — three postures a store can take on its own, plus the pairing of the last two. ## The four postures 1. **Keep nothing.** The keyspace exists only while the process is alive. On restart the store is empty and begins serving immediately. 2. **Keep a periodic whole copy.** Every so often the store writes the whole keyspace out as a point-in-time copy. On restart it loads the most recent copy that finished. 3. **Keep a replayed write log.** Each write that changes state is appended to a log on disk, and at start the store replays that log to rebuild the keyspace. 4. **Keep both.** A copy is used so that only a short tail of the log has to be replayed at start. Not every store in this class offers all four. **Some keep nothing by design and expose no setting that changes it**; others offer a copy, a log, or both. Asserting "an in-memory store writes a copy in the background" as the model of the category is the most common way to get this subject wrong — it is a true statement about some products in this family and a false one about others. ## What is in memory after a restart | Posture | What comes back | What is gone | |---|---|---| | Keep nothing | An empty keyspace | Everything the tier held | | Periodic whole copy | The keyspace as of the last copy that completed | Every write made after that copy | | Replayed write log | The keyspace replayed to the last write flushed to disk | Writes appended but not yet flushed | | Keep both | Copy state with the log's tail replayed on top | Writes not yet flushed | Two shapes do the work in that table: - A copy is **whole and periodic**, so what it loses is *a window of time* — everything written since the last one finished. - A log is **per-write and appended**, so what it loses is *whatever has not reached disk yet*. Stores that offer a log let you choose how often the flush to disk happens, and that choice moves the amount at risk. The pairing exists because those two shapes are complementary: the copy caps how far back a replay has to begin, and the log covers the window the copy leaves open. Stores combine them differently — some load the copy and then replay only the part of the log written after it, others fold a compact copy into the head of the log so that replay starts from a small base. The effect is the same; the arrangement is not universal. ## What each posture costs while the tier is serving A posture is never free, and the bill arrives in two different places: - **Keep nothing** costs nothing on the write path. The entire price is paid at restart, as an empty tier. - **Keep a periodic whole copy** costs memory and some latency while a copy is being written, because the live keyspace keeps accepting writes and diverges from the copy being written out. How large that gets is driven by write rate times how long the copy takes — and the mechanism differs per store, so the size of the effect is not a constant of the category. - **Keep a replayed write log** adds work to every write, and how much depends on how often the log is flushed to disk. The log also grows, and periodically has to be replaced with a shorter one that reconstructs the same state, which is itself a pause or a memory spike. - **Keep both** pays both bills and buys the shortest restart of the three keeping choices. ## Recovery time is the fourth axis Recovery time is how long the process is unavailable because it is rebuilding its keyspace at start. Loading a copy is typically faster than replaying a log, because a copy is *state* and a log is *the operations that produced that state* — replay re-executes them. The gap narrows when the log has been compacted down to a minimal set of operations, and it widens with keyspace size. A tier that keeps nothing has no recovery time at all; it comes back empty and serves at once. ## What none of them make the tier None of these postures turn a volatile tier into a system of record. Even with both, a write can be acknowledged to the caller and then lost — the loss window shrinks, it does not close. What persistence buys here is a **bounded restart**: less to rebuild from scratch, less work thrown at the system of record behind the tier, and a faster return to useful hit ratios. That is the honest summary of the whole subject: the tier is never the only copy of anything that matters.

  • Why does keeping both a periodic copy and a write log make sense rather than being redundant?
    They fail differently. The copy is compact state, so it rebuilds a large keyspace quickly, but it leaves a window of writes behind it. The log covers that window write by write, but replaying a long one is slow because it re-executes operations. Used together, the copy caps how much log has to be replayed while the log covers the copy's window. Stores arrange that pairing differently, but the purpose is the same.
  • If a store keeps nothing across a restart, does that make it a worse choice?
    No — it makes it a different choice. A store that keeps nothing is simple, spends nothing on the write path, and starts instantly. It is the wrong choice only if you were relying on the tier to hold state that has no other home, or if an empty restart puts more load on the system behind than that system can take. Both of those are properties of your design, not defects of the store.
  • A write is acknowledged to the caller and the process dies a moment later. Is it safe under a write log?
    Not necessarily. Acknowledgement means the store accepted the write into memory and appended it to the log; whether the log has reached disk depends on the flush policy in force. A write acknowledged to the caller but not yet flushed to disk is exactly the write that disappears. Tightening the flush policy shrinks that gap; it does not eliminate it.

A text editor offers the same four postures. You can work in a scratch buffer that vanishes when the editor dies; you can let it autosave the whole file every few minutes, so a crash costs you the minutes since the last autosave; you can have it journal every keystroke, so a crash costs only what the journal had not written out; or you can have both, so the file opens from the last autosave and the journal supplies the rest. None of them make the editor your version control system.

saying these in an interview costs you the question

  • Says every store in this class can be configured to keep data across a restart
  • Thinks a write log lets the tier be the only copy of the data
  • Assumes a periodic copy is continuous, so nothing is lost between copies
  • Says the store comes back holding exactly what it held at the crash
  • Describes one product's file names and settings as if they were the postures
open as a page

For an in-memory store, how do a periodic whole copy and a replayed write log differ on writes lost, memory, latency and recovery time?

level: middleimportance: must knowfreq 61%

basics

~20 s

A copy loses a window of time and costs memory while it is written, but reloads fast. A log loses only what missed the disk and costs the write path, but replaying it is slower. Teams pair them to get both.

open as a page

If one candidate in-memory store keeps nothing across a restart by design and another offers a periodic copy and a write log, how should that change your design?

level: seniorimportance: should knowfreq 49%

basics

~20 s

It changes how big an empty restart is, not whether you need a system of record behind the tier. Either way, everything in the tier must be reconstructible or its loss must be acceptable; persistence only shortens and shrinks the restart.

open as a page

For an in-memory tier, when is "keep nothing across a restart and make the restart cheap" the honest choice rather than an unmade decision?

level: principalimportance: should knowfreq 38%

basics

~20 s

It is honest when everything in the tier is reconstructible, the system of record behind it can absorb a full rebuild, and the team has weighed that against what a keeping posture would cost while serving. Otherwise it is a default nobody chose.

open as a page