For an in-memory store, how do a periodic whole copy and a replayed write log differ on writes lost, memory, latency and recovery time?
answer
- two currencies, not one
- writes lost against cost while serving
- state loads, operations re-execute
- divergence is write rate times duration
- each one's weakness is the other's strength
basics
~20 sA copy loses a window of time and costs memory while it is written, but reloads fast. A log loses only what missed the disk and costs the write path, but replaying it is slower. Teams pair them to get both.
solid answer
~50 sPrice each posture in two currencies at once. **Writes lost**: a periodic copy loses every write made since the last one finished, so the interval is the window; a log loses only what the flush policy has not pushed to disk, which is far less. **Cost while serving**: taking a whole copy of a live keyspace costs memory, because the copy and the keyspace diverge as writes continue for as long as the copy takes; a log instead taxes every write, and it grows until it is replaced by a shorter one that rebuilds the same state. **Recovery time**: loading a copy is typically faster, because a copy is state while a log is the operations that produced it — though a compacted log narrows that. That complementarity is why stores offering both are usually run with both.
go deeper
Focus on the first axis: a periodic copy loses everything written since the last copy, while a log loses only what had not reached disk. That single contrast is the foundation everything else in this comparison sits on.
This is your level's material. Be able to give both currencies for each posture without prompting — what it loses at restart and what it charges while the tier is serving — and to say why the two are commonly paired.
Bring the numbers' drivers, not the numbers: divergence is write rate times copy duration, replay time scales with operations rather than state, and compaction is a spike you have to have seen to expect. Then pick one for a stated workload.
Frame the choice as an ongoing tax against a one-off restart cost, and note who pays each. A per-write flush is a permanent latency budget line; an empty restart is a rare, concentrated load event somebody has to have capacity for.
## Two currencies, always The mistake this question exists to catch is answering in one currency. A candidate who has only read about these tiers says "the log loses less data" and stops. A candidate who has run one says "the log loses less data **and here is what it costs me every second the tier is up**". Every posture has a price at restart and a price while serving, and a choice made on one of them alone is not a choice. Four axes are enough to hold the whole comparison: **writes lost**, **memory**, **write-path latency**, and **recovery time**. ## Writes lost - A **periodic whole copy** loses a *window of time*. Whatever was written after the last copy finished is not on disk. Lengthen the interval and the window grows; shorten it and you pay the copy's cost more often. - A **replayed write log** loses *whatever has not reached disk*. Each state-changing write is appended, and the flush policy decides how promptly the appended bytes are flushed to disk — on every write, on a timer of about a second, or left to the operating system's discretion. The tighter the policy, the smaller the amount at risk. Neither reaches zero. A write acknowledged to the caller but not yet flushed to disk is the sharpest object in this whole subject, and it exists under both postures. ## What each costs while the tier is serving - **Taking a whole copy is not free because it runs in the background.** The keyspace carries on accepting writes while the copy is being written out, so the live keyspace and the copy diverge, and that divergence is memory. The term that drives it is *write rate times how long the copy takes*, not keyspace size — a mostly-read keyspace of any size diverges very little. Stores produce the copy differently, so the shape and the ceiling of that effect vary by store. - **A write log taxes the write path.** Every state-changing write does extra work, and the flush policy decides how much of that work is a disk wait. Flushing on every write buys the smallest window by paying per operation. The log also grows without bound until it is replaced with a shorter log that reconstructs the same state, and that replacement is itself a pause or a memory spike. ## Recovery time Recovery time is how long the process is unavailable at start because it is rebuilding the keyspace. - Loading a **copy** is typically the faster of the two: it is state, read and installed. - **Replaying a log** re-executes the operations that produced the state, so it is generally slower for the same keyspace. Compacting the log into a minimal set of operations narrows the gap considerably, and a log that has just been compacted can replay quickly. A tier keeping nothing has no recovery time and comes back empty. ## Side by side | Axis | Keep nothing | Periodic whole copy | Replayed write log | Keep both | |---|---|---|---|---| | Writes lost | Everything | Since the last copy | Since the last flush to disk | Since the last flush to disk | | Memory while serving | None | Divergence while a copy is written | Log buffering, plus the spike when it is compacted | Both | | Write-path latency | None | Largely unaffected between copies | Added per write, sized by the flush policy | Added per write | | Recovery time | None | Usually the fastest rebuild | Usually slower, narrowed by compaction | Copy load plus a short replay | ## Why both is the common answer where it is offered The two postures fail in complementary ways, which is exactly the shape that rewards combining them: the copy caps how much log a restart must replay, and the log covers the window the copy leaves open. That is why, on stores that offer both, running both is a defensible default — not because more persistence is automatically better, but because each one's weakness is the other's strength. Stores differ in how they arrange it. Some load the most recent copy and then replay only the part of the log written after it. Others fold a compact copy into the head of the log so replay begins from a small base. Do not present one arrangement as the mechanism. ## The answer that ends correctly Whichever posture wins, the conclusion is the same: the loss window shrinks, it never closes, and the tier does not become a system of record. What you bought is a shorter, smaller restart — less to rebuild, and less load pushed onto the system of record behind the tier while the tier is coming back.
- Why is the memory cost of taking a whole copy driven by write rate rather than by keyspace size?Because the cost is the *divergence* between the live keyspace and the copy being written, not a second copy of everything. Only what changes while the copy is in progress has to be held separately, so the term is write rate times copy duration. A large, mostly-read keyspace diverges very little; a small, write-heavy one can diverge a lot. Stores implement this differently, so the ceiling varies.
- If a write log loses less than a periodic copy, why would anyone run only the copy?Because the log's advantage is paid for on the write path of every single write, forever, while the copy's cost lands periodically and mostly as memory. For a tier holding reconstructible data where a window of lost writes is harmless, that ongoing tax buys nothing worth having. The copy also reloads a large keyspace faster, which matters more than the window for some workloads.
- Does tightening the flush policy to every write make an acknowledged write safe?It makes the window very small, not zero. The write still has to be appended and flushed after the store accepts it, and there is a moment where the caller has been told it succeeded and the bytes are not yet on disk. It also makes every write wait on the disk, which is the cost side of the trade and often the reason the tightest policy is not chosen.
saying these in an interview costs you the question
- Claims a replayed write log always recovers faster than loading a whole copy
- Treats a background copy as free because it does not block the caller
- Says flushing on every write costs nothing measurable on the write path
- Believes a growing write log never has to be compacted or replaced
- Picks a posture on writes lost alone and ignores the cost while serving