skip to content

A store restarted with a large write log stays unavailable for minutes while replaying it, so what drives that recovery time and what shortens it?

level: seniorimportance: should knowfreq 50%

answer

  1. a start is not instant
  2. time follows the log, not the data
  3. every entry applied, in order
  4. compaction is the lever on start time

basics

~20 s

Recovery time follows the number of entries in the write log and the cost of applying each one, so it grows with everything appended since the last compaction, not with live data. Compaction shortens it.

solid answer

~50 s

Replay at start applies every entry in the log, in order, from the beginning, so the duration is roughly the entry count multiplied by the cost of applying one, with a floor set by how fast the log can be read. The trap is assuming it scales with live data: a keyspace of a few gigabytes can sit behind a log of tens of gigabytes after heavy overwrites, and it is the log that is replayed. The levers are the ones that shorten the history - compacting more often, and holding less in the tier - plus whatever parallelism the store's loader offers, which varies. Most designs keep the process out of service until replay finishes; some answer with an explicit not-ready error rather than serving a half-rebuilt keyspace. The number worth having is measured, from a restart against a realistic log.

go deeper

for a junior

Know that a process keeping a write log is not instantly available after a restart - it has to read the log back and apply it first - and that the bigger the log, the longer that takes.

for a middle

Explain that replay cost follows entry count multiplied by per-entry apply cost, so it tracks history rather than live data, and name compaction as the lever that shortens it.

for a senior

Demonstrate having measured one. Separate read-bound from apply-bound replay, describe what the tier does to callers while it replays, and insist on a readiness signal that does not report healthy on a half-rebuilt keyspace.

for a principal

Treat recovery time as a number the design owes an answer for, watched through log length and re-measured as write volume grows. If replay is long enough to be an outage itself, say so and revisit the posture.

## What replay at start actually does Under the keep-a-replayed-write-log posture, a starting process reads the log from the beginning and applies each entry in order, rebuilding the keyspace operation by operation. It answers no caller while doing this. The duration is the process's **recovery time** - the window in which the tier exists but serves nothing. The important property is what the work is proportional to. Replay applies **entries**, so its cost follows the length of the history, not the size of the result. A keyspace of a few gigabytes sitting behind a log that has absorbed weeks of overwrites can take many times longer to start than a much larger keyspace whose log was compacted this morning. Two tiers with identical memory usage can have restart times an order of magnitude apart. ## What the time is proportional to | Factor | Effect on recovery time | |---|---| | Number of entries since the last compaction | The dominant term; roughly linear | | Cost of applying one entry | Multiplies the above - rebuilding a structured value costs more than assigning a plain one | | Device read throughput | A floor: the log's bytes have to be read before they can be applied | | Parallelism of the loader | Divides the apply cost where the store supports it; several do not, and per-key order constrains what can be parallelised | | Processor speed of the machine | Matters more than people expect, because replay is usually apply-bound rather than read-bound | The last row is the one that catches teams out. Because replay is often limited by applying operations rather than by reading bytes, moving the log to faster storage can leave the start time almost unchanged. Measure which of the two is the constraint before spending on either. ## Levers that shorten it 1. **Compact more often.** This is the direct lever: fewer entries to apply. It has its own price in disk, memory and a latency bump while it runs, so the trade is a shorter start against periodic disruption while serving. 2. **Hold less in the tier.** A compacted log cannot be shorter than what it takes to describe the live keyspace, so entries that did not need to be in this tier at all lengthen every future start. 3. **Use whatever parallel loading the store offers**, knowing it is not universal and that per-key ordering limits how much of the work can be divided. 4. **Keep a periodic copy alongside the log**, in stores that support both, so replay begins nearer the present - the mechanics of that copy are a separate subject. ## What varies between stores - **Behaviour while replaying.** Most designs keep the process out of service until it finishes. Some accept connections and answer with an explicit not-ready error, which is far better than serving a half-rebuilt keyspace but still is not service. - **Loader parallelism**, as above. - **Whether a hosted or managed deployment hides the start** behind something that holds connections, which changes what callers see but not how long the work takes. Because of this variation, "how long does a restart take" has no general answer and a very definite specific one. The specific one is obtainable only by doing it. ## Measure it before you need it The common shape of this incident is that nobody had ever restarted the process with a production-sized log. The tier had been running for months, the log had never been compacted, and the first unplanned restart turned a component people thought of as instant into a multi-minute outage. - **Time a start against a realistic log**, not an empty one, and record the number alongside the log's size so the relationship is visible. - **Make readiness reflect replay.** A health check that reports healthy while the keyspace is half rebuilt is worse than one that reports unhealthy, because callers get answers derived from a keyspace that is missing entries. - **Alert on log length**, not only on disk usage. Log length is the restart time, expressed in a unit you can watch. - **Re-measure after the workload changes.** Recovery time drifts upward silently as write volume grows, and nothing announces it until a restart does. The framing that holds all of this together: persistence in a volatile tier exists to **bound how empty a restart is**, and replay time is what that bound costs. If the replay is long enough to be an outage in its own right, the posture is not paying for itself.

  • Two tiers hold the same amount of live data but start at very different speeds. What explains it?
    The length of their write logs. Replay applies entries, and one of them has accumulated far more history since its last compaction - heavy overwrites of the same keys, or removals whose entries are still in the file. Live data is what the log rebuilds to, not what determines how much work rebuilding takes.
  • Would faster storage fix a long replay?
    Only if reading the log is the constraint. Replay is frequently limited by applying the operations - rebuilding values and inserting them into the keyspace - in which case processor speed matters more than device throughput and faster storage barely moves the number. Establish which of the two binds before spending on either.
  • What should a health check report while the process is replaying?
    Not ready. A keyspace that is half rebuilt answers with entries missing, which is worse than answering nothing, because callers cannot distinguish it from genuine absence. Stores differ in what they do natively - some refuse the connection, some return an explicit not-ready error - and whatever the store does, the readiness signal around it should not say healthy until replay ends.

saying these in an interview costs you the question

  • Assumes recovery time tracks live data rather than log length.
  • Thinks replay happens in the background while the tier serves normally.
  • Believes faster storage always fixes a long replay.
  • Has never timed a restart and assumes it takes seconds.
  • Treats compaction as a disk concern unrelated to start time.
  • Lets a health check report healthy while the keyspace is half rebuilt.