skip to content

When an in-memory store is configured to keep a write log, what is appended while it runs and what happens to that log at start?

level: juniorimportance: must knowfreq 68%

answer

  1. only state changes are recorded
  2. order is the whole mechanism
  3. read back from the beginning at start
  4. appending is not yet flushing

basics

~20 s

A write log records every state-changing operation in the order it was accepted. At start the process replays that log from the beginning, applying each entry again, to rebuild the keyspace it held before it stopped.

solid answer

~40 s

The write log is one of the postures a volatile tier can take toward a restart: keep a replayed write log. While the process serves, every operation that changes state is appended to the end of a file in the order the server accepted it; reads are not appended, because replaying something that changed nothing rebuilds nothing. At start the process reads the log from the beginning and applies each entry again, and when it reaches the end the keyspace matches what the log describes. Stores differ in what they write down - some append the operation as the caller sent it, others rewrite anything depending on a clock or a random source into its resulting effect so replay is reproducible. Several stores of this class keep no write log at all, by design.

go deeper

for a junior

Recall the two halves and say them in order: every state-changing operation is appended while the process runs, and at start the process reads the file from the beginning and applies the entries again to rebuild the keyspace.

for a middle

Explain why order matters, why reads are not logged, and why appending is a different act from flushing to disk. Note that what is written down varies - some stores log the operation as sent, others its resulting effect.

for a senior

Show you have started one of these processes. Mention that the tier is unavailable while it replays, that the duration tracks log length rather than live data, and that a crash can leave a partial trailing entry stores handle differently.

for a principal

Frame it as bounding a restart rather than buying durability. Say plainly that several stores in this class keep nothing on purpose, and that state whose only copy lives here is a design problem the write log does not solve.

## The posture this belongs to A volatile tier takes one of four postures toward a restart: keep nothing, keep a periodic whole copy, keep a replayed write log, or keep both. This question is about the write log. It has exactly two halves - **append on write** while the process is running, and **replay at start** when it comes back - and everything else about it follows from those two. ## Append on write While the process serves traffic, each operation that **changes state** is written to the end of a file, in the order the server accepted it. Three properties carry the whole mechanism: - **Append-only.** Entries go on the end; nothing already written is edited in place. That is why appending is cheap, and also why the file only grows. - **Ordered.** The log is a history, not a set of final values. A key written five times appears five times, and the fifth value wins only because it is applied last. - **Changes only.** A read appends nothing. Replaying a read would rebuild nothing, so logging it would spend write bandwidth for no recovery value. What exactly gets written down varies more than people expect: | Event while serving | Appended? | |---|---| | An operation that changes a value | Yes - this is what the log exists for | | A read, or an operation that changes nothing | No - replaying it would rebuild nothing | | An operation whose result depends on a clock or a random source | Yes, but several stores rewrite it into its resulting effect so that replay is reproducible | | An entry removed because its lifetime lapsed | Varies - some stores append an explicit removal, others record the deadline with the entry and let replay apply it again | Whether the append happens before or after the change is applied in memory also differs between stores, and so does whether the caller is answered before or after the bytes are flushed to disk. Appending is **not** the same act as flushing to disk: the bytes normally sit in the operating system's buffers first, which is exactly why a **flush policy** exists as a separate setting. ## Replay at start Under this posture, a starting process does the following: 1. Begins with an empty keyspace. 2. Opens the log and reads it from the beginning. 3. Applies each entry in order, as if a client had sent it, but answers nobody. 4. Reaches the end, at which point the keyspace matches what the log describes, and opens to traffic. Order is not optional. Applying the log backwards, or in parallel without respecting per-key order, would let an earlier, stale value overwrite a later one. Replaying forward from the beginning is what makes last-write-wins come out the same way it did the first time. Two consequences follow immediately: - The process is **not serving while it replays**. That duration is its recovery time, and it grows with how many entries the log holds - not with how much live data the keyspace ends up containing. - A crash can leave a **trailing partial entry**, because the process died part-way through an append. Stores differ in what they do with it: some truncate the incomplete tail and carry on, others refuse to start until an operator repairs the file. ## What the write log is not - **Not a promise that nothing is lost.** What is at risk is everything appended but not yet flushed to disk, and the size of that depends entirely on the flush policy in force. - **Not a licence to make the tier the only copy.** The log bounds how empty a restart is; it does not give a volatile tier the guarantees a durable engine offers. State with no other home still deserves a hard look. - **Not universal.** Several widely deployed stores in this class keep nothing across a restart and expose no setting that changes it. That is a design choice - cheap restarts, nothing on the write path, no files to operate - not a missing feature, and a design that assumes a log exists will not port to them. - **Not free to keep.** The file grows with every change, so a log nobody ever shortens eventually costs disk and makes replay slow; the repair for that is compacting it.

  • Why is the write log replayed from the beginning rather than from the end?
    Because it is a history, not a set of final values. The same key can appear many times, and the last entry wins only because it is applied last. Reading backwards, or applying entries out of order, would let an earlier stale value overwrite a later one and rebuild a keyspace that never existed.
  • What happens at start if the process died part-way through appending an entry?
    The log ends in a partial entry that describes no complete operation. Stores differ: some truncate the incomplete tail and replay what precedes it, others refuse to start until an operator repairs or truncates the file. Either way that last operation is gone, and the caller may already have been told it succeeded.

saying these in an interview costs you the question

  • Thinks reads are appended to the write log alongside writes.
  • Believes an appended entry is already on disk.
  • Assumes every in-memory store keeps a write log of some kind.
  • Treats the log as making the tier safe as the only copy of data.
  • Expects the tier to answer normally while it is still replaying.