skip to content

Log Flush, fsync and OS Page Cache Durability

Why Kafka relies on replication rather than fsync for durability, and the loss window that leaves on correlated power failure. A sharp question, because many candidates assume every write is flushed to disk.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

After a Kafka write returns successfully, where does the data physically live, and why does that distinction matter?

level: juniorimportance: must knowfreq 60%

answer

  1. success = in RAM (page cache), not on disk yet
  2. page cache is volatile
  3. OS flushes to disk later
  4. safety = many RAM copies on different brokers
  5. correlated power loss is the danger

basics

~20 s

After a successful write the data is in the broker's RAM (the OS page cache), and with acks=all also in the RAM of other replicas. It isn't necessarily on the physical disk yet — the operating system writes it to disk a bit later.

solid answer

~40 s

When a Kafka write succeeds, the record has been appended to the leader's log, which means it's in that broker's OS page cache (RAM) — and with acks=all, also in the page cache of every in-sync replica. It is NOT guaranteed to be on the physical disk at that moment; the kernel flushes dirty page-cache pages to disk asynchronously. This distinction matters because page cache is volatile: if a broker loses power, anything in its page cache that hasn't been flushed is lost. Kafka accepts this because the record also lives in the page cache of other brokers, so as long as one survives, the data survives. Durability comes from having multiple in-memory copies on independent machines, not from each copy being flushed to disk.

go deeper

for a junior

Know that after a successful write the data is in RAM/page cache (and replicas' RAM), not necessarily on disk yet.

for a middle

Explain the write → page cache → async OS flush path and why multiple RAM copies provide durability.

for a senior

Distinguish volatile vs. stable storage and reason about correlated failure across replicas.

for a principal

Design replica placement (racks/AZs) so independent power makes the in-RAM model safe at fleet scale.

## The journey of a record 1. A producer sends a record to the **leader** broker for a partition. 2. The leader does a normal file `write()` appending the record to the partition's active log segment. On Linux, this `write()` puts the bytes into the **OS page cache** — a region of RAM the kernel uses to buffer file data — and marks those pages "dirty." 3. With **acks=all**, the leader waits for every **in-sync replica** to fetch and append the same record (also into *their* page cache) before acknowledging the producer. 4. **Later**, asynchronously, the kernel writes the dirty pages out to the physical disk (an SSD or HDD). This is the actual *flush*. Until that happens, the only persisted copies are in volatile RAM. ## Page cache vs. disk - **Page cache** = RAM. Fast, but **volatile** — its contents vanish on power loss or a hard crash. - **Disk** = stable storage. Survives power loss, but slower to write to. When a Kafka write returns, the data is in page cache (RAM). It reaches disk only when the OS flushes, or when Kafka's optional `flush.messages`/`flush.ms` thresholds force an `fsync`. ## Why the distinction matters Because acknowledged data may be only in RAM, a sudden power loss on a broker can lose page-cache data that hadn't been flushed. Kafka tolerates this by keeping **multiple copies in RAM on independent brokers**. If one broker loses power, the others still hold the record, and a surviving in-sync replica becomes the new leader without data loss. The dangerous case is a **correlated failure** — e.g., a whole rack losing power so that *all* replicas drop their page-cache copies before flushing. To avoid this, replicas are spread across racks/availability zones with independent power. ## Key takeaway "The write succeeded" means "enough replicas have it in memory," not "it's safely on disk." Durability is bought with redundancy across machines, not with disk flushes per message.

  • What component actually moves the data from page cache to disk?
    The operating system kernel's background writeback flushes dirty page-cache pages to disk; Kafka can also force it via fsync when flush.messages/flush.ms thresholds are hit.

saying these in an interview costs you the question

  • Saying a successful write means the data is guaranteed on physical disk.
  • Confusing page cache (volatile RAM) with persistent storage.
  • Assuming a single broker's copy is durable on its own.

context

open as a page

How do acks=all and min.insync.replicas work together to provide durability, and how would you configure them with a replication factor of 3?

level: middleimportance: must knowfreq 65%

basics

~20 s

acks=all makes the producer wait until all in-sync replicas have the record. min.insync.replicas sets how many replicas must be in sync for that write to be allowed. With replication factor 3, set min.insync.replicas=2 and acks=all so you can lose one broker without losing data and without halting writes.

open as a page

Why does Kafka rely on replication for durability instead of fsync-ing every message to disk before acknowledging it?

level: middleimportance: must knowfreq 70%

basics

~20 s

Kafka acknowledges writes once enough replicas have the data in memory, not once it's flushed to disk. Replication across machines is faster and safer than a per-message fsync, which would be slow and still vulnerable to a single disk failure.

open as a page

Explain the data-loss window that exists with Kafka's default flush settings, and what conditions are required to actually lose acknowledged data.

level: seniorimportance: should knowfreq 40%

basics

~20 s

Because Kafka acknowledges writes when replicas have the data in memory (not on disk), acknowledged records can be lost only if every in-sync replica loses power at nearly the same time before any of them flushes to disk. That's the correlated power-loss window.

open as a page

What do the flush.messages and flush.ms settings control, and why are they usually left at their defaults?

level: seniorimportance: should knowfreq 45%

basics

~20 s

flush.messages and flush.ms tell a broker to fsync its log to disk after a certain number of messages or after a time interval. They're usually left at defaults because Kafka relies on replication for durability, so forcing extra fsyncs just hurts throughput.

open as a page