skip to content

Why does Kafka rely on the OS page cache instead of flushing every write to disk, and what is the role of log.flush.interval.messages / log.flush.interval.ms?

level: middleimportance: must knowfreq 65%

answer

  1. Write to RAM, OS flushes later
  2. Durability = replication, not fsync
  3. flush.interval.* default ~off
  4. Reads served from cache → zero-copy
  5. Tune vm.dirty_ratio not Kafka flush

basics

~20 s

Kafka writes to the OS page cache and lets the OS flush to disk in the background, which is fast. Durability comes from replication, not fsync. log.flush.interval.* would force periodic fsyncs, but Kafka leaves them effectively off by default.

solid answer

~50 s

Kafka appends records to log segment files via the page cache — a region of RAM the OS uses to buffer file data. Writes return as soon as they hit cache; the kernel flushes dirty pages to disk asynchronously. This gives near-memory write speed and means recently produced and consumed data is served from RAM, so reads usually never touch disk. Kafka deliberately does not fsync on each write; durability is provided by replicating to other brokers (acks=all, min.insync.replicas) rather than by forcing data to platters on one node. log.flush.interval.messages and log.flush.interval.ms can force Kafka to fsync after N messages or every T ms, but their defaults are effectively infinite/disabled — enabling them trades throughput for marginal single-node durability and is almost always the wrong tradeoff. The flush cadence is therefore left to the OS (tuned via vm.dirty_ratio / vm.dirty_background_ratio).

go deeper

for a junior

Know writes go to RAM (page cache) and the OS flushes later; Kafka is fast because it doesn't wait for disk.

for a middle

Articulate that durability = replication, and that log.flush.interval.* defaults are off and shouldn't be turned on lightly.

for a senior

Connect page cache to zero-copy reads, no in-heap cache, and OS vm.dirty_* tuning instead of Kafka flush configs.

for a principal

Reason about the full durability tradeoff: acks/min.insync.replicas vs fsync, correlated power-loss failure domains, and RAM sizing for cache hit rate.

## What the page cache is The **page cache** is RAM the operating-system kernel uses to hold copies of file data. When a process writes to a file, the bytes first land in dirty page-cache pages in RAM; the write call returns immediately, and the kernel later **flushes** (writes back) those dirty pages to the physical disk. When a process reads, the kernel serves from cache if the data is already there. Kafka's storage engine is built entirely around this: appends to a partition's active log segment go into the page cache, and fetches read from it. ## Why Kafka leans on it 1. **Write speed.** Returning after a memory copy is orders of magnitude faster than waiting for a disk `fsync()`. Sequential appends + page cache let a single broker sustain very high write throughput. 2. **Read speed without an app cache.** Because consumers typically read data shortly after it's produced, that data is still in page cache, so most fetches are served from RAM. Kafka therefore keeps no large JVM-heap data cache — it delegates caching to the OS, which avoids GC pressure and double-buffering. (This synergizes with zero-copy `sendfile`, which streams cached file bytes straight to the socket.) 3. **Survives broker process restart.** Page cache lives in the kernel, not the JVM. If the Kafka process restarts (but the OS does not), the cache is still warm. ## Durability model: replication, not fsync The key mental shift: **Kafka does not equate 'written' with 'fsynced to disk'.** A produce with `acks=all` is acknowledged when the record has been written (to page cache) on the leader and replicated to all in-sync replicas (`min.insync.replicas`). Durability comes from having the data in the page cache / log of *multiple independent brokers*, so a single machine losing power (and its un-flushed cache) does not lose acknowledged data. This is why per-write fsync is unnecessary and avoided. ## `log.flush.interval.messages` and `log.flush.interval.ms` These two settings let Kafka itself force an `fsync` of a log to disk: - `log.flush.interval.messages` — fsync after this many messages have been appended to a log. - `log.flush.interval.ms` — fsync at most every this many milliseconds. **Defaults are effectively disabled** (Long.MAX_VALUE / unset), meaning Kafka does not proactively fsync and leaves flushing to the OS. Enabling aggressive values forces frequent fsyncs, which serializes writes against disk latency and tanks throughput, in exchange for only single-node durability that replication already provides better. The standard guidance is to leave them off and rely on replication. ## Tuning the OS flush instead Because the OS owns flush timing, you tune *it*: `vm.dirty_background_ratio` / `vm.dirty_ratio` (or the `_bytes` variants) control how much dirty data accumulates before the kernel starts/forces writeback. Too high → a huge backlog flushes at once, causing latency spikes; too low → frequent small writebacks. Operators tune these to smooth flush behavior rather than touching Kafka's flush configs. ## Edge cases - A simultaneous power loss across enough replicas can still lose un-flushed data — that's the accepted tradeoff, mitigated by spreading replicas across racks/AZs and using `acks=all`. - Page cache is shared with everything else on the box; co-locating other memory-hungry processes evicts hot Kafka pages and forces cold disk reads. - A broker with too little free RAM for page cache will see read I/O spike because fetches miss cache and hit disk.

  • If you don't fsync per write, how is acknowledged data not lost on a crash?
    Because acks=all replicates it to min.insync.replicas other brokers' logs/page caches; durability comes from multiple independent copies, not from forcing one node's data to platters.
  • What OS knobs control flush cadence when Kafka's flush configs are off?
    vm.dirty_background_ratio and vm.dirty_ratio (or the *_bytes variants) govern when the kernel starts and forces writeback of dirty pages.

saying these in an interview costs you the question

  • Saying Kafka fsyncs every message for durability — it does not by default.
  • Claiming Kafka keeps a large in-heap data cache — it relies on the OS page cache.
  • Thinking enabling log.flush.interval.messages improves throughput — it hurts it.
  • Equating 'acked' with 'fsynced to disk on the leader'.

context