skip to content

Page Cache and Zero-Copy I/O

Why Kafka keeps no records in its own heap and leans on the OS page cache plus sendfile for reads. A favourite 'why is Kafka fast' question, and the reason broker memory tuning looks nothing like a normal JVM service.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

Why does Kafka rely on the operating system page cache instead of maintaining its own in-process (JVM heap) record cache?

level: juniorimportance: must knowfreq 70%

answer

  1. page cache = kernel RAM for file pages
  2. small heap (~6GB), rest to OS cache
  3. no double buffering, no GC pauses
  4. cache survives broker restart
  5. enables sendfile/zero-copy

basics

~20 s

Kafka writes data to files and lets the OS keep recently used file pages in RAM (the page cache). It avoids a JVM heap cache to dodge GC pressure, double-buffering, and to reuse the OS cache that survives broker restarts.

solid answer

~50 s

Kafka deliberately keeps almost no record data on the JVM heap. When a broker writes to a partition log it writes to a regular file; the OS caches recently touched pages in the page cache (free RAM). Reads for recent data are served straight from that cache without Kafka allocating Java objects. The reasons: (1) a heap cache would double-store data already in the page cache, wasting RAM; (2) large heaps cause long GC pauses, so Kafka runs a small heap (often 5-6 GB) and gives the rest of RAM to the page cache; (3) the page cache is managed in compact, off-heap memory by the kernel and survives a broker process restart (warm cache); (4) it lets Kafka use sendfile/zero-copy. Net effect: Kafka's throughput tracks the OS's ability to cache hot data and do sequential I/O.

go deeper

for a junior

Know that Kafka relies on the OS page cache, not a Java-heap cache, and roughly why (GC, double caching).

for a middle

Explain the small-heap/large-page-cache tuning and that the cache survives restart.

for a senior

Connect the page cache decision to zero-copy, GC tuning, and co-location pitfalls; reason about cold vs warm reads.

for a principal

Teach the full memory model trade-off and capacity-plan RAM around hot-data working set and consumer lag patterns.

## The core idea Kafka stores each partition as an append-only log on disk: a sequence of *segment* files. When a producer sends records, the broker appends serialized bytes to the active segment file. When a consumer reads, the broker reads bytes back from those files. The key architectural decision is that Kafka does **not** keep a cache of records on the **JVM heap** — it leans entirely on the **operating system page cache**. ## What is the page cache? The **page cache** is RAM that the OS kernel uses to hold copies of recently accessed file data (in fixed-size **pages**, typically 4 KB). When any process reads or writes a file, the data flows through this cache. If the data is already cached, a read is satisfied from RAM with no disk access. This cache is **transparent** — applications just do normal file I/O and benefit automatically. It is **off-heap** (kernel memory, not counted against the JVM), and it **persists across a process restart** because it belongs to the kernel, not the Kafka process. ## Why not a heap cache? - **Avoid double caching.** If Kafka cached records on the heap, the same bytes would live twice in RAM: once in the JVM, once in the page cache. That halves effective cache capacity. - **Avoid GC pressure.** The JVM garbage collector must scan and manage heap objects. A multi-gigabyte object cache causes long GC pauses and unpredictable latency. Kafka instead runs a **small heap** (commonly 5-6 GB regardless of machine size) and lets the kernel use the remaining RAM as page cache. - **Compact representation.** Object headers and references in the JVM can inflate memory use ~2x versus the raw compact bytes the kernel stores. - **Warm cache after restart.** Because the page cache is the kernel's, restarting the broker process does not lose it; the broker comes back with a warm cache. A heap cache would be cold after every restart. - **Enables zero-copy.** Serving reads from files lets Kafka use the `sendfile` syscall to ship bytes from page cache straight to the network socket, bypassing user space entirely. ## The consequence Kafka's read/write performance is effectively the OS's performance: sequential appends are cheap, and reads of recent data (the common consumer pattern) hit the page cache and never touch disk. This is why Kafka docs say it is built to take advantage of the **'free' memory the OS already manages**. ## Edge cases - If consumers fall far behind (read old, cold data), reads miss the page cache and hit disk — historically random-ish reads, mitigated by sequential segment layout. - Co-locating other memory-hungry apps on the broker steals page cache and tanks Kafka throughput. - A huge JVM heap is an anti-pattern: it starves the page cache. The guidance is small heap, large page cache.

  • Roughly how big should the Kafka broker JVM heap be, and why not larger?
    Typically ~5-6 GB even on large machines. A bigger heap steals RAM the OS would use for page cache and lengthens GC pauses; Kafka stores almost no data on-heap so it does not need a large heap.
  • What happens to performance if you co-locate a memory-hungry app on the broker?
    It competes for RAM with the page cache, shrinking the cache, causing more disk reads and page evictions, and degrading Kafka throughput and latency.

saying these in an interview costs you the question

  • Claiming Kafka keeps a large in-JVM cache of recent records.
  • Recommending a very large broker heap 'for caching'.
  • Thinking the page cache is lost when the broker process restarts.
  • Confusing page cache (kernel) with the producer/consumer client buffers.

context

open as a page

What is zero-copy in Kafka, which syscall provides it, and what copies/context-switches does it eliminate on the consumer read path?

level: seniorimportance: must knowfreq 65%

basics

~20 s

Zero-copy lets Kafka send file data from the page cache straight to the network socket using the sendfile syscall (via Java's FileChannel.transferTo). It skips copying the bytes into the application (user space) and back, saving CPU and memory bandwidth.

open as a page

Explain how Kafka's append-only log and sequential disk writes make it fast even on spinning disks.

level: middleimportance: should knowfreq 55%

basics

~20 s

Kafka only appends to the end of partition log files instead of writing in random places. Sequential writes are far faster than random writes (no disk seeks), so even HDDs can sustain high throughput, and writes go through the page cache first.

open as a page

How does Kafka use memory-mapped files (mmap) for its offset and time indexes, and how do consumers find a record by offset?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Each log segment has sparse index files (.index for offsets, .timeindex for timestamps) that Kafka memory-maps (mmap) so lookups happen in RAM. To find an offset, Kafka binary-searches the sparse index to get a nearby file position, then scans forward in the .log file.

open as a page

What do flush.messages and flush.ms control, and why does Kafka discourage relying on broker fsync for durability?

level: principalimportance: should knowfreq 45%

basics

~20 s

flush.messages and flush.ms force the broker to fsync log data to disk after N messages or T milliseconds. By default both are effectively unset, so Kafka lets the OS flush dirty pages lazily and relies on replication, not local fsync, for durability — forcing frequent fsync hurts throughput.

open as a page