skip to content

Broker Memory, Threads and Disk Layout

How to split memory between the JVM heap and the page cache, and how logs are laid out across data dirs with recovery checkpoints. This is the practical sizing question for anyone claiming production Kafka experience.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

How should you size a Kafka broker's JVM heap versus the OS page cache, and why does Kafka deliberately keep its heap small?

level: middleimportance: must knowfreq 72%

answer

  1. small heap (~6 GB), big page cache
  2. Xms == Xmx, no resize
  3. zero-copy sendfile bypasses heap
  4. OS flushes dirty pages, not Kafka
  5. RAM is zero-sum: heap steals from cache

basics

~20 s

Give the Kafka broker a small JVM heap (commonly 6 GB) and leave most of the machine's RAM free for the OS page cache, which caches log file data. Kafka relies on the page cache, not heap, for read/write throughput.

solid answer

~40 s

Kafka writes messages to append-only log segment files and lets the Linux page cache hold the hot data. Producers write into the page cache (flushed to disk by the OS later), and consumers reading recent data are served from the same cache, often without touching disk. So the broker itself needs only a modest heap (the classic guidance is around 6 GB, set with -Xms/-Xmx equal) for connections, request buffers, replication state, and metadata; everything else should stay free RAM for the page cache. A large heap is counterproductive: it steals RAM from the cache and causes long GC pauses. With zero-copy (sendfile) Kafka transfers bytes straight from the page cache to the network socket, bypassing the heap entirely. You tune for cache hit rate, not heap size.

go deeper

for a junior

Know that Kafka brokers use a small heap and rely on the OS to cache file data in RAM.

for a middle

Explain the heap-vs-page-cache split, the ~6 GB guidance with Xms==Xmx, and that the OS flushes data asynchronously.

for a senior

Reason about GC pause impact, cache hit rate, zero-copy/sendfile, and when many partitions justify raising the heap.

for a principal

Frame broker memory sizing as a capacity-planning trade-off across cache hit rate, GC behavior, replication state, and host co-tenancy; set org-wide defaults and monitoring.

## The two memory pools on a Kafka broker Every Kafka broker process has access to the machine's RAM, which is effectively split into two pools that matter here: 1. **The JVM heap** — memory the Java process manages directly, garbage-collected, holding Java objects: network connection state, in-flight request/response buffers, the replica fetcher state, in-memory indexes metadata, and producer/consumer session bookkeeping. 2. **The OS page cache** — RAM the operating system (Linux) uses to cache file contents. When any process reads or writes a file, the OS keeps recently touched file pages in this cache so subsequent access avoids the physical disk. ## Why Kafka leans on the page cache Kafka stores each partition as a sequence of **append-only log segment files** on disk. It does NOT maintain its own in-process cache of message data. Instead: - **On write:** a producer's records are appended to the active segment file. The bytes land in the page cache; the OS flushes them to the physical disk asynchronously (controlled by the kernel's dirty-page writeback, and optionally by Kafka's `log.flush.interval.messages`/`log.flush.interval.ms`, which are usually left at their defaults so the OS decides). - **On read:** a consumer reading recent data hits the same pages still in cache — no disk I/O. Kafka further uses **zero-copy** via the `sendfile` system call (`TransferableChannel`/`FileChannel.transferTo`), shuttling bytes directly from the page cache to the network socket buffer without ever copying them into the JVM heap. Because the hot path (recent messages) is served from RAM the OS already manages, the broker gains almost nothing from a large heap — and pays for it. ## Why a small heap is the right call - **RAM is zero-sum.** Memory you give the heap is memory the OS cannot use for the page cache. A bloated heap shrinks the cache, lowers the cache hit rate, and forces more disk reads. - **GC pauses.** Larger heaps mean longer garbage-collection pauses (even with G1, the default modern collector). A multi-second stop-the-world pause can cause the broker to miss ZooKeeper/Raft session heartbeats and get fenced, or cause replica fetch lag. - **Empirically, brokers are happy with little heap.** The common production guidance is a heap of roughly **6 GB** (set via `KAFKA_HEAP_OPTS="-Xmx6g -Xms6g"`, with Xms == Xmx to avoid heap resizing). Even very large brokers rarely need more than ~8-12 GB. The rest of a 64 GB machine is left for page cache. ## Edge cases and nuance - **Many partitions / many connections** push heap usage up (each partition and connection carries some bookkeeping), so heavily loaded brokers may need more than 6 GB — but you raise it deliberately, watching GC logs, not by default. - **Off-heap:** Kafka uses some off-heap/direct memory for network buffers; this is separate from the heap and the page cache and is usually small. - **Cold reads:** consumers reading old data (backfill, lagging consumers) miss the cache and hit disk; this is expected and is why disk throughput still matters. - **Don't run other memory-hungry processes** on the broker host — they evict Kafka's page cache. ## Bottom line Tune for page-cache hit rate. Keep the heap small and fixed; give the OS the RAM. The broker's performance model is 'the OS caches my files for me,' not 'I cache everything in heap.'

  • Why set -Xms equal to -Xmx on a broker?
    To pre-allocate the full heap at startup and prevent the JVM from growing/shrinking the heap at runtime, which causes GC variability and gives more predictable memory layout — leaving the rest of RAM stably available to the page cache.
  • What is zero-copy and how does it relate to the page cache?
    Zero-copy uses the sendfile syscall to transfer bytes directly from the page cache (kernel space) to the network socket without copying them into the JVM heap (user space). It only works because data is served from the page cache, reinforcing why heap caching is unnecessary.

saying these in an interview costs you the question

  • Saying 'give Kafka as much heap as possible for caching' — Kafka caches in the OS page cache, not the heap.
  • Claiming Kafka maintains its own large in-process message cache.
  • Thinking a bigger heap improves read throughput — it shrinks the page cache and worsens it.

context

open as a page

What are the checkpoint files in a Kafka log directory (recovery-point-offset-checkpoint and replication-offset-checkpoint), and what does each track?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Each log directory holds small checkpoint files. recovery-point-offset-checkpoint records, per partition, the offset up to which data has been flushed to disk (so recovery can skip it). replication-offset-checkpoint records the high watermark, the offset up to which messages are replicated/committed.

open as a page

How does a Kafka broker know on startup whether it was shut down cleanly, and why does that distinction matter?

level: juniorimportance: should knowfreq 35%

basics

~20 s

On a clean shutdown the broker flushes all data, updates the checkpoint files, and writes a clean-shutdown marker file in each log dir. On startup, if that marker is present it skips log recovery; if it is missing (a crash), it runs recovery.

open as a page

What does configuring multiple log.dirs (JBOD) give you on a Kafka broker, and how are partitions placed across the directories?

level: middleimportance: should knowfreq 55%

basics

~20 s

Setting multiple paths in log.dirs (JBOD = Just a Bunch Of Disks) lets one broker spread partition data across several independent disks for more total capacity and I/O. Kafka assigns each new partition to the directory with the fewest partitions.

open as a page

What does num.recovery.threads.per.data.dir control, and when does increasing it help?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It sets how many threads, per log directory, Kafka uses to load and recover partition logs at startup and to flush them at shutdown. Raising it speeds up startup recovery on brokers with many partitions, especially after an unclean shutdown.

open as a page