Why is OS page cache central to Kafka sizing, and how do you reserve page-cache and network headroom when provisioning brokers?
answer
- small heap (~5-6 GB), rest = page cache
- tail reads = zero-copy sendfile from cache
- lagging consumers → cache miss → disk reads
- NIC load = produce + (RF-1) replication + fan-out fetch
- headroom for broker-failure re-replication
basics
~20 sKafka serves recent reads from the OS page cache (RAM) instead of disk, so consumers that keep up stay memory-fast. Give the JVM a modest heap (~5-6 GB) and leave most RAM free for page cache. For network, size NICs above peak replication + produce + fetch traffic, which RF multiplies.
solid answer
~50 sKafka deliberately keeps a small JVM heap and lets the OS page cache hold hot log data; producers write to cache (flushed by the OS) and consumers reading the tail are served from cache via sendfile/zero-copy, avoiding disk seeks. So you provision RAM = small heap (typically 5-6 GB, rarely >8) + enough free memory to cache the working set (roughly the data younger than your consumers' lag). If consumers fall far behind, reads miss cache and hit disk, so capacity planning must consider read patterns, not just writes. Network is often the real bottleneck: a broker handles inbound produce traffic, outbound consumer fetches (x fan-out), and replication — both leader→follower egress and follower ingress. Replication multiplies cross-broker traffic by ~(RF-1). Size NIC bandwidth for peak total (produce + replication + fetch x consumer groups) with headroom for a broker failure when survivors absorb extra leadership and re-replication. Watch for hitting NIC saturation before CPU/disk.
go deeper
Know Kafka serves recent data from RAM (page cache) and uses a small heap.
Size RAM as small heap plus free memory for the cached working set; recognize lag causes disk reads.
Compute network load including replication and fan-out, and reserve headroom for broker-failure re-replication.
Define broker hardware profiles (RAM/NIC/disk) and capacity buffers across normal, peak, and failure modes.
## Page cache: Kafka's secret to speed The **OS page cache** is the part of RAM the kernel uses to cache file data. Kafka is built around it: - **Writes**: a produce appends to the active segment; the data lands in page cache and the OS flushes to disk asynchronously (Kafka relies on replication, not fsync-per-message, for durability). So writes are RAM-speed. - **Reads**: consumers reading the **tail** (recent data) hit page cache. Kafka uses **zero-copy** (`sendfile`) to send bytes straight from page cache to the network socket, bypassing the JVM heap entirely. That's why tail reads are extremely cheap. ### Sizing implication Give the JVM a **small heap** (typically **5-6 GB**, rarely above 8) — Kafka itself doesn't need much heap, and a big heap steals RAM from page cache and lengthens GC. Leave the **majority of RAM free** so the kernel can cache the **working set**: roughly the volume of data younger than your consumers' typical lag. Rule of thumb some use: cache enough to cover ~30 seconds to a few minutes of writes (or your largest consumer's lag window). If consumers lag far behind, their reads fall out of cache and become **random disk reads**, spiking latency and disk IOPS — so a backfill or a recovering consumer group is a capacity event. ## Network headroom A broker's network load is the sum of: - **Inbound produce**: client writes = ~T MB/s. - **Replication out (leader)**: each leader ships every write to (RF-1) followers → up to (RF-1) x (its leader traffic). - **Replication in (follower)**: it receives writes for partitions it follows. - **Outbound fetch**: every consumer group reading the topic multiplies egress by the **fan-out** (number of independent groups). So total cross-broker + client traffic can be several times the raw ingest. With RF=3 and, say, 3 consumer groups, egress alone is large. **Size the NIC (e.g. 10/25 GbE) for peak total with headroom**, because: - During a **broker failure**, survivors take over the dead broker's leadership and re-replicate its partitions — a burst of extra network + disk I/O on top of normal load. - Bursty producers and consumer catch-up create spikes. ## File descriptors and other headroom (related) While RAM and NIC dominate, also size: file descriptors (`nofile` ulimit for many segments/partitions), and disk **IOPS/throughput** for the (rarer) cache-miss reads and flushes. SSDs help recovery and lagging-consumer reads, though sequential HDD throughput can suffice for pure tail workloads. ## Putting it together Provision a broker as: small heap + abundant free RAM for page cache (sized to lag/working set), NIC bandwidth >= peak (produce + (RF-1) replication + fan-out fetch) x failure-headroom, plus disk and fd headroom. Monitor page-cache hit behavior (disk read activity), NIC utilization, and under-replicated partitions to confirm the headroom holds.
- Why does Kafka recommend a small JVM heap rather than giving it most of the RAM?Kafka offloads caching to the OS page cache and uses zero-copy for reads, so it needs little heap. A large heap wastes RAM that would otherwise cache log data and worsens GC pauses. 5-6 GB heap with the rest free for page cache is typical.
- A new analytics team starts re-reading a topic from the beginning. Why might broker latency spike?Historical reads miss the page cache (which holds only recent data), forcing random disk reads and consuming disk IOPS and NIC egress. This backfill is a capacity event that can degrade latency for tail consumers sharing the broker.
saying these in an interview costs you the question
- Giving Kafka a huge JVM heap 'for performance' (starves page cache, worsens GC).
- Sizing network only for ingest and ignoring replication and consumer fan-out.
- Assuming reads are always cheap — lagging/backfill consumers hit disk.
- Ignoring failure-mode re-replication when sizing NIC bandwidth.