skip to content

A broker shows periodic p99 produce-latency spikes every 20-30 seconds that line up with LocalTimeMs jumps. How do you tell whether GC pauses or page-cache misses are the cause?

level: seniorimportance: should knowfreq 55%

answer

  1. LocalTimeMs spike = leader-local stall
  2. GC: STW pauses, all threads, heap/CPU, regular cadence
  3. cache miss: disk reads, iostat await/util, lagging consumers
  4. Kafka leans on OS page cache, no internal cache
  5. overlay GC log vs disk-read timeline on the spike

basics

~20 s

Periodic LocalTimeMs spikes point at the leader's local work stalling. Check GC pause logs/JMX for stop-the-world pauses lining up with the spikes; if GC is clean, look at page-cache misses forcing disk reads (rising disk read I/O and read latency while consumers fetch old data).

solid answer

~50 s

LocalTimeMs spiking means the I/O thread is stalling during leader-local processing. The two prime suspects are JVM GC and page-cache misses. For GC: enable GC logging (or read the GarbageCollector MXBeans / G1 pause metrics) and check whether stop-the-world pauses align in time and magnitude with the latency spikes — a 200 ms p99 jump every ~25 s that matches a 200 ms G1 mixed/full pause is conclusive. For page cache: Kafka relies on the OS page cache for both writes and reads; when consumers fall behind and fetch cold data, reads miss cache and hit disk, so I/O threads block on disk reads. Look for rising disk read throughput/IOPS, higher device read latency (iostat await/util), and a falling page-cache hit ratio while LocalTime climbs. GC pauses are CPU/heap-correlated and bursty/regular; page-cache misses correlate with disk read I/O and lagging/backfilling consumers. Often you fix GC by tuning heap/G1, and cache misses by adding RAM, reducing retention reads, or smoothing consumer lag.

go deeper

for a junior

Know that periodic latency spikes can come from GC pauses or slow disk reads and that GC freezes the JVM.

for a middle

Connect LocalTimeMs to leader-local stalls and know to check GC logs and disk read metrics.

for a senior

Independently rule GC and page-cache misses in or out by overlaying GC pause and disk-read timelines on the LocalTime spike, and prescribe the right fix.

for a principal

Set broker heap/page-cache and GC policy, design cache-aware capacity and tiered-storage strategy, and teach the page-cache-centric performance model.

## Where the symptom points **LocalTimeMs** is the time an I/O (request-handler) thread spends doing *leader-local* work for a request — appending to the log, validating/decompressing, or reading from the log. When LocalTime spikes, the thread itself stalled mid-work. Two dominant causes: **JVM garbage collection** (a process-wide pause that freezes the I/O thread) and **page-cache misses** (the thread blocks on a synchronous disk read). ## Background: how Kafka uses the page cache Kafka does **not** maintain its own data cache. It writes records to the OS page cache (the kernel's in-RAM cache of file pages) and lets the OS flush them to disk; reads are served from page cache when the data is hot. This is why a healthy Kafka broker shows most reads served from RAM with little disk read I/O. **Page-cache miss** = the page a consumer wants isn't in RAM, so the read becomes an actual disk read, which is orders of magnitude slower and blocks the I/O thread. ## Distinguishing GC from page-cache misses **Signature of GC pauses:** - Enable GC logging (`-Xlog:gc*` on modern JVMs) or read `java.lang:type=GarbageCollector` MXBeans / G1 metrics. - Look for **stop-the-world (STW)** pause durations that match the latency spike magnitude and timing. A regular ~25 s cadence often matches G1 mixed collections or a heap filling at a steady allocation rate. - GC pauses freeze *all* threads, so you'd see correlated jumps across many request types simultaneously, and CPU/heap-after-GC trends moving. - Fixes: right-size heap (Kafka brokers usually want a *modest* heap, e.g. 6–8 GB, leaving most RAM for page cache), tune G1 (`-XX:MaxGCPauseMillis`), avoid huge heaps that cause long mixed/full GCs. **Signature of page-cache misses:** - LocalTime spikes correlate with **disk read** activity: rising read IOPS/throughput, high device `%util` and `await` in `iostat`, growing read latency. - Correlates with **consumer behavior**: a consumer group falling behind and backfilling old segments, a new consumer reading from the beginning, or a batch/analytics job sweeping historical data — all force cold reads. - Falling page-cache hit ratio; the working set exceeds RAM. - Fixes: add RAM (more page cache), reduce the amount of cold-read traffic (tiered storage, smoothing lag), avoid retention/segment configs that thrash cache, isolate batch readers. ## Key discriminator GC is a **CPU/heap-and-time-correlated, all-threads-frozen** event with no necessary disk activity. Page-cache miss is a **disk-read-correlated, consumer-driven** event. Pull both timelines and overlay them on the LocalTime spike — only one will align. ## Edge cases - Both can coexist; a long GC can also evict/cool caches indirectly. Rule each in or out independently. - A regular cadence strongly hints GC (allocation-rate-driven), but a periodic batch consumer job can also produce a regular cadence of cache misses — check what's running on schedule. - fsync stalls (background flush, `log.flush.interval`) are a third disk cause that also raises LocalTime; iostat write latency separates it from read-driven cache misses. - Don't forget non-GC JVM pauses (e.g. safepoint, page faults from heap swapping) — ensure swappiness is low so the JVM heap is never swapped.

  • Why do Kafka brokers typically run a relatively small JVM heap?
    Kafka relies on the OS page cache for read/write performance, so leaving most RAM to the kernel page cache (rather than a large heap) maximizes cache hits and minimizes disk reads. A small heap also keeps GC pauses short, reducing tail-latency stalls.
  • How would you distinguish a page-cache read miss from an fsync/flush write stall, since both raise LocalTimeMs?
    Look at iostat: read-driven cache misses show high read IOPS/await, whereas flush stalls show high write IOPS/await and correlate with flush intervals or background fsync. Page-cache misses also correlate with lagging/backfilling consumers; flush stalls correlate with produce write volume.

saying these in an interview costs you the question

  • Assuming any LocalTime spike is automatically GC without checking disk-read metrics
  • Recommending a huge JVM heap for Kafka (it starves the page cache and lengthens GC pauses)
  • Believing Kafka keeps its own in-process record cache (it relies on the OS page cache)
  • Ignoring that a periodic cadence can come from a scheduled batch consumer, not only from GC

context