A broker shows periodic p99 produce-latency spikes every 20-30 seconds that line up with LocalTimeMs jumps. How do you tell whether GC pauses or page-cache misses are the cause?
answer
- LocalTimeMs spike = leader-local stall
- GC: STW pauses, all threads, heap/CPU, regular cadence
- cache miss: disk reads, iostat await/util, lagging consumers
- Kafka leans on OS page cache, no internal cache
- overlay GC log vs disk-read timeline on the spike
basics
~20 sPeriodic LocalTimeMs spikes point at the leader's local work stalling. Check GC pause logs/JMX for stop-the-world pauses lining up with the spikes; if GC is clean, look at page-cache misses forcing disk reads (rising disk read I/O and read latency while consumers fetch old data).
solid answer
~50 sLocalTimeMs spiking means the I/O thread is stalling during leader-local processing. The two prime suspects are JVM GC and page-cache misses. For GC: enable GC logging (or read the GarbageCollector MXBeans / G1 pause metrics) and check whether stop-the-world pauses align in time and magnitude with the latency spikes — a 200 ms p99 jump every ~25 s that matches a 200 ms G1 mixed/full pause is conclusive. For page cache: Kafka relies on the OS page cache for both writes and reads; when consumers fall behind and fetch cold data, reads miss cache and hit disk, so I/O threads block on disk reads. Look for rising disk read throughput/IOPS, higher device read latency (iostat await/util), and a falling page-cache hit ratio while LocalTime climbs. GC pauses are CPU/heap-correlated and bursty/regular; page-cache misses correlate with disk read I/O and lagging/backfilling consumers. Often you fix GC by tuning heap/G1, and cache misses by adding RAM, reducing retention reads, or smoothing consumer lag.
go deeper
Know that periodic latency spikes can come from GC pauses or slow disk reads and that GC freezes the JVM.
Connect LocalTimeMs to leader-local stalls and know to check GC logs and disk read metrics.
Independently rule GC and page-cache misses in or out by overlaying GC pause and disk-read timelines on the LocalTime spike, and prescribe the right fix.
Set broker heap/page-cache and GC policy, design cache-aware capacity and tiered-storage strategy, and teach the page-cache-centric performance model.
## Where the symptom points **LocalTimeMs** is the time an I/O (request-handler) thread spends doing *leader-local* work for a request — appending to the log, validating/decompressing, or reading from the log. When LocalTime spikes, the thread itself stalled mid-work. Two dominant causes: **JVM garbage collection** (a process-wide pause that freezes the I/O thread) and **page-cache misses** (the thread blocks on a synchronous disk read). ## Background: how Kafka uses the page cache Kafka does **not** maintain its own data cache. It writes records to the OS page cache (the kernel's in-RAM cache of file pages) and lets the OS flush them to disk; reads are served from page cache when the data is hot. This is why a healthy Kafka broker shows most reads served from RAM with little disk read I/O. **Page-cache miss** = the page a consumer wants isn't in RAM, so the read becomes an actual disk read, which is orders of magnitude slower and blocks the I/O thread. ## Distinguishing GC from page-cache misses **Signature of GC pauses:** - Enable GC logging (`-Xlog:gc*` on modern JVMs) or read `java.lang:type=GarbageCollector` MXBeans / G1 metrics. - Look for **stop-the-world (STW)** pause durations that match the latency spike magnitude and timing. A regular ~25 s cadence often matches G1 mixed collections or a heap filling at a steady allocation rate. - GC pauses freeze *all* threads, so you'd see correlated jumps across many request types simultaneously, and CPU/heap-after-GC trends moving. - Fixes: right-size heap (Kafka brokers usually want a *modest* heap, e.g. 6–8 GB, leaving most RAM for page cache), tune G1 (`-XX:MaxGCPauseMillis`), avoid huge heaps that cause long mixed/full GCs. **Signature of page-cache misses:** - LocalTime spikes correlate with **disk read** activity: rising read IOPS/throughput, high device `%util` and `await` in `iostat`, growing read latency. - Correlates with **consumer behavior**: a consumer group falling behind and backfilling old segments, a new consumer reading from the beginning, or a batch/analytics job sweeping historical data — all force cold reads. - Falling page-cache hit ratio; the working set exceeds RAM. - Fixes: add RAM (more page cache), reduce the amount of cold-read traffic (tiered storage, smoothing lag), avoid retention/segment configs that thrash cache, isolate batch readers. ## Key discriminator GC is a **CPU/heap-and-time-correlated, all-threads-frozen** event with no necessary disk activity. Page-cache miss is a **disk-read-correlated, consumer-driven** event. Pull both timelines and overlay them on the LocalTime spike — only one will align. ## Edge cases - Both can coexist; a long GC can also evict/cool caches indirectly. Rule each in or out independently. - A regular cadence strongly hints GC (allocation-rate-driven), but a periodic batch consumer job can also produce a regular cadence of cache misses — check what's running on schedule. - fsync stalls (background flush, `log.flush.interval`) are a third disk cause that also raises LocalTime; iostat write latency separates it from read-driven cache misses. - Don't forget non-GC JVM pauses (e.g. safepoint, page faults from heap swapping) — ensure swappiness is low so the JVM heap is never swapped.
- Why do Kafka brokers typically run a relatively small JVM heap?Kafka relies on the OS page cache for read/write performance, so leaving most RAM to the kernel page cache (rather than a large heap) maximizes cache hits and minimizes disk reads. A small heap also keeps GC pauses short, reducing tail-latency stalls.
- How would you distinguish a page-cache read miss from an fsync/flush write stall, since both raise LocalTimeMs?Look at iostat: read-driven cache misses show high read IOPS/await, whereas flush stalls show high write IOPS/await and correlate with flush intervals or background fsync. Page-cache misses also correlate with lagging/backfilling consumers; flush stalls correlate with produce write volume.
saying these in an interview costs you the question
- Assuming any LocalTime spike is automatically GC without checking disk-read metrics
- Recommending a huge JVM heap for Kafka (it starves the page cache and lengthens GC pauses)
- Believing Kafka keeps its own in-process record cache (it relies on the OS page cache)
- Ignoring that a periodic cadence can come from a scheduled batch consumer, not only from GC