An app is slow and you suspect GC. Using jstat alone, how do you confirm GC churn or a heap leak, and what are jstat's limits here?
answer
- Poll and read deltas — counters are cumulative
- Leak signal = Old stays high AFTER full GC, not before
- GCT ÷ uptime = GC overhead %; jstat -t adds a timestamp
- Old recovers = pressure/undersize; Old ratchets = leak
- jstat shows aggregates only → heap dump + MAT for the culprit
basics
~20 sRun jstat -gcutil <pid> 1000 and watch over time. Frequent full GCs (FGC climbing) with growing FGCT and an Old region that stays high after each full GC point to GC churn or a leak. To find the actual culprit objects you still need a heap dump, because jstat only shows aggregate numbers, not what's filling the heap.
solid answer
~50 sStart with `jstat -gcutil <pid> 1000` and observe across many samples. Key signals: **FGC** (full-GC count) incrementing rapidly, **FGCT** growing so that GC eats a large share of wall-clock time, and crucially **O** (old gen) staying near full *after* each full GC instead of dropping — that 'collected but not reclaimed' shape means live objects keep accumulating (a leak) rather than transient pressure. Compare GCT to process uptime to quantify GC overhead. If instead Old drops nicely after full GCs but full GCs are merely frequent, you may just have an undersized heap or high allocation rate — fix with `-Xmx`/allocation reduction, not a leak hunt. jstat's limit: it shows *aggregate* counters only — it can't tell you *which* objects or allocation sites are responsible. To go further you need a heap dump (jmap + Eclipse MAT for dominator trees), GC logs, or a profiler / JFR. jstat confirms the *symptom*; those tools find the *cause*.
go deeper
Can run jstat -gcutil and notice that lots of full GCs and a full Old generation look unhealthy.
Watches deltas over time and knows frequent FGC with growing FGCT indicates GC pressure; aware a heap dump is the next step.
Distinguishes leak (Old won't recover) from undersized-heap pressure (Old recovers), quantifies GC overhead, and knows jstat's aggregate-only limit forces escalation to heap dumps/GC logs/JFR.
Frames jstat as ad-hoc triage within a broader observability strategy, decides when GC logs/JFR/metrics pipelines are the right production tooling, and separates JVM-heap from native-memory diagnoses.
## Diagnosing GC trouble with jstat — and where it stops When an app is sluggish and you suspect garbage collection, jstat is the cheapest confirmation step. But it's important to know exactly what it *can* prove and what it *can't*. ### Step 1 — sample over time, not once The GC counters (YGC, FGC, YGCT, FGCT, GCT) are **cumulative since JVM start**, so a single snapshot is nearly useless. Poll: ``` jstat -gcutil <pid> 1000 ``` and watch *deltas* across many lines. A 1-second interval is a good default for human watching. ### Step 2 — read the three diagnostic signals 1. **Full-GC frequency (FGC delta).** If FGC increments several times per second, the collector is repeatedly doing the expensive, often stop-the-world major collection. Frequent full GCs almost always correlate with latency spikes. 2. **GC overhead (GCT vs elapsed).** Take GCT (total GC seconds) and divide by the process's elapsed wall-clock time (from `jstat -t` which prepends a Timestamp column, or the OS). A healthy server spends a low single-digit percentage in GC; tens of percent means GC is your bottleneck. This is essentially the same condition the JVM itself watches for `GC overhead limit exceeded` (≈98% time in GC reclaiming <2%). 3. **The leak shape — Old not falling.** This is the decisive one. Watch **O** immediately *after* each full GC (FGC ticks up). - If O **drops back down** after a full GC and the app stabilizes → not a leak; it's transient pressure / undersized heap / a burst of allocation. - If O **stays high / ratchets upward** after every full GC, climbing toward 100% over minutes → live (reachable) data keeps growing. The collector ran but had nothing it was *allowed* to reclaim. That is the signature of a **heap leak**, heading for `OutOfMemoryError: Java heap space`. The same logic applies to **M** (metaspace) climbing toward its max → a likely **classloader leak** ending in `OutOfMemoryError: Metaspace`. ### Step 3 — distinguish leak vs. mere pressure The jstat output lets you separate two commonly-confused situations: - **Undersized heap / high allocation rate:** lots of GCs, but Old *recovers* after each full GC. Remediation: bigger `-Xmx`, reduce allocation, tune the collector. - **Leak:** Old never recovers. Remediation: find what's retaining objects — and a bigger heap only delays the OOM. ### Where jstat hits its ceiling jstat reports **aggregate counters only**. It will tell you *that* the old gen is full and *that* full GCs aren't reclaiming it — but it cannot tell you **which classes, which objects, or which allocation sites** are responsible. It has no per-object, per-thread, or per-stack-trace detail. So once jstat has confirmed the symptom, you escalate to tools that show the *cause*: - **Heap dump** via `jmap -dump:live,format=b,file=heap.hprof <pid>` (or `-XX:+HeapDumpOnOutOfMemoryError` to capture one automatically at the crash), then analyze with **Eclipse MAT** or **VisualVM** to find **dominator trees** and **retained sizes** — i.e., *what* is holding memory and *who* keeps it reachable. - **GC logs** (`-Xlog:gc*` on modern JVMs) for a complete, timestamped history with pause durations and causes — richer and more permanent than jstat sampling. - **Java Flight Recorder (JFR)** / async-profiler for allocation profiling — *where in the code* the allocations originate. Other practical limits of jstat: it samples the perfdata file so it can miss very short-lived spikes between samples; it's a manual, interactive tool not a metrics pipeline (for fleets you want GC logs or a metrics agent feeding a dashboard); and it shows the JVM's view, not OS-level RSS, so native-memory issues outside the Java heap won't appear. ### Summary jstat is the **stethoscope**: cheaply confirm GC churn (FGC/FGCT growth, high GC overhead) and the **leak shape** (Old refusing to fall after full GCs). It's the wrong tool to identify the offending objects — for that you take a heap dump and open it in MAT, or capture GC logs / a JFR recording. Use jstat to decide *whether* to dig, then dig with the heavier tools.
- After jstat confirms Old never recovers, what's your next concrete step?Capture a heap dump — `jmap -dump:live,format=b,file=heap.hprof <pid>` or rely on `-XX:+HeapDumpOnOutOfMemoryError` — and open it in Eclipse MAT (or VisualVM). Use the dominator tree and 'retained size' to find which object subtree holds most of the heap and which GC root keeps it reachable; that points at the leak (e.g. an unbounded static cache or a never-deregistered listener).
- How do you turn the cumulative GCT into an overhead percentage with jstat?Run `jstat -t -gcutil <pid> <interval>`: the `-t` flag prepends a Timestamp (seconds since JVM start). GC overhead ≈ GCT / Timestamp × 100. Watching it over an interval gives the recent-rate overhead; a sustained high value (e.g. >10–20%) means GC is the bottleneck.
saying these in an interview costs you the question
- Calling it a leak just because full GCs are frequent — you must check whether Old recovers afterward
- Reading a single jstat snapshot instead of watching the trend
- Expecting jstat to name the leaking class or allocation site — it only reports aggregates
- Ignoring metaspace growth, which is a classloader-leak path distinct from heap leaks
- Confusing JVM heap utilization with OS process memory (RSS) — native leaks won't show in jstat