What does 'OutOfMemoryError: GC overhead limit exceeded' mean, and how does it differ from a plain 'Java heap space' error?
answer
- ~98% time in GC, <2% heap reclaimed
- Fail-fast instead of GC thrashing / death spiral
- Same root cause as 'Java heap space', caught earlier
- Throughput/Parallel collector's tripwire
- -XX:-UseGCOverheadLimit only delays a harder crash
basics
~20 sIt means the garbage collector is working almost constantly (about 98% of the time) but barely freeing anything (less than 2% of the heap). The JVM gives up early instead of thrashing. It's an early-warning version of a heap problem: the heap is nearly full and the program is spending all its time collecting instead of doing real work.
solid answer
~50 s'GC overhead limit exceeded' is a heap-pressure OOME thrown by the throughput (Parallel) collector when it detects the JVM is spending **~98% of total time in GC while reclaiming less than ~2% of the heap** across recent collections. Rather than let the process **thrash** — endlessly GC'ing, freeing crumbs, immediately filling up again, making no real progress — the JVM fails fast with this dedicated message. It's essentially the same underlying condition as 'Java heap space' (the heap is effectively exhausted by reachable objects), but caught **earlier**, at the point where the app is alive yet useless because it's all GC and almost no work. Causes mirror heap-space: a memory leak, an under-sized heap, or sustained allocation right at the ceiling. Diagnosis and fix are the same — take a heap dump, decide leak vs. under-size, fix reachability or raise `-Xmx`. You *can* disable just this check with `-XX:-UseGCOverheadLimit`, but that only converts it into a later, harder 'Java heap space' crash, so it's rarely the right move.
go deeper
Knows it means the program is spending almost all its time garbage collecting and freeing almost nothing, and that it's a memory problem.
States the ~98%/<2% thresholds, recognizes it as a fail-fast against thrashing, and knows the fix path mirrors a heap-space OOME (dump, then leak vs. resize).
Frames it as the same heap-exhaustion root cause caught earlier by the throughput collector, reads GC logs to confirm the spiral, and explains why -XX:-UseGCOverheadLimit only postpones a harder failure.
Connects it to capacity/SLO management: GC-overhead alerts as early signals, collector choice and heap sizing strategy, and why fail-fast-with-heap-dump beats letting a service thrash at 100% CPU in a fleet.
## The setup: GC, thrashing, and 'progress' The garbage collector frees unreachable heap objects so your program can keep allocating. Normally GC is a small fraction of runtime. But imagine the heap is **almost entirely full of objects that are still reachable** (a leak, or a genuinely too-small heap). Now: - The app allocates a little → heap fills → GC runs. - GC frees only a tiny sliver (most objects are still in use). - The app immediately re-fills that sliver → GC runs again. The program is technically running but accomplishing almost nothing — it's **thrashing**: spending nearly all CPU on garbage collection, freeing almost no memory each cycle. Without intervention this could limp along for a long time, pinned at near-100% CPU, doing no useful work. ## What the error means The **throughput / Parallel collector** has a built-in tripwire for this. By default it throws `OutOfMemoryError: GC overhead limit exceeded` when, over recent collections, **more than ~98% of total time is spent in GC** and **less than ~2% of the heap is recovered**. The exact heuristic: this condition must hold across several consecutive full GCs. The JVM is saying: 'I'm spending essentially all my time collecting and reclaiming essentially nothing — there's no point continuing; fail now.' It's a **fail-fast** mechanism: better to die promptly with a clear signal than to hang in a GC death-spiral. ## How it differs from 'Java heap space' They describe the **same root condition** — the heap is effectively exhausted by reachable objects — but at **different moments**: - **GC overhead limit exceeded**: caught *early*, when the app is still alive but has become all-GC-no-work. The collector recognizes the futility before a single allocation literally cannot be satisfied. - **Java heap space**: thrown when a *specific* allocation cannot be fulfilled even after GC — the hard wall. In practice an app under a leak may throw either, depending on allocation pattern and collector. 'GC overhead' often appears slightly sooner, as a symptom of imminent exhaustion. Both are heap problems; both have the same cause families (leak / under-size / sustained ceiling allocation) and the same fix path. ## Diagnosis and fix Identical to heap-space OOME: 1. `-XX:+HeapDumpOnOutOfMemoryError`, capture the dump. 2. In MAT/VisualVM, inspect dominator tree and retained sizes; compare over time. 3. **Leak** (a structure's retained size keeps growing) → fix reachability (evict caches, deregister listeners, `remove()` ThreadLocals). 4. **Under-size** (heap legitimately full, flat over time) → raise `-Xmx` (and check container limits). 5. Watch GC logs (`-Xlog:gc`) — long, frequent full GCs reclaiming little confirm the overhead spiral. ## The tempting wrong fix `-XX:-UseGCOverheadLimit` disables *just this check*. It does **not** fix anything — it removes the early tripwire, so the JVM will instead thrash longer and eventually throw plain 'Java heap space' (or hang at high CPU). Almost always the wrong move; the message is doing you a favor by failing fast. Also note: this tripwire is specific to the throughput/Parallel collector's accounting. Other collectors (G1, ZGC) may surface heap exhaustion differently, but the underlying 'GC busy, reclaiming nothing' condition is the same warning sign. ## Key takeaways - Meaning: ~98% time in GC, <2% heap reclaimed → JVM fails fast instead of thrashing. - Same root cause as 'Java heap space' (heap effectively full of reachable objects), caught earlier. - Causes and fixes are identical: heap dump → leak vs. under-size → fix reachability or raise -Xmx. - Disabling the check (-XX:-UseGCOverheadLimit) hides the warning and just delays a harder crash.
- Is 'GC overhead limit exceeded' a fundamentally different problem from 'Java heap space'?No — it's the same underlying heap exhaustion, just detected earlier by the throughput collector when GC becomes nearly all-effort-no-reward. Diagnosis (heap dump) and fix (leak vs. resize) are the same.
- When would using -XX:-UseGCOverheadLimit be reasonable?Rarely. It only suppresses the early tripwire; the heap problem remains and you'll get a plain heap-space OOME (or a CPU-pinned hang) later. It's not a fix — only a last resort if the early failure interferes with capturing diagnostics, and even then a proper heap dump is better.
It's like a bailing crew on a sinking boat scooping nonstop but the water keeps rising — at some point the captain calls it rather than have everyone bail forever for nothing. The JVM 'calls it' when GC effort is total and the payoff is near zero.
saying these in an interview costs you the question
- Treating it as a totally separate bug from heap exhaustion — it's the same root cause caught earlier.
- Fixing it with -XX:-UseGCOverheadLimit and considering the problem solved (it only delays a harder crash).
- Assuming it means the GC is broken — the GC is working overtime; the problem is too little reclaimable memory.
- Forgetting the thresholds: ~98% time in GC and <2% heap reclaimed.