skip to content

A service running with -XX:+UseShenandoahGC normally pauses for under 5 ms, but its logs occasionally show a 'Pause Degenerated GC' of several hundred milliseconds and, rarely, a multi-second full GC. What causes these and how would you address them?

level: seniorimportance: should knowfreq 24%

answer

  1. Degenerated = finish the in-flight cycle STW; log names the phase
  2. Full GC = whole-heap sliding compaction, last-ditch
  3. Root cause: allocation beat the collector / cycle started too late
  4. Fixes: headroom first, earlier trigger, ConcGCThreads, allocate less
  5. Humongous allocations + fragmentation also force full GC

basics

~20 s

The collector lost the race against allocation: the heap ran out before the concurrent cycle finished. It then finishes the remaining work stop-the-world (degenerated GC), or falls back to a full stop-the-world compaction if that also fails. Fix by giving more heap headroom, starting cycles earlier, adding concurrent GC threads, or reducing allocation rate.

solid answer

~60 s

Both are **failure modes of the concurrent cycle**, not normal operation. - **Degenerated GC**: an allocation request fails while a concurrent cycle is in progress. Rather than throwing the partial work away, the collector stops the world and finishes the *current* cycle's remaining phases (mark, evacuation, or update-refs — the log says which) with all threads helping. The pause is proportional to the work left, so it is far longer than a normal pause but far shorter than a full GC. - **Full GC**: the last-ditch fallback — a stop-the-world, whole-heap sliding compaction. It happens when even a degenerated cycle cannot free enough, or when fragmentation blocks a large (humongous) allocation. **Root cause is almost always the same**: mutators allocate faster than the collector reclaims, or the cycle started too late, or there is not enough free heap to evacuate into. **Remedies, in order of effect**: increase heap size so there is real headroom; make the heuristic start cycles earlier; raise `-XX:ConcGCThreads` if cores are available; reduce allocation rate in the application; check for humongous allocations causing fragmentation. Investigate a rising live set as a leak before tuning anything.

code

text · 6 lines
text
[311.4s] GC(88) Trigger: Free (48M) is below minimum threshold (96M)
[311.4s] GC(88) Pause Init Mark 0.401ms
[311.7s] GC(88) Cancelling GC: Allocation Failure
[311.7s] GC(89) Pause Degenerated GC (Mark) 512.442ms
...
[402.9s] GC(94) Pause Full (Allocation Failure) 3184.771ms

go deeper

for a junior

Know that these are fallbacks that stop the world when the collector cannot keep up with allocation.

for a middle

Distinguish degenerated (finish the current cycle stop-the-world) from full GC (whole-heap compaction), and name headroom and trigger timing as the levers.

for a senior

Diagnose from the log — trigger reasons, cycle duration versus allocation rate, live-set trend, humongous allocations — and order remedies by leverage.

for a principal

Set fleet policy: alert on full and degenerated GCs, define heap-headroom floors for the collector, and separate leak investigations from tuning work.

## Why a concurrent collector can be forced to stop the world A concurrent collector is in a race. On one side, application threads allocate; on the other, the collector marks, evacuates, and frees regions. As long as the collector finishes a cycle before the free space runs out, pauses stay short. When allocation wins the race, there is nowhere to allocate from, and the only remaining options involve stopping the application. Shenandoah has two escalating fallbacks, and reading the log to tell them apart is the first diagnostic step. ## Degenerated GC When an allocation cannot be satisfied while a concurrent cycle is running, the collector *degenerates*: it stops the world and completes the remainder of the in-flight cycle with the world stopped, using all available threads. The log names the phase at which it degenerated: ``` Pause Degenerated GC (Mark) 512.4ms Pause Degenerated GC (Evacuation) 233.1ms Pause Degenerated GC (Update Refs) 118.7ms ``` This is a deliberate, well-behaved fallback: the work already done concurrently is kept, so the pause covers only what was left. Degenerating at Mark is the most expensive (most work remains); degenerating at Update Refs is the cheapest. An occasional degenerated GC under an extreme allocation spike is survivable; a regular one means the collector is chronically behind and the configuration is wrong. ## Full GC If a degenerated cycle still cannot produce enough free space — or if a **humongous** allocation (an object spanning multiple contiguous regions) cannot be satisfied because free regions are not contiguous — the collector falls back to a full stop-the-world compaction of the entire heap. This is the last-ditch mechanism that guarantees the JVM does not throw OutOfMemoryError while memory is still reclaimable. On a large heap it can take seconds, which is exactly the outcome the collector was chosen to avoid. Any full GC in a latency-sensitive service warrants investigation. ## Diagnosing the root cause Run with detailed logging and read four things: 1. **Allocation rate** versus **cycle duration**. If a concurrent cycle takes 300 ms and the application can fill the free heap in 200 ms, degeneration is arithmetic, not bad luck. 2. **The trigger line.** Shenandoah logs why it started a cycle ("Trigger: Free (…) is below minimum threshold", "Trigger: Average GC time … is above the time for allocation rate"). A trigger that fires only when free space is nearly gone means the cycle starts too late. 3. **Live set over time.** Occupancy after each cycle. If it climbs steadily, you have a leak or a growing cache, and no GC flag will fix it. 4. **Humongous allocation counts.** Large arrays that span regions can fragment the heap and force full GCs even when total free space looks adequate. ## Remedies, ordered by how much they actually help **1. More heap headroom.** A concurrent evacuating collector must copy live objects *somewhere*, so it cannot run near full occupancy the way a stop-the-world compactor can. Running Shenandoah at 90%+ occupancy is asking for degeneration. Raising `-Xmx` (and giving the container the memory to back it) is usually the highest-leverage fix. **2. Start cycles earlier.** The heuristic decides when to begin. `-XX:ShenandoahGCHeuristics=adaptive` is the default and learns from recent cycles; `static` starts on a fixed free-space threshold; `compact` runs cycles continuously to minimise footprint; `aggressive` collects back-to-back and is a diagnostic tool, not a production setting. Related knobs (diagnostic/experimental, so verify on your build) include `-XX:ShenandoahMinFreeThreshold`, `-XX:ShenandoahAllocSpikeFactor`, and `-XX:ShenandoahGarbageThreshold`. Start earlier and you trade a little throughput for a much bigger safety margin. **3. More collector threads.** `-XX:ConcGCThreads` raises concurrent capacity — but only if the machine has spare cores. On a CPU-saturated container this steals from the application and can make matters worse. **4. Allocate less.** The unglamorous but permanent fix: eliminate per-request garbage, reuse buffers, avoid gratuitous copies and boxing. Reducing allocation rate helps every collector and costs no resources. **5. Pacing.** Shenandoah includes an allocation **pacer** that briefly stalls threads allocating far faster than the collector can keep up, buying the cycle time to finish. It is on by default and shows up as small allocation stalls in the log; those stalls are the collector trying to prevent a degeneration and are usually preferable to one. ## What good looks like afterwards Zero full GCs, degenerated GCs rare or absent, cycles completing with meaningful free space still available, and a flat live set across cycles. Alert on `Pause Full` and on degenerated-GC frequency — they are the early warning that a service has outgrown its heap or started leaking.

  • Why can't the collector simply run at 95% heap occupancy the way a stop-the-world compactor can?
    Because it compacts by copying live objects into free regions while the application runs, so it needs free space to copy into for the duration of the cycle, plus room for whatever the application allocates during that time. Occupancy that high leaves no destination space and no slack for the allocation rate, so the cycle cannot finish and degenerates.
  • You see frequent degenerated GCs at the Update Refs phase specifically. Is that better or worse than degenerating at Mark?
    Better, in the sense that most of the cycle's work was already completed concurrently, so the resulting pause is much shorter. It still signals that the collector is finishing too close to exhaustion. The remedy is the same — more headroom or an earlier trigger — but the urgency is lower than for degeneration at Mark, where nearly the whole cycle must be redone with the world stopped.
  • The live set grows steadily across cycles. What do you do before touching any GC flag?
    Treat it as a suspected memory leak or an unbounded cache rather than a tuning problem. Capture a heap dump at a high-water mark and inspect what retains the growing objects. No collector configuration can reclaim memory that is still reachable; tuning would only delay the failure and obscure the cause.

A cleaning crew working around guests: if the rubbish piles up faster than they can carry it out, they eventually have to clear the room — and if that still isn't enough, close the building.

saying these in an interview costs you the question

  • Treating degenerated GC as a normal, expected part of the cycle
  • Reaching for -XX:ConcGCThreads on a CPU-saturated container, which steals from the application
  • Assuming a low-pause collector removes the need for heap headroom
  • Blaming the collector for a steadily growing live set instead of investigating a leak
  • Believing a full GC in this collector is also concurrent — it is stop-the-world whole-heap compaction

context