skip to content

VM Runtime Tuning

Configuring the runtime itself — heap size, generational layout, collector choice, and GC logging — to balance throughput, latency, and footprint. Interviewers raise it in production-facing rounds, where the real question is less "which flag" and more "what did you measure first".

on this pageshow

explore

questions

19

A HotSpot JVM prints this garbage-collection log line: `[3.216s][info][gc] GC(7) Pause Young (Normal) (G1 Evacuation Pause) 512M->96M(1024M) 14.221ms`. Walk through what each field on that line means.

level: juniorimportance: must knowfreq 55%

answer

  1. uptime · level · tags · GC(id)
  2. Pause = stop-the-world; Concurrent = not a pause
  3. before->after(capacity)
  4. "Allocation Failure" is normal, not an error
  5. cause in parentheses answers *why now*

basics

~20 s

[3.216s] is uptime, info the log level, gc the tag, GC(7) the collection's id. It was a young pause caused by G1 evacuation. Heap used went 512M before to 96M after, of 1024M total, taking 14.221 ms stop-the-world.

solid answer

~50 s

Reading left to right: `[3.216s]` is a decorator — JVM uptime when the line was written; `[info]` is the log level and `[gc]` the tag selector that produced it. `GC(7)` is the collection id, which lets you correlate every sub-line (phases, heap breakdown, CPU) belonging to the same collection. `Pause` means stop-the-world, as opposed to a `Concurrent` line that runs alongside the application. `Young (Normal)` is what was collected — the young generation, in the ordinary (not mixed, not initial-mark) flavour. `(G1 Evacuation Pause)` is the *cause*: G1 filled eden and evacuated the live objects out. Then `512M->96M(1024M)`: heap **used** before the collection, heap used after, and current heap **capacity**. So ~416M of garbage died, 96M survived. `14.221ms` is the wall-clock pause duration. Nothing here indicates a problem: this is a healthy, cheap young collection.

go deeper

for a junior

Be able to name every field in order and state that before->after(capacity) is used-heap before, used-heap after, and current capacity, and that the pause is stop-the-world.

for a middle

Add the causes you have seen and what each implies, and distinguish Pause lines from Concurrent lines when totalling pause time.

for a senior

Read a sequence rather than a line: derive interval, survivor volume and trend, and know to pull gc+cpu when Real exceeds User+Sys.

for a principal

Frame it as observability: which decorators and tags a fleet standardises on so GC lines can be correlated with request traces and incident timelines.

## Where the line comes from Since JDK 9 all HotSpot logging goes through *unified logging*, controlled by `-Xlog`. A GC line has three parts: decorators in square brackets, then the message. Decorators are chosen by you (`time`, `uptime`, `level`, `tags`, `pid`, `tid`); the default for `-Xlog:gc` is uptime, level and tags, which is exactly what the example shows. ## Field by field **`[3.216s]`** — the `uptime` decorator: seconds since JVM start. Add the `time` decorator if you need wall-clock timestamps to line up GC events with application logs or an incident timeline. Uptime alone is fine for measuring intervals between collections. **`[info]`** — the log level. GC lines you normally read are `info`; `debug` and `trace` levels expose per-phase and per-region detail and are opt-in because they are much noisier. **`[gc]`** — the tag set. Sub-systems tag their output: `gc`, `gc+heap`, `gc+cpu`, `gc+age`, `gc+ergo`, `gc+phases`. `-Xlog:gc*` selects all of them; `-Xlog:gc` selects only the top-level summary lines like this one. **`GC(7)`** — the collection sequence number, counting from 0 for the life of the JVM. Every line emitted for that collection carries the same id, so with `gc*` enabled you can gather the phase breakdown, the region counts and the CPU line for collection 7 even though they are interleaved with other output. **`Pause`** — this collection stopped every application thread. The counterpart is `Concurrent`, e.g. `Concurrent Mark Cycle`, which runs while application threads keep going and therefore is *not* a pause even though it appears in the log and takes wall-clock time. Beginners routinely add up concurrent durations and report a pause budget that never happened. **`Young (Normal)`** — the scope. In a G1 log you will also see `Young (Mixed)`, where some old regions are collected alongside the young ones, `Young (Prepare Mixed)`, `Young (Concurrent Start)`, and `Full`, which collects and compacts the entire heap. In Serial and Parallel logs the same slot reads simply `Young` or `Full`. **`(G1 Evacuation Pause)`** — the *cause*, i.e. why the JVM decided to collect now. Common causes and what they mean: - `G1 Evacuation Pause` / `Allocation Failure` — normal: a thread wanted memory and eden had none left. The word *failure* alarms people; it is the ordinary trigger of every young collection, not an error. - `G1 Humongous Allocation` — an object at least half a region in size forced a collection. - `Metadata GC Threshold` — class metadata space, not the Java heap, needed room. - `System.gc()` — application or a library called it explicitly. Worth chasing down. - `GCLocker Initiated GC` — a collection was deferred because a thread was inside a JNI critical section. - `Ergonomics` — the collector's own heuristics decided to act. **`512M->96M(1024M)`** — used-before → used-after (current capacity). The drop tells you how much died; the after value approximates the *live set* at that moment for a young collection (plus whatever old-generation data was already there). The capacity in parentheses is the currently committed heap, which can grow toward `-Xmx` or shrink; if it is far below `-Xmx`, the heap has not expanded yet. **`14.221ms`** — wall-clock duration of the stop-the-world portion. This is what your latency percentiles feel. It is *not* CPU time: the `gc+cpu` tag prints `User`, `Sys` and `Real` separately, and a Real much larger than User+Sys means the machine, not the collector, was the problem (CPU starvation, swapping, slow disk on the log write). ## Reading it as a whole One line rarely means anything; a sequence does. From consecutive lines you get the interval between collections, how much was allocated in between, how much survived each time, and whether the after-values trend upward. This single line says: at 3.2 seconds in, a routine young collection reclaimed ~416 MB in 14 ms. Healthy.

  • The cause on a line reads `(Allocation Failure)`. Does that mean the JVM is running out of memory?
    No. Allocation Failure simply means a thread requested space in eden and eden was full, which is the normal trigger for every young collection. A healthy application produces thousands of them. Genuine exhaustion shows up as repeated Full GCs that reclaim almost nothing, and ultimately as an OutOfMemoryError.
  • What is the difference between the number in parentheses and `-Xmx`?
    The parenthesised value is the currently *committed* heap capacity, which the JVM grows and shrinks between `-Xms` and `-Xmx` according to its heuristics. `-Xmx` is only the ceiling. Seeing a small capacity early in a run usually means the heap has simply not expanded yet, not that your maximum was ignored.

saying these in an interview costs you the question

  • Reading `Allocation Failure` as an error or an imminent OutOfMemoryError.
  • Treating `512M->96M(1024M)` as though the last number were the heap *used*, rather than the committed capacity.
  • Adding `Concurrent ...` durations into a pause-time budget — they do not stop application threads.
  • Assuming the duration is CPU time consumed by GC rather than wall-clock stop-the-world time.

context

open as a page

Which garbage collector does a modern HotSpot JVM use if you specify nothing, and what makes it pick a different one on its own?

level: juniorimportance: must knowfreq 55%

basics

~20 s

On JDK 9 and later the default is G1 on any server-class machine (roughly two or more CPUs and enough memory). On a smaller machine or container, JVM ergonomics falls back to Serial. On JDK 8 the default was the Parallel collector.

open as a page

What do the Java command-line options -Xms and -Xmx set, and what does the HotSpot JVM choose for them when you specify neither?

level: juniorimportance: must knowfreq 55%

basics

~20 s

-Xms is the initial heap size the JVM commits at startup; -Xmx is the maximum it may grow to. With neither set, HotSpot defaults to roughly 1/64 of available memory for initial and 1/4 for maximum, reading the container limit rather than host RAM when running under one.

open as a page

How do you enable garbage-collection logging on a HotSpot JVM running JDK 9 or later, and which options control the log destination, the level of detail, and file rotation?

level: middleimportance: must knowfreq 60%

basics

~10 s

Use the unified logging flag -Xlog. Its four colon-separated parts are selectors, output, decorators, output options: -Xlog:gc*:file=gc.log:time,uptime,level,tags:filecount=10,filesize=20M. It replaced -XX:+PrintGCDetails and -Xloggc, and is cheap enough to leave on in production.

open as a page

Reading a HotSpot garbage-collection log, how do you tell a young (minor) collection from a mixed collection and from a full collection, and why is spotting a full collection important?

level: middleimportance: must knowfreq 58%

basics

~20 s

Young collections log Pause Young, G1's mixed ones Pause Young (Mixed) (young plus some old regions), and full ones Pause Full, which collects and compacts the whole heap in one long stop. Under G1 or a concurrent collector, a Full GC signals the concurrent machinery failed to keep up.

open as a page

When would you choose HotSpot's throughput-oriented Parallel collector (-XX:+UseParallelGC) over G1, and what do you give up by doing so?

level: middleimportance: must knowfreq 60%

basics

~20 s

Choose Parallel when total work per unit time matters and pauses do not — batch jobs, ETL, offline analytics, benchmarks. It usually delivers a few percent more throughput because it does no concurrent work and no write barriers. You give up pause control: its collections stop everything for a time that scales with the heap.

open as a page

In a generational Java heap, how does enlarging the young generation change the frequency and the duration of minor collections, and when does a larger young generation actually reduce total GC overhead?

level: middleimportance: must knowfreq 60%

basics

~20 s

A bigger young generation fills more slowly, so minor collections happen less often. Each collection copies only the survivors, so its pause tracks surviving data, not eden size. Because most objects die fast, a bigger young space usually means fewer collections and fewer survivors per collection.

open as a page

Why is it common practice to set -Xms equal to -Xmx on long-running server JVMs, and what is the argument against doing it?

level: middleimportance: must knowfreq 60%

basics

~20 s

Pinning them equal stops the JVM repeatedly committing and releasing memory, which costs collection work and page faults, and it makes the process footprint predictable so the peak cannot surprise a container limit later. Against it: you pay for memory you may never use, and elastic services lose the ability to give it back.

open as a page

A JVM service normally shows 15 ms garbage-collection pauses but occasionally stops for nearly a second. Working only from GC logs, how do you find out what causes the outliers?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Sort collections by duration, then read the outliers' lines: their scope and cause, whether To-space exhausted or a Full GC preceded them, the phase breakdown from gc+phases, the gc+cpu User/Sys/Real split, and safepoint timings. Cause first, then flags.

open as a page

When is a fully concurrent low-pause collector such as ZGC or Shenandoah the right choice for a JVM service, and what does it cost you compared with G1?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Choose them when a strict tail-latency SLO must hold on a large heap — pauses stay sub-millisecond to low-millisecond largely independent of heap size, because marking and relocation run concurrently. You pay barrier overhead on application code, extra CPU for concurrent work, and more heap headroom.

open as a page

A Java service in a container with a 2 GiB memory limit keeps being killed by the kernel out-of-memory killer, yet it never throws a Java heap OutOfMemoryError. How do you size the JVM so this stops happening?

level: seniorimportance: must knowfreq 60%

basics

~20 s

The container limit covers the whole process, not just the heap. Set the maximum heap to a fraction of the limit - commonly 50-75% depending on thread count and off-heap use - leaving room for metaspace, code cache, stacks, direct buffers and GC structures, then verify the real resident footprint with Native Memory Tracking.

open as a page

What does the JVM flag -XX:MaxGCPauseMillis actually do, and what happens if you set it to 5 milliseconds on a G1 heap?

level: middleimportance: should knowfreq 45%

basics

~20 s

It is a soft goal, not a guarantee. G1 uses it to size each collection set: to hit a smaller target it collects less per pause, so pauses become shorter but more frequent, throughput drops, and old-generation reclamation can fall behind. An unrealistic 5 ms mostly costs throughput without delivering 5 ms.

open as a page

Inside a HotSpot young generation, how is space divided between eden and the two survivor spaces, what does the flag -XX:SurvivorRatio change, and what goes wrong when the survivor spaces are too small?

level: middleimportance: should knowfreq 45%

basics

~20 s

The young generation is eden plus two equal survivor spaces; only one survivor holds data at a time. -XX:SurvivorRatio=N makes eden N times one survivor. If the surviving objects do not fit in the target survivor space, they are promoted to the old generation immediately, regardless of age.

open as a page

Why can raising a HotSpot maximum heap setting from 31 GB to 33 GB leave a Java service with less usable capacity than it had before?

level: middleimportance: should knowfreq 35%

basics

~20 s

Below roughly 32 GB, HotSpot stores object references as 32-bit compressed pointers. Above that threshold it must use full 64-bit references, so every object with references grows and the whole live set expands - typically enough to swallow the extra 2 GB and more.

open as a page

How do you compute an application's allocation rate and promotion rate from a HotSpot garbage-collection log, and what do those two numbers tell you?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Allocation rate = (heap used before a collection − heap used after the previous one) ÷ the interval between them, averaged over many collections. Promotion rate = old-generation growth per collection ÷ the same interval. High allocation means frequent young pauses; high promotion means old-generation pressure and eventual long collections.

open as a page

Beyond the Java heap, what else consumes memory in a running JVM process, and how do you put an explicit bound on each part when budgeting a service's total memory footprint?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Metaspace and compressed class space, the JIT code cache, one stack per thread, direct and mapped byte buffers, garbage-collector bookkeeping, and native library allocations. Bound them with -XX:MaxMetaspaceSize, -XX:ReservedCodeCacheSize, -Xss plus bounded thread pools, and -XX:MaxDirectMemorySize; measure the rest with Native Memory Tracking.

open as a page

You own a fleet of JVM services with mixed workloads — request-serving APIs, streaming consumers, nightly batch jobs. How do you decide which garbage collector each should run, and how do you validate the decision?

level: principalimportance: should knowfreq 40%

basics

~20 s

Classify each service by its governing metric — tail latency, throughput, or footprint — then map: latency SLO plus a large heap goes to a concurrent collector, batch to a throughput collector, everything else to the pause-targeting default. Validate by running the real workload under load with GC logging and comparing pause percentiles, throughput and CPU.

open as a page

HotSpot ships a no-op garbage collector, Epsilon (-XX:+UseEpsilonGC), which allocates memory but never reclaims it. What is it actually for?

level: middleimportance: nice to knowfreq 20%

basics

~20 s

Epsilon allocates and never collects; when the heap is exhausted the JVM exits with an OutOfMemoryError. It exists for performance testing where GC must be excluded, for measuring an application's true allocation footprint, for last-drop latency experiments, and for very short-lived processes that finish before the heap fills.

open as a page

Region-based collectors such as G1 size the young generation adaptively to meet a pause-time goal. On a production service, when would you override that with a fixed young-generation size, and what do you give up by doing so?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Rarely. Pinning the young size with -Xmn disables the adaptive sizing that serves the pause goal, so the collector can no longer trade young size for latency as the workload shifts. Justify it only for throughput-only batch work, benchmarking reproducibility, or when adaptation demonstrably misbehaves - and prefer raising the floor instead.

open as a page