skip to content

Why is it common practice to set -Xms equal to -Xmx on long-running server JVMs, and what is the argument against doing it?

level: middleimportance: must knowfreq 60%

answer

  1. MinHeapFreeRatio 40 / MaxHeapFreeRatio 70 drive resize
  2. resize costs page faults inside request latency
  3. pinning turns a peak-time OOM kill into a startup failure
  4. AlwaysPreTouch: honest footprint, slower startup
  5. against: density, cost, elastic uncommit

basics

~20 s

Pinning them equal stops the JVM repeatedly committing and releasing memory, which costs collection work and page faults, and it makes the process footprint predictable so the peak cannot surprise a container limit later. Against it: you pay for memory you may never use, and elastic services lose the ability to give it back.

solid answer

~50 s

Between -Xms and -Xmx the JVM resizes the heap, expanding when free space after a collection falls below -XX:MinHeapFreeRatio and shrinking when it rises above -XX:MaxHeapFreeRatio. On a server whose load varies, that produces a grow-shrink cycle: memory is uncommitted and then re-committed, each expansion faulting in fresh pages during request processing, and in some collectors resize decisions are tied to collections. Pinning -Xms to -Xmx removes the cycle. The footprint is decided at startup, so capacity planning is exact and you discover a too-large heap immediately instead of at peak traffic, when the container limit turns it into a kernel kill. Adding -XX:+AlwaysPreTouch faults the pages in at startup so the cost is not paid during requests, at the price of slower startup. The counter-argument is density and cost: a pinned heap holds memory the service may never need, and modern collectors can return unused memory to the operating system. For scale-to-zero, bursty, or heavily co-located workloads, a lower -Xms is deliberate.

code

text · 5 lines
text
# long-running dedicated service: fixed footprint, pages faulted in at startup
java -Xms3g -Xmx3g -XX:+AlwaysPreTouch -jar app.jar

# elastic/idle-heavy service: floor at steady state, allow return of memory
java -Xms1g -Xmx3g -XX:MinHeapFreeRatio=20 -XX:MaxHeapFreeRatio=40 -jar app.jar

go deeper

for a junior

Know that a differing -Xms and -Xmx lets the heap grow and shrink, and that servers often pin them equal for stability.

for a middle

Name the free-ratio flags, explain the page-fault and resize cost, and state the predictability argument.

for a senior

Argue both sides with the failure modes: late OOM kills at peak versus wasted committed memory, and place AlwaysPreTouch correctly.

for a principal

Make it a platform policy tied to workload class - pinned for long-running dedicated services, elastic for bursty or co-located ones - and tie both to a documented whole-process memory budget.

## What the JVM does when the two differ The heap lives between an initial committed size (-Xms) and a maximum (-Xmx). After collections, HotSpot compares the proportion of free heap against -XX:MinHeapFreeRatio (default 40) and -XX:MaxHeapFreeRatio (default 70). Too little free and it expands by committing more of the reserved address range; too much free and it shrinks, returning pages to the operating system. G1 additionally uncommits at concurrent-cycle boundaries as well as at full collections, and -XX:-ShrinkHeapInSteps changes whether shrinking happens gradually or in one move. So a heap that is not pinned tracks demand. That sounds purely good, and on a laptop it is. ## Why servers pin them equal First, resizing is not free. Committing memory means the operating system must supply pages, and the first write to each page takes a fault; on a busy service that cost lands inside request latency. Shrinking then re-growing repeats it. Under load spikes - exactly when you care - the JVM is doing the most resizing. Second, resize decisions are entangled with collection. Heap expansion and, in the older collectors, most shrinking happen at collection boundaries, and heap-occupancy-driven heuristics such as when to start a concurrent cycle are computed against the current committed size. A heap that keeps changing size keeps moving those thresholds, so behaviour is less reproducible and tuning done at one size may not hold at another. Third, and most important operationally: predictability. With -Xms equal to -Xmx the process footprint is settled at startup. If that footprint does not fit the container, the service fails immediately and visibly. With a small -Xms it starts happily and grows into the limit hours later at peak traffic, where exceeding it means a kernel OOM kill of the whole process rather than a Java-level error. Pinning converts a latent production failure into a startup failure. Fourth, stability of layout. A stable committed heap gives collectors a stable region or generation layout; repeatedly returning and re-acquiring memory means re-establishing that structure and can leave the address space and the collector's bookkeeping in a less tidy state on a process that runs for weeks. Many teams pair the pinned heap with -XX:+AlwaysPreTouch, which writes to every heap page during startup so the pages are really backed before traffic arrives. This trades a longer, memory-hungry startup for the absence of first-touch faults later, and it makes the resident footprint honest immediately rather than creeping upward. ## The case against The strongest counter-argument is money and density. A pinned heap holds its full size whether or not the application needs it, so co-locating many services on one node, or running functions that idle most of the time, wastes real capacity. Modern collectors are markedly better at giving memory back - G1 uncommits during concurrent cycles, and ZGC and Shenandoah have explicit uncommit behaviour with idle delays - which makes elasticity a genuine option rather than a theoretical one. There are also environments where a large initial commitment is actively harmful: short-lived CLI or serverless invocations where startup time dominates, development machines, and platforms that bill on observed memory rather than on the declared limit. ## A defensible middle position Set -Xms to the size the service actually needs in steady state, measured as live set plus collection headroom, and -Xmx to that same value on a dedicated, long-running server. Where elasticity has real value, keep -Xms at the honest steady-state floor and allow -Xmx some room above it, accepting occasional resize cost in exchange for returning memory when idle. What you should not do is leave -Xms at a tiny default on a service that will certainly grow: you get all of the resize cost and all of the late-failure risk with none of the density benefit. Whichever you choose, the heap bounds must be justified against the whole process budget, since the heap is only part of what the container is charged for.

  • What does -XX:+AlwaysPreTouch do, and what does it cost?
    It writes to every page of the committed heap during startup so the operating system backs them with real memory before the application runs. The benefit is that no first-touch page faults occur during request processing and resident memory is truthful immediately. The cost is a slower startup proportional to heap size, and the process holds its full footprint from second zero.
  • Which HotSpot flags govern whether the heap grows or shrinks when -Xms is below -Xmx?
    -XX:MinHeapFreeRatio and -XX:MaxHeapFreeRatio, defaulting to 40 and 70. After a collection, if free heap is below the minimum the JVM expands; if it is above the maximum it may shrink. Narrowing the band makes the heap track demand more closely at the cost of more resizing; widening it makes the size stickier.

saying these in an interview costs you the question

  • Claiming -Xms equals -Xmx improves throughput because 'the GC runs less'
  • Believing the JVM never returns heap memory to the operating system on modern collectors
  • Assuming a pinned -Xms is always right, including for idle or serverless workloads
  • Thinking pre-touch is free rather than a startup-time and footprint tradeoff
  • Setting -Xmx to the whole container limit because the heap is pinned anyway

context