skip to content

Why does a GraalVM native-image Spring Boot app often have lower peak throughput than the same app on the JVM, and what is the headline lever to close that gap?

level: juniorimportance: should knowfreq 45%

answer

  1. No JIT = no re-profiling
  2. PGO instrument -> run -> optimize
  3. -O2 default, -O3 aggressive
  4. Serial GC default -> G1 for throughput
  5. PGO + G1 = Oracle GraalVM only

basics

~20 s

A native image is compiled ahead-of-time and has no JIT, so it can't re-optimize hot code from runtime profiles. Peak throughput lags the JVM. Profile-Guided Optimization (PGO) feeds a runtime profile back into the build to close most of the gap.

solid answer

~40 s

On the JVM the JIT watches the running program and recompiles hot methods with speculative, profile-driven optimizations, so throughput climbs after warm-up. A GraalVM native image is fully ahead-of-time compiled: there is no JIT, no runtime re-profiling, so it starts fast but its peak throughput is typically lower because the compiler had to guess without runtime data. The main lever to close this is Profile-Guided Optimization (PGO), available in Oracle GraalVM: you build an instrumented binary, run a representative workload to collect a profile (default.iprof), then rebuild using that profile so the AOT compiler optimizes the actually-hot paths. Secondary levers are the optimization level (-O2 default, -O3 aggressive) and switching the garbage collector from the default Serial GC to G1 for throughput-bound, larger-heap services.

go deeper

for a junior

Know the one-liner: native has no JIT, so no runtime re-optimization; PGO feeds a profile back in to recover throughput.

for a middle

Be able to name all three levers (PGO, -O level, GC choice) and that PGO is a 3-step build.

for a senior

Explain the closed-world/AOT reason for the gap and that PGO/G1 are Oracle-GraalVM-only, with licensing/CI implications.

for a principal

Frame it as a cost/benefit decision: whether the throughput recovery justifies Oracle GraalVM licensing and a more complex, slower CI pipeline.

## The throughput gap **JIT (Just-In-Time) compilation** on the HotSpot JVM starts your app in interpreted/lightly-compiled mode, then continuously profiles execution (which branches are taken, which types actually flow through a call site) and recompiles hot methods with aggressive, *speculative* optimizations (inlining, branch pruning, monomorphic dispatch). This is why a JVM service is slow for the first seconds (warm-up) but reaches high steady-state throughput. A **GraalVM native image** is produced by `native-image` (invoked from Spring Boot via the `org.graalvm.buildtools.native` plugin). All compilation happens **ahead-of-time (AOT)** at build time under the **closed-world assumption** — the whole reachable program is known and compiled up front. At runtime there is **no JIT, no re-profiling, no recompilation**. The upside is near-instant startup and low, flat memory. The downside: the compiler optimized the code *blind*, without knowing which paths are actually hot, so **peak (steady-state) throughput is usually lower than the warmed-up JVM**. ## The levers that close the gap 1. **PGO (Profile-Guided Optimization)** — the biggest lever. A three-step build: (a) build an **instrumented** binary (`--pgo-instrument`), (b) run it under a **representative load** to emit a profile file `default.iprof`, (c) rebuild feeding that profile (`--pgo=default.iprof`). The AOT compiler now knows the real hot paths, branch frequencies and type profiles, so it inlines and lays out code the way a JIT eventually would — recovering much of the lost throughput. **PGO is an Oracle GraalVM feature; GraalVM Community Edition does not have it.** 2. **Optimization level `-O`** — `-O0` (no optimization, for debugging), `-O1`, `-O2` (**default**, recommended for production), `-O3` (more aggressive, Oracle GraalVM), and `-Ob` (**optimize *build* time**, i.e. faster/cheaper builds for local dev — NOT for production throughput). Passed as a build arg. 3. **Garbage collector** — a native image defaults to the **Serial GC** (single-threaded, small footprint, great for short-lived CLIs and memory-constrained containers, but stop-the-world pauses hurt throughput under load). Oracle GraalVM offers **G1 GC** (`--gc=G1`, Linux only) for throughput-bound, larger-heap, long-running services. There is also **Epsilon GC** (`--gc=epsilon`), a no-op collector for extremely short-lived jobs. ## How you set these in a Spring Boot build All three are `native-image` build arguments configured through the Native Build Tools plugin's `buildArgs` (Gradle `graalvmNative { binaries { named("main") { buildArgs.add(...) } } }`) or `<buildArgs>` in the Maven `native-maven-plugin`. ## When to bother - Latency/startup-driven serverless or CLI: stock native image (Serial GC, -O2) is usually fine; the throughput gap doesn't matter. - Long-running, load-bearing microservices where you switched to native for startup/footprint but still need throughput: invest in PGO + G1, and license Oracle GraalVM. ## Gotchas - PGO and G1 require **Oracle GraalVM**; on Community Edition you only get Serial/Epsilon GC and no PGO. - The PGO profile is only as good as the workload you replay — an unrepresentative load produces a misleading profile. - G1 raises memory footprint and build complexity — you may lose part of native image's footprint advantage. - `-Ob` is a build-time optimization, not a runtime one; shipping it to prod hurts throughput.

  • Does GraalVM Community Edition support PGO?
    No. PGO (and the G1 GC for native) are Oracle GraalVM features. Community Edition gives you Serial and Epsilon GC and no PGO, so its native images keep the throughput gap.
  • If startup is all you care about, do you need any of this?
    No. Stock native image with Serial GC and -O2 already gives fast startup and low footprint; PGO/G1 matter only when steady-state throughput under sustained load is the goal.

saying these in an interview costs you the question

  • Claiming native images still JIT-compile hot methods at runtime
  • Thinking PGO is available in GraalVM Community Edition
  • Assuming native is always faster than the JVM at steady state

context