skip to content

Why can a warmed-up JVM out-perform a GraalVM native image on peak throughput, and what mitigates that gap?

level: middleimportance: must knowfreq 60%

answer

  1. JIT C2 uses runtime profile → speculate + deoptimize
  2. AOT frozen at build time, no live profile
  3. warm-up cost vs steady-state win
  4. PGO: instrument → collect → rebuild (Oracle GraalVM)
  5. G1 vs Serial GC for throughput

basics

~20 s

The JVM's JIT compiler profiles running code and re-optimizes hot paths, so after warm-up it can run faster. A native image is compiled once ahead-of-time with no runtime profiling, so its peak throughput is often lower. Profile-Guided Optimization (PGO) narrows the gap.

solid answer

~50 s

On the JVM, the JIT (C2) observes runtime behavior — which branches are taken, which types actually appear, which methods are hot — and aggressively optimizes and re-optimizes machine code using that live profile. A native image is AOT-compiled once at build time with no runtime profile, so it can't specialize to the real workload; peak (steady-state) throughput is therefore typically lower than a fully warmed JVM, even though startup is far faster. The main mitigation is **Profile-Guided Optimization (PGO)** in Oracle GraalVM: you run an instrumented build against representative load, collect a profile, then rebuild using it — recovering much of the throughput. Choosing a throughput-oriented GC (G1 instead of the default Serial GC) also helps. The strategic point: native trades peak throughput for startup and footprint, so it suits short-lived or scale-heavy workloads more than long-running max-throughput services.

code

kotlin · 16 lines
kotlin
// PGO is a two-phase native build (Oracle GraalVM only).
// Phase 1 — instrumented build, then run under representative load to emit default.iprof:
//   ./gradlew nativeCompile -Pinstrumented   (adds --pgo-instrument)
//   ./build/native/nativeCompile/app          (drive real traffic, produces default.iprof)
// Phase 2 — optimized rebuild using the collected profile:
//   ./gradlew nativeCompile -Ppgo             (adds --pgo=default.iprof)

// Configure the flags in build.gradle.kts:
graalvmNative {
    binaries {
        named("main") {
            // e.g. throughput-oriented GC (Oracle GraalVM): buildArgs.add("--gc=G1")
            // PGO flags wired via project properties as above.
        }
    }
}

go deeper

for a junior

Know that the JVM 'warms up' via JIT and can get faster over time; native is fixed at build time.

for a middle

Explain C2's use of a runtime profile and name PGO as the mitigation; know it's Oracle GraalVM only.

for a senior

Reason about workload shape (CPU- vs I/O-bound, short- vs long-lived) and edition/GC choices when deciding.

for a principal

Weigh throughput loss against fleet cost and cold-start SLAs; factor Oracle vs Community edition licensing into the call.

## The JIT advantage The HotSpot JVM runs bytecode through tiered compilation: 1. first interpreted, 2. then compiled by C1 (fast, light optimization), 3. and the hottest code by **C2** (heavy optimization). Crucially, C2 uses a *runtime profile*: - counts of how often each branch is taken, - which concrete classes actually flow through a virtual call site (for inlining and devirtualization), - which loops dominate. It can make *speculative* optimizations (e.g., assume a call is monomorphic and inline it) and **deoptimize** back to the interpreter if the assumption breaks. After 'warm-up' — the seconds/minutes it takes to gather profiles and compile — a JVM reaches high steady-state throughput tuned to the *actual* workload. ## Why native gives up some peak throughput A GraalVM native image is compiled **once, ahead of time (AOT)**, with a **closed-world** view but no knowledge of real runtime frequencies. It can do static/heuristic optimization but cannot observe the live profile, cannot speculate-and-deoptimize, and its code quality is frozen at build time. So for CPU-bound, long-running hot loops, a warmed-up JVM often wins on requests/sec. (For I/O-bound services the difference is frequently small because the CPU isn't the bottleneck.) ## Mitigation 1 — Profile-Guided Optimization (PGO) Available in **Oracle GraalVM** (not the free GraalVM Community Edition). Workflow: 1. build an *instrumented* executable (`--pgo-instrument`), 2. run it under representative traffic to produce an `.iprof` profile file, 3. then rebuild passing that profile (`--pgo`). The AOT compiler now optimizes using real frequencies — inlining hot paths, laying out branches — and typically recovers a large fraction of the JVM's throughput advantage. The cost is a two-phase build and the need for a realistic load profile. ## Mitigation 2 — GC choice The native default is **Serial GC** (single-threaded, low footprint) which can cap throughput under allocation-heavy load. Oracle GraalVM offers **G1** (`--gc=G1`), a parallel/concurrent collector that improves throughput and pause behavior at the cost of more memory. ## Gotchas - (1) Don't benchmark cold — compare native against a *warmed-up* JVM to see the real steady-state gap, and against a *cold* JVM to see native's startup win. - (2) PGO profiles must resemble production load; a bad profile helps little. - (3) Community Edition lacks PGO and G1, so its throughput ceiling is lower — an important licensing/edition decision. ## When it matters Peak-throughput loss is acceptable — even irrelevant — for serverless functions, CLIs, and bursty autoscaled fleets where each instance is short-lived or I/O-bound. It matters most for steady, CPU-bound, long-running services, where a JVM's warm-up cost is amortized over days.

  • Is native always slower than the JVM at steady state?
    No. For I/O-bound services the CPU isn't the bottleneck, so throughput is often comparable. And with PGO + G1 on Oracle GraalVM the gap narrows substantially. The clear JVM win is CPU-bound, long-running hot loops.
  • What's the catch with PGO?
    It's Oracle GraalVM only (not Community Edition), requires a two-phase build, and needs a representative load profile — an unrepresentative profile yields little benefit.
  • Why can the JVM 'speculate' but native can't?
    The JVM can deoptimize — fall back to the interpreter if a speculative assumption (e.g., a monomorphic call) proves wrong. AOT code is fixed at build time with no interpreter to fall back to, so it can't take those risks.

saying these in an interview costs you the question

  • Claiming native is always faster than the JVM
  • Thinking there's a runtime JIT in native images
  • Assuming PGO is free/Community Edition
  • Confusing startup speed with peak throughput

context