skip to content

Native Throughput Tuning

Profile-guided optimization, higher optimization levels and switching the native image from serial to G1 close much of the no-JIT throughput gap. Interviewers ask how you would answer a benchmark showing native is slower under sustained load.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

5

Walk through the Profile-Guided Optimization (PGO) workflow for a Spring Boot native image. What are the steps and what file is produced?

level: middleimportance: must knowfreq 50%

answer

  1. instrument -> run -> optimize
  2. default.iprof file
  3. --pgo-instrument then --pgo=
  4. representative workload
  5. Oracle GraalVM only

basics

~10 s

Three steps: build an instrumented binary with --pgo-instrument, run it under a realistic workload to produce a profile file (default.iprof), then rebuild passing --pgo=default.iprof so the compiler optimizes the hot paths. Requires Oracle GraalVM.

solid answer

~50 s

PGO is a two-build process on Oracle GraalVM. First, build an instrumented native image by adding the build arg --pgo-instrument; this binary records execution counters. Second, run that instrumented binary against a representative workload (a load test or replayed production traffic); on exit it writes a profile file, default.iprof, in the working directory. Third, do the real production build passing --pgo=<path-to>/default.iprof (bare --pgo picks up default.iprof from the CWD). The AOT compiler now uses the recorded branch frequencies, hot-method data and type profiles to inline and lay out code the way it's actually exercised, recovering most of the throughput the JVM's JIT would have found at runtime. The critical caveat: the profile only reflects the workload you replayed, so an unrepresentative run yields misleading optimizations. In Spring Boot you pass these via the Native Build Tools buildArgs.

code

kotlin · 29 lines
kotlin
// build.gradle.kts — GraalVM Native Build Tools
// Toggle instrument vs. optimize via a project property so CI can do two builds.
import org.graalvm.buildtools.gradle.dsl.GraalVMExtension

plugins {
    id("org.springframework.boot") version "3.4.0"
    id("org.graalvm.buildtools.native") version "0.10.3"
}

configure<GraalVMExtension> {
    binaries {
        named("main") {
            // Step 1 (instrument):  -Ppgo=instrument  -> adds --pgo-instrument
            // Step 3 (optimize):    -Ppgo=optimize    -> adds --pgo=default.iprof + -O3
            when (project.findProperty("pgo")) {
                "instrument" -> buildArgs.add("--pgo-instrument")
                "optimize"   -> {
                    buildArgs.add("--pgo=${'$'}{rootDir}/default.iprof")
                    buildArgs.add("-O3")
                }
            }
        }
    }
}
// CI:
//   ./gradlew nativeCompile -Ppgo=instrument
//   ./build/native/nativeCompile/app   # run load test -> writes default.iprof
//   cp default.iprof .
//   ./gradlew nativeCompile -Ppgo=optimize

go deeper

for a junior

Memorize the three-step order and the default.iprof filename.

for a middle

Explain each step, the instrument-vs-optimize builds, and that the profiling workload must be representative.

for a senior

Discuss CI orchestration (two builds + a load stage), profile merging, and Oracle GraalVM licensing.

for a principal

Weigh the doubled build time and load-test infrastructure against the throughput gain; decide whether PGO earns its keep for this service.

## What PGO is **Profile-Guided Optimization (PGO)** lets the ahead-of-time `native-image` compiler use a *recorded runtime profile* to make the same kind of profile-driven decisions the HotSpot JIT makes dynamically. It is an **Oracle GraalVM** feature (not in Community Edition). ## The three steps **Step 1 — Instrument build.** Build a native image with the build arg `--pgo-instrument`. The resulting binary is larger and slower because it carries instrumentation that counts how often each branch/method runs and what types flow through call sites. **Step 2 — Profiling run.** Run the instrumented binary against a **representative workload** — ideally the same request mix, payload sizes and concurrency you expect in production (a load test, or replayed production traffic). When the process exits cleanly it writes a profile file named **`default.iprof`** into the current working directory. (You can point it elsewhere; you can also merge multiple profiles.) **Step 3 — Optimizing build.** Do the real build passing `--pgo=<path>/default.iprof` (or just `--pgo`, which looks for `default.iprof` in the CWD). The compiler consumes the profile and produces a binary optimized for the hot paths — better inlining, code layout that keeps hot code together, and profile-driven branch decisions. ## Wiring it into Spring Boot Spring Boot native builds go through the **GraalVM Native Build Tools** plugin (`org.graalvm.buildtools.native`). You inject the build args through `buildArgs` (Gradle) or `<buildArgs>` (Maven `native-maven-plugin`). Because you need *two different builds* (instrument, then optimize), teams typically parameterize the build or run two Gradle/Maven invocations with different args, capturing `default.iprof` between them in CI. ## Edge cases & gotchas - **Representativeness is everything.** The profile is a snapshot of one workload. Optimize for the wrong traffic shape and you can *lose* throughput on real traffic. - **The instrumented binary is not for production** — it's slower and bigger; only its profile output matters. - **The profiling run must exercise the hot paths** you care about, and exit cleanly so the profile is flushed. - **CI cost**: PGO doubles native build time (already the slow part) and adds a load-test stage. Budget for it. - **Oracle GraalVM required** — verify the distribution; on Community Edition `--pgo-instrument`/`--pgo` are unavailable. - PGO stacks with `-O` level and GC choice; it does not replace them.

  • What file does the instrumented run produce and where?
    default.iprof, written to the current working directory when the instrumented binary exits cleanly. You feed it back with --pgo=<path>/default.iprof (or bare --pgo to read it from the CWD).
  • What's the single biggest risk with PGO?
    An unrepresentative profiling workload. The compiler optimizes for whatever traffic you replayed, so a wrong or thin workload can misplace optimizations and even reduce real-world throughput.
  • How does this fit a Spring Boot build?
    Via the org.graalvm.buildtools.native plugin's buildArgs (Gradle) / <buildArgs> (Maven). You run two builds — instrument, then optimize — capturing default.iprof in between, usually scripted in CI.

saying these in an interview costs you the question

  • Believing native-image self-profiles at build time without a runtime run
  • Thinking the JVM/JFR profile feeds native PGO
  • Shipping the instrumented binary to production

context

open as a page

Why does a GraalVM native-image Spring Boot app often have lower peak throughput than the same app on the JVM, and what is the headline lever to close that gap?

level: juniorimportance: should knowfreq 45%

basics

~20 s

A native image is compiled ahead-of-time and has no JIT, so it can't re-optimize hot code from runtime profiles. Peak throughput lags the JVM. Profile-Guided Optimization (PGO) feeds a runtime profile back into the build to close most of the gap.

open as a page

What do the native-image optimization levels (-O0 through -O3, and -Ob) mean, and which would you use for a production throughput-bound Spring service versus local development?

level: middleimportance: should knowfreq 38%

basics

~20 s

-O sets compiler optimization: -O0 none (debugging), -O2 the production default, -O3 more aggressive (Oracle GraalVM), -O1 in between. -Ob optimizes build time, not runtime — use it for fast local dev builds, never for production throughput.

open as a page

A native-image Spring service is throughput-bound under sustained load. Why is the default GC a problem, and what would you change?

level: seniorimportance: should knowfreq 40%

basics

~10 s

Native images default to the Serial GC: single-threaded, low footprint, but stop-the-world pauses that hurt throughput on larger heaps under load. For a throughput-bound service on Oracle GraalVM (Linux), switch to G1 with --gc=G1.

open as a page

You migrated a load-bearing Spring service to native image for startup and footprint, but steady-state throughput regressed versus the JVM. As the tech lead, how do you decide whether and how to close the gap?

level: principalimportance: should knowfreq 30%

basics

~20 s

First confirm throughput actually matters here. If it does, apply the levers in order of payoff: G1 GC, -O3, and PGO — but they require Oracle GraalVM (licensing) and heavier CI. If the cost outweighs the benefit, keep this service on the JVM.

open as a page