You migrated a load-bearing Spring service to native image for startup and footprint, but steady-state throughput regressed versus the JVM. As the tech lead, how do you decide whether and how to close the gap?
answer
- First ask: is throughput actually binding?
- Levers by payoff: G1 -> -O3 -> PGO
- All need Oracle GraalVM (licensing)
- PGO doubles CI + needs load stage
- Willing to route back to JVM
basics
~20 sFirst confirm throughput actually matters here. If it does, apply the levers in order of payoff: G1 GC, -O3, and PGO — but they require Oracle GraalVM (licensing) and heavier CI. If the cost outweighs the benefit, keep this service on the JVM.
solid answer
~50 sI'd treat it as a cost/benefit decision, not a reflex. Step one: quantify — is steady-state throughput actually a binding constraint for this service, or did we go native purely for startup/footprint (serverless, scale-to-zero)? If throughput isn't binding, do nothing. If it is, apply the levers: switch the default Serial GC to G1 (--gc=G1) to remove single-threaded GC as a ceiling; raise optimization to -O3; and invest in PGO — instrument, replay representative load to capture default.iprof, rebuild with --pgo. These recover most of the JIT gap. But they all require Oracle GraalVM, which means licensing cost, and PGO roughly doubles an already-slow native build plus adds a load-test stage in CI, and G1 raises footprint. So I weigh the recovered throughput against Oracle licensing, CI complexity, and lost footprint. If the balance is poor, the honest answer may be to keep this particular service on the JVM and reserve native for the workloads where startup/footprint dominate.
go deeper
Understand there are levers (GC, -O, PGO) but the lead's job is deciding whether to use them.
Name the levers and that they need Oracle GraalVM.
Sequence the levers by payoff and articulate the licensing/CI/footprint tradeoffs.
Drive it as a per-service cost/benefit decision with governance (build profiles, native-vs-JVM classification) and the willingness to keep throughput-bound services on the JVM.
## Frame it as a decision, not a checklist The worst move is to reflexively pile on tuning flags. The throughput gap exists because native image is **AOT-compiled with no JIT** — it can't re-optimize hot code from runtime profiles. Whether that matters depends entirely on the service. ### Step 1 — Is throughput even the binding constraint? Ask *why* you went native. If it was **startup latency** (serverless, scale-to-zero, rapid autoscaling) or **memory footprint** (dense packing, small containers), and the service's throughput SLO is comfortably met, then the regression is irrelevant — **do nothing**. Native was the right call for its actual reason. If throughput *is* a binding SLO under sustained load, proceed — but measure the actual gap first with representative load testing, don't guess. ### Step 2 — Apply levers in order of payoff / cost 1. **GC: Serial -> G1** (`--gc=G1`). The default **Serial GC** is single-threaded — under sustained allocation it's often the first ceiling. **G1** parallelizes and mostly-concurrently collects, scaling with cores. Relatively cheap to adopt; costs footprint. *(Oracle GraalVM, Linux.)* 2. **Optimization: `-O3`.** More aggressive AOT optimization than the `-O2` default. Costs build time. *(Oracle GraalVM.)* 3. **PGO** — the biggest lever and the most work: build with `--pgo-instrument`, replay a **representative** workload to produce `default.iprof`, rebuild with `--pgo`. Recovers most of the JIT gap by giving the compiler real hot-path/branch/type data. Costs a doubled build plus a load-test stage in CI, and is only as good as the workload you replay. These stack: G1 addresses GC overhead, `-O3` + PGO address code quality. ### Step 3 — Weigh the true cost - **Licensing:** PGO, G1, and `-O3` are **Oracle GraalVM** features. Community Edition can't do them. Oracle GraalVM under production carries licensing considerations — a real budget line, not a footnote. - **CI cost & complexity:** native builds are already the slowest CI stage; PGO roughly **doubles** build time and adds a representative-load stage that must be maintained and kept representative as traffic evolves. - **Footprint erosion:** G1 raises memory use, partly undoing the footprint win that may have justified native in the first place. - **Profile staleness:** a PGO profile drifts as traffic changes; someone owns re-capturing it. ### Step 4 — Decide, and be willing to say no If recovered throughput clears the SLO and the Oracle licensing + CI + footprint costs are acceptable, ship G1 + `-O3` + PGO with a maintained profiling pipeline. If not, the mature answer is often **route this service back to the JVM** (where the JIT gives peak throughput for free) and reserve native image for the workloads where **startup and footprint** genuinely dominate. Native vs JVM is a per-service decision, not a platform mandate. ### Governance angle Establish **build profiles** (dev uses `-Ob` for fast iteration; release uses `-O3` + PGO) and a **classification** of which services are native (startup/footprint-driven) versus JVM (throughput-driven), so the choice is deliberate and documented rather than ad hoc.
- When is 'do nothing' the right answer to the throughput regression?When the service went native for startup/footprint (serverless, scale-to-zero, dense packing) and its throughput SLO is still met. The regression is irrelevant if throughput isn't the binding constraint.
- What's the hidden ongoing cost of PGO beyond the initial build?The profile drifts as traffic evolves; someone must own re-capturing a representative default.iprof, and CI must maintain the load-test stage. It's a maintained pipeline, not a one-time flag.
- Why might keeping a service on the JVM be the mature choice?The JIT delivers peak steady-state throughput for free with no Oracle licensing, no doubled builds, and no footprint tradeoff. If native's startup/footprint benefits don't apply to this service, the JVM is simply the better fit.
saying these in an interview costs you the question
- Reflexively tuning when throughput isn't the binding constraint
- Ignoring Oracle GraalVM licensing cost of PGO/G1/-O3
- Treating native-vs-JVM as a platform mandate rather than a per-service decision
- Forgetting PGO profiles need ongoing maintenance