skip to content

Why does a Java application compiled ahead of time into a native binary usually reach lower peak throughput than the same application on a warmed-up JVM, and what narrows that gap?

level: seniorimportance: should knowfreq 30%

answer

  1. JIT has counters; AOT has none
  2. speculation needs a deoptimization fallback
  3. no interpreter in the image, so no bail-out
  4. PGO: instrumented build → run → final build
  5. profiles cannot adapt after deployment

basics

~20 s

A build-time compiler has no runtime profile, so it cannot speculate on the branches and receiver types actually taken, and it has no way to recover if a guess is wrong. Profile-guided optimization from an instrumented run recovers much of the gap.

solid answer

~60 s

A profiling compiler in a running JVM knows things static analysis cannot: which branches are taken, which loops are hot, which receiver types actually appear at a call site, which exceptions never occur. It uses that to devirtualize and inline aggressively, prune cold paths, and specialize code — all guarded by cheap checks. If a guard fails, the code is discarded and execution falls back to the interpreter, then recompiles with the new information. That safety net is what makes aggressive speculation profitable. A build-time compiler must be conservative because it has neither the profile nor the fallback: a megamorphic-looking call stays an indirect call, cold error handling is compiled as if it mattered, and inlining decisions are made from heuristics rather than measurement. Profile-guided optimization closes much of the gap: build an instrumented binary, run a representative workload, feed the recorded profile back into a second build. The costs are a two-phase build, profiles that must stay representative of production, and a still-missing ability to adapt to a workload that changes shape after deployment.

go deeper

for a junior

Say the runtime compiler can watch the program run and optimize for what it actually does, while a build-time compiler cannot.

for a middle

Name specific profile-driven optimizations such as devirtualization and hot-path inlining, and mention profile-guided optimization as the mitigation.

for a senior

Add the deoptimization argument, and reason about which workloads ever reach peak throughput in the first place.

for a principal

Weigh the throughput ceiling against fleet-level metrics — instance count, startup, memory cost — and against the delivery complexity a two-stage profiled build imposes.

## The information asymmetry Optimizing compilers make decisions that are only good if certain facts hold. Inlining is worth it if the callee is small and actually called often. Devirtualizing an interface call into a direct call is worth it if one implementation dominates. Skipping a bounds check is valid if the index is provably in range. Laying out code so the hot path falls through is worth it if you know which path is hot. A JVM's optimizing compiler runs after the program has been executing, and it has counters: per-branch taken and not-taken counts, per-call-site observed receiver types, per-loop trip counts, whether a given exception path ever executed, whether a null was ever seen. It optimizes against those facts. A build-time compiler for a native binary has the source and the whole reachable program, but not one bit of runtime behaviour. It can prove things (this call has exactly one possible implementation in the closed world) but it cannot measure things. ## Why the fallback matters as much as the profile The subtler half of the advantage is deoptimization. A JVM can compile code that is only correct under an assumption — for example that a call site only ever sees one receiver class, or that a class is never subclassed — provided it inserts a check and can bail out. When the assumption breaks, the frame is converted back to interpreter state and execution continues correctly, then the method is recompiled without the invalid assumption. A native binary has no interpreter and no compiler, so there is nothing to fall back to. Every optimization must be unconditionally valid or guarded by a check that is always paid. This makes speculative optimization structurally unavailable rather than merely uninformed, which is why simply feeding a profile in does not make the two models equivalent. ## The concrete gaps - **Call sites** stay indirect where a JVM would inline the single hot receiver and guard it. - **Cold code** — error handling, rarely-taken branches — is compiled and laid out as ordinary code rather than being pruned or moved out of the hot instruction stream. - **Inlining budget** is spent on heuristics rather than on the calls that actually dominate. - **Class-hierarchy assumptions** are safer in a closed world (the whole hierarchy is known) which does claw some of this back — a genuinely single-implementation interface can be devirtualized with certainty at build time. - **Garbage collection** in the image may be a simpler collector than a fully-tuned HotSpot collector, which affects sustained throughput independently of code quality. ## Profile-guided optimization The standard remedy is to give the build-time compiler the missing measurements. The flow is: build an instrumented binary, run it against a workload that resembles production, collect the recorded profile, then build the final binary with that profile. The compiler then places branches, chooses inlining, and specializes code using real frequencies. Reported gains commonly recover a large share of the difference to a warmed JVM on throughput-oriented benchmarks. Its limits should be stated honestly. It adds a two-stage build with a workload run in the middle, which complicates continuous delivery. The profile is only as good as the workload used, and a profile taken from an unrepresentative load can misdirect optimization. And because there is still no deoptimization, the binary cannot adapt when production traffic changes shape — a JVM would simply reprofile and recompile. ## The judgment to express If the process lives long enough to warm up and throughput is the objective, the JVM's adaptive model has a structural advantage and you should expect to keep it. If processes are short-lived, numerous, or memory-constrained, the throughput ceiling is largely irrelevant because the JVM would rarely reach it, and native compilation wins on the metrics that actually bind. Profile-guided optimization is the tool for the middle ground: a long-running service that nevertheless needs fast startup and small footprint.

  • If the closed world is fully known at build time, why can the compiler not just devirtualize everything?
    It can devirtualize a call site that provably has one implementation in the whole program, and that is a real advantage of closed-world analysis. But many call sites have several possible implementations even though only one is common at runtime. Without a profile the compiler cannot know which, and without deoptimization it cannot bet on one and recover if it is wrong, so it emits an indirect call.
  • What operational cost does profile-guided optimization add to a delivery pipeline?
    The build becomes two-stage with a workload run in between, so build time roughly doubles and the pipeline needs a repeatable, representative load generator. The profile becomes an artifact that must be versioned and refreshed as the application changes, otherwise it slowly drifts away from actual behaviour and its guidance degrades.

saying these in an interview costs you the question

  • Claiming ahead-of-time compilation is slower because machine code is generated by a weaker compiler — it is the same optimizing compiler technology without profiles or a fallback.
  • Forgetting deoptimization and attributing the entire gap to missing profiles.
  • Assuming profile-guided optimization fully eliminates the difference in all workloads.
  • Comparing native startup performance against warmed-up JVM throughput as if they were the same metric.

context