skip to content

Method Inlining & Devirtualization

Inlining replaces a call with the callee's body and unlocks most other optimizations, bounded by method size, call-site hotness, and depth; devirtualization plus monomorphic, bimorphic, and megamorphic inline caches decide whether a virtual call can be inlined at all. Interviewers use it to test why "virtual calls are slow" is mostly false in practice.

on this pageshow

questions

5

When a JIT compiler inlines a method call, what happens to the generated machine code, and why is inlining called the optimization that unlocks most of the others?

level: middleimportance: must knowfreq 60%

answer

  1. Call replaced by callee body
  2. Barrier removal, not cycle saving
  3. Enables folding, escape analysis, hoisting
  4. Compilation unit = inlining tree
  5. Cost: compile time, code size, deopt

basics

~20 s

Inlining replaces a call with a copy of the callee's body inside the caller. Removing call overhead is the small win; the real win is that both bodies become one compilation unit, so constant folding, null and bounds-check removal, escape analysis and loop optimizations can work across the boundary that disappeared.

solid answer

~50 s

Inlining pastes the callee's intermediate representation into the caller at the call site: arguments become ordinary values, the return becomes a value, and the frame setup, argument marshalling, call instruction and return jump all vanish. The direct saving is small — a well-predicted direct call costs a few cycles. What matters is that the callee's code now sits in the caller's context. The compiler can propagate caller constants into callee branches and fold them, delete null checks and array-bounds checks the caller already proved, hoist callee field loads out of the caller's loop, and run escape analysis on an object that previously escaped merely by being passed as an argument — enabling scalar replacement and lock elision. So the real unit of optimization in HotSpot is not a method but an inlining tree: the hot root method plus everything pulled into it. The costs are compile time, code size, instruction-cache pressure and more code to discard on deoptimization, which is why inlining is bounded by heuristics.

code

text · 6 lines
text
java -XX:+UnlockDiagnosticVMOptions -XX:+PrintInlining -XX:+PrintCompilation App

@ 7   Point::getX (5 bytes)   accessor
@ 12  Point::distance (28 bytes)   inline (hot)
@ 21  Formatter::render (612 bytes)   too big
@ 33  List::add (0 bytes)   no static binding

go deeper

for a junior

Say what it is: the JIT copies the small method's body into the caller so there is no call, and it happens automatically once the code is hot.

for a middle

Explain that the call was an optimization barrier, and name two things inlining enables — constant folding into the callee and escape analysis on arguments.

for a senior

Frame the compilation unit as an inlining tree, connect inlining to scalar replacement and check elimination, and mention reading PrintInlining to verify an abstraction collapsed.

for a principal

Discuss the tradeoff curve: bigger trees buy optimization but cost compile time, code cache, icache and deopt payload; treat inlining budget as a resource the design of hot paths should respect.

## The transformation A call site such as `int y = box.compute(x);` compiles to an invoke bytecode. When the JIT inlines it, no call instruction is emitted at all: the compiler copies the callee's intermediate representation into the caller, substitutes the actual argument values for the callee's parameters, and replaces the callee's return with a value the caller consumes. Everything the calling convention required — spilling caller-saved registers, laying arguments into registers or stack slots, pushing a frame, jumping, returning — is gone. ## Why the direct saving is the small part On modern hardware a direct, well-predicted call plus return is only a handful of cycles. If that were the whole benefit, inlining would be a minor peephole trick. It is not, because a call is also an *optimization barrier*: on either side of it the compiler must assume the worst about values it cannot see through. ## What inlining actually unlocks - **Constant propagation and folding.** If the caller passes a literal or a value the compiler has proved constant, that constant flows into the callee's branches; branches fold, and the untaken side is deleted as dead code. A generic, flag-driven helper collapses into the one path this caller uses. - **Redundant-check elimination.** The caller may already have null-checked a reference or bounds-checked an index. Once the callee's checks are in the same unit, the compiler sees them as redundant and removes them. - **Escape analysis.** An object passed as an argument to an opaque call must be assumed to escape. After inlining, the compiler can often prove it never leaves the compiled region, and then scalar-replace it — the allocation disappears entirely and the fields live in registers — or elide a lock on it. - **Loop optimizations.** Field loads and invariant computations inside the callee can be hoisted out of a caller loop, unrolled with it, or vectorized with it, none of which is possible while the body is behind a call. - **Register allocation** happens over the merged region, so values stop being spilled at every boundary. ## The compilation unit is an inlining tree Because of this, HotSpot's C2 does not compile a method in isolation; it compiles a root method together with the transitive callees it decided to inline. The compiled code (an *nmethod*) therefore represents many source methods, and its metadata records which frames were folded in so a deoptimization can rebuild the interpreter frames that never physically existed. ## The costs, and why heuristics exist A bigger inlining tree means better optimization but longer compile time, larger machine code (instruction-cache and branch-predictor pressure), and more work thrown away when a speculative assumption fails. So the compiler bounds inlining by callee size, call-site hotness and inlining depth, and refuses call sites whose target it cannot pin down to one method. ## Observing it `-XX:+UnlockDiagnosticVMOptions -XX:+PrintInlining` prints the decision tree with reasons (`inline (hot)`, `too big`, `not compilable`, `no static binding`), which is how you confirm that the abstraction you care about actually collapsed.

  • If inlining removes so little call overhead, why do micro-benchmarks that call an empty method in a loop still show inlining mattering enormously?
    Because the benchmark is usually measuring the barrier effect, not the call. Once the empty callee is inlined the whole loop body becomes provably dead and the loop is deleted, so the measurement collapses to nothing. That is also the classic reason such benchmarks need a consumed result or a blackhole to stay honest.
  • What does inlining cost the runtime, and where does that cost show up?
    It costs JIT compile time and code cache space, and it increases instruction-cache and branch-predictor pressure when the compiler over-inlines. It also enlarges the deoptimization payload: an nmethod containing many inlined frames must be able to rebuild all of them, and invalidating it throws away more optimized code.

Reading a recipe that says "prepare the sauce as on page 40" versus having the sauce steps written inline: only when the steps are in front of you can you notice you already have the pan hot and skip two of them.

saying these in an interview costs you the question

  • Saying inlining is only about avoiding the cost of the call/return instructions
  • Claiming javac inlines methods — the source compiler does not; the JIT does at run time
  • Believing every small method is always inlined, regardless of hotness or dispatch
  • Assuming inlining is free and that more inlining is monotonically better

context

open as a page

Java calls through interfaces and overridable methods are virtual. How can a JIT compiler still inline such a call, and what must be proved or guarded for that to be legal?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Two routes. Proof: class-hierarchy analysis shows only one loaded implementation, so the compiler emits a direct inlined call and records a dependency — loading a second implementation invalidates that compiled code. Speculation: the receiver-type profile shows a dominant type, so the compiler emits a class check, inlines behind it, and traps on mismatch.

open as a page

Does wrapping a Java field in a getter method make code slower at run time? Explain what the JVM does with a one-line accessor in hot code.

level: juniorimportance: should knowfreq 42%

basics

~20 s

In the interpreter the getter really is a call. Once the code is hot and JIT-compiled, a tiny accessor is inlined and the call disappears, leaving just the field load. Caveats: it only applies after warm-up, and only if the JIT can pin the call to one target.

open as a page

HotSpot does not inline every call. Which properties of the call site and of the target method decide the outcome, and what practical size and depth limits apply?

level: middleimportance: should knowfreq 45%

basics

~20 s

Three families of limits: callee size (tiny methods almost always, small ones if the site is hot, big ones never), call-site hotness from the collected invocation profile, and inlining depth plus total expansion budget. On top of that, the target must be statically bindable and compilable.

open as a page

A hot Java call site dispatches on many different receiver classes. What does HotSpot do at such a megamorphic site, why is it much slower than a site that sees one class, and what can you change?

level: seniorimportance: should knowfreq 40%

basics

~20 s

With many receiver types the inline cache gives up and the call becomes a table dispatch — a virtual-table index, or a slower interface-table search for interface calls. Nothing is inlined, so all downstream optimization is lost, and the indirect branch mispredicts. Fix by restoring one type per hot call site.

open as a page