skip to content

Java calls through interfaces and overridable methods are virtual. How can a JIT compiler still inline such a call, and what must be proved or guarded for that to be legal?

level: seniorimportance: must knowfreq 50%

answer

  1. CHA: single implementer ⇒ direct call + dependency
  2. Dependency invalidated by later class loading ⇒ deopt
  3. Type profile per call site ⇒ guarded speculation
  4. Guard = load class word, compare, uncommon trap
  5. Bimorphic = two guards, two bodies

basics

~20 s

Two routes. Proof: class-hierarchy analysis shows only one loaded implementation, so the compiler emits a direct inlined call and records a dependency — loading a second implementation invalidates that compiled code. Speculation: the receiver-type profile shows a dominant type, so the compiler emits a class check, inlines behind it, and traps on mismatch.

solid answer

~50 s

Devirtualization turns a dynamically dispatched call into a direct one the inliner can consume. **Proved (class-hierarchy analysis).** If exactly one implementation of the method is loaded, HotSpot compiles a direct call and registers a dependency on that assumption. If a class loader later defines a second implementer, the JVM invalidates every nmethod holding that dependency and the affected frames deoptimize. Correctness is preserved because the compiler never assumes anything it cannot revoke. **Speculated (type profiling).** The interpreter and the profiling compiler tier record receiver classes per call site in the method data. If one type dominates, C2 emits a guard — compare the receiver's class word with the expected class — and inlines that target on the fast path; a mismatch takes an uncommon trap back to the interpreter. With two dominant types HotSpot can inline both behind a two-way type check (bimorphic inlining). The guard is a load and a compare, cheap enough that the inlined body pays for it many times over.

code

text · 5 lines
text
; receiver in rax, speculated class = ArrayList
mov   r10d, [rax + 8]        ; load compressed class word
cmp   r10d, <ArrayList klass>
jne   uncommon_trap           ; profile said otherwise -> deoptimize
; ---- ArrayList.add body inlined here, fully optimizable ----

go deeper

for a junior

Know the headline: the JVM watches which class actually shows up at a call site and can inline that method behind a quick type check.

for a middle

Distinguish the two routes — a single loaded implementation proved by hierarchy analysis versus a profile-based speculation with a guard — and note both stay correct via deoptimization.

for a senior

Explain dependency invalidation on class loading, the per-call-site profile, bimorphic inlining, and diagnose a slow site as profile pollution rather than blaming the interface.

for a principal

Reason about system shape: keep hot call sites monomorphic by construction (specialized pipelines, per-type paths), and accept interfaces elsewhere; weigh recompilation churn from late class loading in plugin-style architectures.

## Why this is hard Most Java calls are `invokevirtual` or `invokeinterface`. The target depends on the run-time class of the receiver, and Java allows classes to be loaded at any moment, so no compiler can decide the target once and for all by looking at source. Yet without a single known target there is nothing to inline, and without inlining most other optimizations stall. Devirtualization is how a JIT escapes this. ## Route 1: prove it, then take back the proof if needed Class-hierarchy analysis (CHA) asks the runtime: among the classes currently loaded, how many override this method? If the answer is one — an extremely common situation, since most interfaces have a single implementation in a given process — the compiler emits a direct call to that body and inlines it, with no run-time check at all. This is only sound because the assumption is *revocable*. The compiled method (nmethod) records a dependency of the form "no other implementer of this method exists". Class loading consults these dependencies: when a second implementer is defined, the JVM marks every dependent nmethod as not-entrant and deoptimizes any frames currently running it, falling back to the interpreter and later recompiling with a weaker assumption. This is the same machinery that makes speculative optimization safe generally. A closely related, cheaper case is exact-type inference: if the compiler can see the allocation (`List<X> l = new ArrayList<>(); ... l.add(v);` in one compiled region), the receiver's exact class is known outright, no CHA and no guard needed. And calls that are non-virtual in the first place — `invokestatic`, private and `super` calls, constructors — never need any of this. ## Route 2: speculate on the profile, guard the speculation When several implementations exist, the compiler falls back to what the site has actually observed. While a method is interpreted or running in a profiling compiled tier, HotSpot records receiver classes per call site in that method's profile data (a small number of type slots, governed by `TypeProfileWidth`, plus a counter for "something else"). If one type dominates, C2 emits a **monomorphic guard**: load the receiver's class word, compare it to the expected class, and on match run the inlined body; on mismatch execute an uncommon trap that deoptimizes to the interpreter, updates the profile and eventually triggers recompilation. If two types dominate, it can perform **bimorphic inlining**: two checks and two inlined bodies, with a trap for the rest. Beyond that, the site is megamorphic and the compiler stops trying. ## Why the guard is worth it A guard costs a load and a compare-and-branch that predicts perfectly when the speculation holds. In exchange, the callee's body is exposed to constant folding, check elimination and escape analysis. The economics are strongly in favour of guarding, which is why HotSpot speculates aggressively rather than falling back to dispatch. ## What breaks it - A truly polymorphic hot site with many receiver types: no dominant type, no inlining. - Profile pollution: a shared helper whose inner call site is reached from many contexts records the union of all their types, so even callers that are monomorphic in reality inherit a megamorphic profile. - Late class loading that invalidates a CHA-based assumption in steady state, causing recompilation churn — usually harmless, occasionally visible in latency after a plugin or framework loads new classes. ## Design consequences Single-implementation interfaces cost nothing; abstraction is not the enemy. The enemy is many receiver types flowing through *one* hot call site. Splitting a loop per type, specializing a pipeline at construction time so each compiled path sees one class, or duplicating a hot shared helper to give it its own clean profile are the standard remedies. `final` and `sealed` help the compiler's reasoning marginally but do not rescue a site that genuinely sees many types.

  • If class-hierarchy analysis lets the compiler emit a direct call with no guard, what happens when a new subclass is loaded an hour later?
    The compiled code recorded a dependency on "only one implementer". When the class loader defines a second implementer, the JVM invalidates every nmethod carrying that dependency, makes it not-entrant, and deoptimizes any running frames back to the interpreter. Execution stays correct; the affected methods are simply recompiled later with a guarded or dispatching version.
  • Two call sites call the same interface method; one is fast and one is not. What is the usual explanation?
    Different receiver-type profiles. The fast site has seen one dominant class and is inlined behind a cheap guard; the slow site has seen many classes and is dispatched, losing the inlined body and every optimization that depended on it. Profile pollution through a shared helper is the common cause, and specializing or duplicating the path is the usual fix.

saying these in an interview costs you the question

  • Claiming the JVM cannot inline virtual or interface calls at all
  • Thinking final or sealed is required for devirtualization
  • Believing a speculative inline is unsound because a new class could be loaded — it is guarded and revocable
  • Assuming the type check before an inlined body is expensive
  • Confusing devirtualization with source-level static binding done by javac

context