skip to content

A hot Java call site dispatches on many different receiver classes. What does HotSpot do at such a megamorphic site, why is it much slower than a site that sees one class, and what can you change?

level: seniorimportance: should knowfreq 40%

answer

  1. Inline cache: unlinked → mono → bi → megamorphic
  2. Megamorphic = vtable index, itable search
  3. Indirect branch mispredicts
  4. Lost inlining dominates the cost
  5. Profile pollution from shared helpers/lambdas

basics

~20 s

With many receiver types the inline cache gives up and the call becomes a table dispatch — a virtual-table index, or a slower interface-table search for interface calls. Nothing is inlined, so all downstream optimization is lost, and the indirect branch mispredicts. Fix by restoring one type per hot call site.

solid answer

~60 s

HotSpot links each call site through an inline cache. It starts unlinked; on first execution it becomes **monomorphic** — a class check plus a direct jump, and in compiled code the target is usually inlined behind that check. Seeing a second dominant type can yield **bimorphic** inlining (two checks, two bodies). When more types keep appearing, the cache goes **megamorphic** and the site is patched to dispatch through a table: a vtable index for virtual calls, or an itable lookup for interface calls, which also has to search for the interface's table in the receiver's class. The table lookup itself costs only a few cycles plus an indirect branch that mispredicts. The real loss is the inlining: no constant folding into the callee, no escape analysis on arguments, no check elimination — often several times the direct dispatch cost. Remedies are structural: specialize the loop or the pipeline so each compiled site sees one type, duplicate a shared helper whose profile is polluted, or hoist the type decision out of the inner loop.

code

java · 8 lines
java
// Megamorphic: one call site sees every Handler implementation
for (Event e : batch) handlerFor(e).handle(e);

// Specialized: each loop's call site sees one class after grouping
for (var entry : batch.groupedByHandler().entrySet()) {
    Handler h = entry.getKey();          // dispatch hoisted out
    for (Event e : entry.getValue()) h.handle(e);
}

go deeper

for a junior

Know that when many different classes go through the same call, the JVM has to look the method up instead of taking a shortcut, and the code gets slower.

for a middle

Name the inline-cache states and say that the main loss is inlining, not the lookup itself.

for a senior

Diagnose it: identify profile pollution from shared helpers or lambdas, and propose specializing the loop or duplicating the helper, verified with PrintInlining and a profile.

for a principal

Design so hot inner loops are monomorphic by construction — dispatch chosen per batch or per pipeline instance — and treat a small amount of deliberate duplication on a proven hot path as an acceptable, documented trade.

## Inline caches and their three states A call site in HotSpot is not a fixed instruction sequence; it is patched as the runtime learns about it. - **Unlinked** — not yet executed; the first call resolves it. - **Monomorphic** — one receiver class seen. The site holds a check against that class and a direct jump to its compiled entry. In optimized code the compiler usually goes further and inlines the body behind the check. This is the fast case and, in most programs, the overwhelmingly common one. - **Bimorphic** — the profile shows two dominant classes. C2 can emit two class checks and inline both bodies, with an uncommon trap for anything else. - **Megamorphic** — many classes have flowed through. Caching one or two targets is now counterproductive (each miss costs a patch or a trap), so the site is switched to a general dispatch. ## What megamorphic dispatch actually does For `invokevirtual`, the receiver's class holds a virtual method table; the compiler knows the slot index, so dispatch is: load class word, load the vtable slot, indirect call. For `invokeinterface` it is more work: interface methods do not have a fixed vtable index across unrelated classes, so the receiver's class carries an interface method table (itable) that must first be searched for the right interface, then indexed. HotSpot mitigates this but interface megamorphic dispatch remains the more expensive of the two. ## Why the slowdown is bigger than the dispatch cost Count only the instructions and megamorphic dispatch looks cheap — a couple of loads and an indirect jump. Two effects multiply it: 1. **Branch misprediction.** An indirect branch that alternates among many targets defeats the branch-target predictor; each mispredict costs a pipeline flush, easily tens of cycles. 2. **Lost inlining, which is the dominant term.** The callee body stays opaque. Arguments must be assumed to escape, so allocations that would have been scalar-replaced are materialised; caller constants cannot fold into callee branches; null and bounds checks cannot be removed; the loop cannot be unrolled or vectorized through the call. The call site becomes an optimization wall in the middle of a hot loop. ## The usual root cause: profile pollution Megamorphism is frequently *not* inherent to the algorithm. A generic helper — a mapper, a validator, a `forEach` body, a logging façade — is invoked from dozens of contexts. Its own internal call site records the union of every context's receiver types, so a caller that only ever passes one type still inherits a hopeless profile. Lambdas and method references make this easy to hit because each lambda body is a distinct class flowing through the same functional-interface call site. ## What to change - **Specialize the hot path.** Choose the implementation once, outside the loop, and run a loop per type. Each compiled loop then sees one class. - **Duplicate the shared helper** on the hot path so it gets its own clean profile; the code duplication is deliberate and should be commented as such. - **Hoist dispatch out of the inner loop** — decide per batch, not per element. - **Reduce the number of implementations reaching that site** at all: merge implementations, or route different kinds through different code paths higher up. - Do **not** expect `final`, `sealed`, or flag tuning to help; they do not change how many classes reach the site. ## Confirming rather than guessing `-XX:+UnlockDiagnosticVMOptions -XX:+PrintInlining` will report the site as not statically bindable; `-XX:+LogCompilation` gives per-site profile data showing the receiver-type spread. A profiler that attributes time to an indirect call inside a hot loop, with no inlined children, is the same signal seen from the outside.

  • Why is a megamorphic interface call typically more expensive than a megamorphic virtual call in HotSpot?
    A virtual call has a fixed vtable index the compiler knows statically, so dispatch is two loads and an indirect jump. An interface method has no such uniform index across unrelated classes, so the receiver's interface method table must be located for that interface before indexing, adding a search step. Both also pay the indirect-branch misprediction.
  • You duplicate a small shared helper method and the hot path gets noticeably faster with no algorithmic change. What happened?
    The original helper's internal call site had a polluted receiver-type profile because many callers with many types shared it. The duplicate is invoked from one context only, so its profile is monomorphic, it devirtualizes and inlines, and the downstream optimizations return. It is a legitimate technique on a proven hot path, and worth a comment explaining why the duplication exists.

A receptionist who always sends visitors to the same office memorises the route; once visitors come for fifty different offices, the receptionist stops memorising and looks every one up in the directory — and the visitor loses the shortcuts the memorised route came with.

saying these in an interview costs you the question

  • Assuming the cost of megamorphism is just the extra indirect jump
  • Trying to fix a megamorphic site with final or sealed
  • Blaming interfaces in general rather than the number of types at one specific site
  • Believing the JIT will eventually 'learn' all the types and inline them all
  • Reaching for flag tuning instead of restructuring dispatch

context