Selecting a method from the receiver's runtime type costs more than calling a fixed function. Compare how these four runtimes claw that cost back, and what each gives up: Rust's choice between monomorphised generics and dyn Trait objects; a C++ compiler that can devirtualise only under link-time optimisation or an explicit final marker; the HotSpot JVM's inline caches and speculative devirtualisation backed by deoptimisation; and the per-class method caches in Objective-C's objc_msgSend and in Ruby.
answer
- cost = lost inlining, not the jump
- monomorphise vs dyn Trait = speed vs code size
- C++ needs final / exact type / LTO to prove it
- mono to bi to megamorphic inline cache, then deopt
- runtime class mutation flushes method caches
basics
~20 sThe indirect jump is cheap; losing inlining is the real cost. Rust chooses per call site: monomorphise (fast, code bloat) or dyn Trait (indirect, one copy). C++ devirtualises only when it can prove the exact type. HotSpot speculates from profiles and deoptimises when wrong. Objective-C and Ruby cache lookups and flush on class mutation.
solid answer
~60 sThe indirect branch is nearly free on a modern CPU. What you lose is **inlining**, and everything inlining unlocks: constant propagation, escape analysis, loop transforms. Four strategies: - **Rust** makes it a source-level choice. A generic parameter is monomorphised: one specialised copy per concrete type, direct and inlinable, paid for in code size, instruction-cache pressure and compile time. `dyn Trait` is a fat pointer (data + vtable), one copy of the code, indirect call. - **C++** must *prove* exactness. Another translation unit may hold an override, so devirtualisation needs `final`, a provably exact object, or link-time optimisation; compilers also devirtualise speculatively behind a type guard. - **HotSpot** guesses instead of proving. Class-hierarchy analysis plus a per-call-site inline cache inlines the one or two types actually seen; when a second implementation is later loaded the assumption is invalidated and live frames are deoptimised and recompiled. - **Objective-C and Ruby** cannot prove anything, because classes are mutable at runtime. They cache the lookup per class or per call site and invalidate on any method (re)definition. So final-by-default in Kotlin or C# is a gift to an AOT compiler and near-redundant for a profiling JIT.
code
text · 10 linescall site: shape.area()
t0 receiver Circle only
MONOMORPHIC if (recv.type == Circle) { inlined Circle.area } else -> slow path
t1 Square starts arriving
BIMORPHIC two guarded branches, both may still inline
t2 Triangle, Hexagon, Blob, ...
MEGAMORPHIC load table from recv.type; indirect call; nothing inlinedgo deeper
Know that calling through an abstraction means the implementation is picked from the object's runtime type via a table, that this is what lets new subtypes work at old call sites, and that it is slightly more expensive than a fixed call.
Be able to say that the dominant cost is the lost inlining rather than the jump, and that compilers and runtimes try to recover it by proving or guessing the single likely target.
Compare the strategies concretely and diagnose: recognise a megamorphic site, know that loading a second implementation can trigger deoptimisation, know why a microbenchmark overstates dispatch speed, and know that runtime class mutation invalidates method caches.
Frame it as a platform choice. Decide whether the language should let programmers pick per call site (Rust), require build-wide whole-program analysis (C++ LTO), or buy adaptivity with a deoptimisation mechanism (HotSpot), and weigh sealing hierarchies for AOT and start-up predictability against keeping extension points open.
## What a virtual call actually costs A call resolved from the receiver's runtime type usually compiles to: load a pointer to the type's method table from the object, load the slot for this method, jump to it. On a modern out-of-order CPU with a branch target buffer, a call site that keeps hitting the same target predicts well and the extra loads are hidden. The cost that matters is **the optimisations you never get**. A direct call can be inlined; once inlined, the compiler sees the callee's body in the caller's context and can fold constants, prove an object never escapes and stack-allocate or eliminate it, unroll a loop, or delete a branch. An unresolved call is an opaque wall: arguments must be materialised, registers spilled, the heap assumed clobbered. Measured naively the indirect jump looks like a couple of cycles; measured in a hot loop, the missing inlining can be an order of magnitude. Every technique below exists to recover inlining, not to shave the jump. A second cost appears only at **megamorphic** sites (many receiver types through one call). There the indirect branch stops predicting and you eat mispredictions plus method-table cache misses. ## Rust: the choice is in the source Rust splits the decision and hands it to the programmer. Writing `fn draw<T: Shape>(s: &T)` monomorphises: the compiler emits a separate copy of `draw` for every concrete `T` used, each with a direct, inlinable call. Writing `fn draw(s: &dyn Shape)` builds a fat pointer (data pointer plus vtable pointer) and does one indirect call from one copy of the code. Monomorphisation is not free: binary size grows with the number of instantiations, compile times grow with it, and a large instantiated blob can thrash the instruction cache badly enough that the `dyn` version wins. This is why the choice is a real engineering decision, not a default. ## C++: devirtualisation as a proof obligation C++ has no runtime to observe behaviour, so the compiler may only devirtualise when it can *prove* the target. It can when the object's exact type is visible (`Circle c; c.area();`), when the class or the method is marked `final`, or when it sees the whole program. Without link-time optimisation, each translation unit must assume some other unit defines an override, so an ordinary virtual call through a base pointer stays virtual. With LTO or whole-program visibility the compiler can see the closed hierarchy, and GCC and Clang additionally do *speculative* devirtualisation: emit a guarded direct call to the likeliest target with a fallback to the indirect call. Note the divergence from Rust: in C++ the optimisation is a property of your build configuration, not of your source. ## HotSpot: speculate, verify, and be able to undo A JIT can do something no ahead-of-time compiler can: look at what has actually happened. HotSpot uses class-hierarchy analysis (if an interface currently has exactly one loaded implementor, the target is unambiguous *right now*) and per-call-site inline caches recording the receiver types observed. A call site that has only ever seen one type is monomorphic and gets a type guard plus a fully inlined body; two types can be handled bimorphically with two guarded branches; beyond that the site is declared megamorphic and falls back to a table lookup with no inlining. Because these are guesses, they need an undo. Compiled code records dependencies ("interface I has one implementor"). When a plugin loads a second implementation, the class loader invalidates the dependency, marks the compiled method not-entrant and **deoptimises** any live frames back to the interpreter, and the method is recompiled with a guard or a table lookup. The operational consequence: a hot path can get slower with no code change and no configuration change, purely because a new type was loaded. It also makes microbenchmarks lie: warm one implementation, measure inlined code, then ship into a process where three exist. ## Objective-C and Ruby: cache, then flush Where classes are mutable at runtime, nothing can be proved at all. Objective-C's `objc_msgSend` resolves a selector against the receiver's class and keeps a per-class cache of recent selector to implementation mappings; a hit is a few instructions. Ruby uses inline caches at call sites, validated against class and global serial numbers. In both, defining, replacing or swizzling a method invalidates caches: per class in Objective-C, and via serial bumps that can invalidate broadly in Ruby. Mutating classes after start-up therefore costs far more than the mutation itself, which is why frameworks do their patching once during boot rather than lazily. ## Why final-by-default helps AOT more than JIT Kotlin classes and C# methods are non-virtual unless you opt in. For an ahead-of-time backend (Kotlin/Native, .NET NativeAOT, Android's AOT compilation) that sealing is exactly the proof the compiler lacked, the same benefit as C++ `final`. On HotSpot it buys much less, because class-hierarchy analysis already discovers that a class has a single implementor and inlines on that basis; the JIT knows the *actually loaded* hierarchy, which is strictly more information than the declaration. What sealing still buys on a JIT is stability: fewer invalidations, fewer deoptimisation storms when new code loads. ## Operating it Diagnose before you restructure: check whether the hot site is monomorphic or megamorphic, whether compiled code is being repeatedly discarded, and whether the regression coincided with a class or plugin being loaded. Then the fixes are ordinary: split a megamorphic site so each copy sees one type, hoist the dispatch out of the loop, or seal the hierarchy. Not "remove polymorphism".
- A benchmark shows a virtual call costing roughly the same as a direct call, yet removing polymorphism from the real service gave a large speedup. How do you reconcile that?The benchmark almost certainly measured the indirect branch in isolation, where a well-predicted target costs a couple of cycles. The real gain came from inlining: once the callee body is visible, the compiler can fold constants, eliminate allocations via escape analysis, and reshape the loop. It is also likely the benchmark's call site was monomorphic while production's was megamorphic, so the branch predictor helped in one case and not the other.
- Your service's hot path degrades minutes after start-up, only in production, and profiles show time moving into the interpreter. What do you suspect?Repeated deoptimisation. The JIT compiled the hot path on an assumption about the loaded class hierarchy or the observed receiver types, and production loads a second implementation (a plugin, a lazily initialised strategy, a proxy) that falsifies it. Compiled code is discarded and frames fall back to the interpreter before recompiling with a guard or table lookup. Confirm with compilation and deoptimisation logging, and correlate against class-loading events.
- If Rust's monomorphisation is faster, why would you ever pass a dyn Trait object?Because monomorphisation costs one copy of the code per concrete type. That inflates binary size and compile time, and a large set of instantiations can evict each other from the instruction cache, making the supposedly faster version slower. dyn Trait also lets you store heterogeneous implementations in one collection and keeps the concrete type out of your public signatures, which matters for compile-time firewalls and dynamic plugin loading.
An inline cache is a receptionist who has learned that everyone asking for "the architect" means the person in room 3, and walks them straight there while still glancing at the badge. Hire a second architect and the shortcut is void: the receptionist has to go back to reading the directory for everyone.
saying these in an interview costs you the question
- Saying virtual calls are slow because of the pointer chase, when the real cost is the inlining and downstream optimisations that never happen.
- Claiming a JIT cannot inline a call that dispatches on the runtime type; HotSpot inlines monomorphic and bimorphic sites behind type guards all the time.
- Believing that marking classes final is a major JVM speedup; class-hierarchy analysis already gets that result for a single-implementor class, though sealing does help AOT compilers.
- Assuming C++ devirtualises virtual calls automatically; without final, an exactly known type, or link-time optimisation the compiler must assume an override exists in another translation unit.
- Treating monomorphisation as strictly better than dynamic dispatch, ignoring binary size, compile time and instruction-cache pressure.
- Thinking a performance change must follow a code change, missing that loading a new implementation can invalidate compiled assumptions on its own.