Interpreting bytecode instruction by instruction is portable but slow. Mechanically, where does the slowness come from, and why does a runtime that interprets first still keep an interpreter around after it can compile code to native instructions?
answer
- decode + indirect dispatch per opcode
- stack traffic instead of registers
- no inlining / no cross-instruction optimization
- cold code dominates → compile only hot
- interpreter = profiler + deopt landing site
basics
~20 sEach bytecode costs decode plus an indirect dispatch on top of its actual work, values shuttle through the operand stack instead of registers, and nothing is optimized across instruction boundaries. The interpreter is kept because it starts instantly, costs no compile time for cold code, and is the fallback when compiled code must be abandoned.
solid answer
~60 sPer instruction, an interpreter pays overhead the work itself does not need: - **Dispatch.** Fetch opcode, jump through a table to the handler, advance the program counter — an indirect branch per bytecode, often mispredicted. - **Operand traffic.** Values move through the operand stack rather than staying in CPU registers, so trivial arithmetic becomes loads and stores. - **No cross-instruction optimization.** Each bytecode is executed in isolation: no constant folding, no common subexpression elimination, no inlining, no loop-invariant hoisting. Every virtual call is a real dispatch. The usual figure is an order of magnitude or more versus optimized native code, which is why hot methods get JIT-compiled. The interpreter is nonetheless permanent, because compilation is not free: - it starts executing immediately, with no compile pause — better startup and better behavior for code run once; - most methods in a large application are cold; compiling them all would waste CPU and code cache; - it collects the profile (invocation counts, branch and receiver-type data) that makes later compilation aggressive; - it is the landing site for **deoptimization** when a speculative assumption in compiled code turns out to be wrong.
code
text · 2 lines$ java -Xint -jar bench.jar # interpreter only: no JIT compilation at all
$ java -jar bench.jar # default: interpret, profile, compile hot methodsgo deeper
Know that the interpreter executes one bytecode at a time with per-instruction overhead, that hot code is compiled to native code later, and that this is why the first runs are slower.
Name the three cost sources — dispatch, operand-stack traffic, no cross-instruction optimization — and explain that only hot methods are worth compiling.
Emphasize the interpreter's two structural roles beyond execution: gathering the profile that licenses speculative optimization, and being the destination for deoptimization. Connect this to warmup effects in production and in benchmarks.
Frame it as an economic policy: compilation is an investment with a payback period, so the runtime spends CPU only where evidence says it repays, and keeps a correct-but-slow tier so speculation stays safe.
## What an interpreter does per instruction Conceptually the execution loop is: ``` while (true) { opcode = code[pc]; switch (opcode) { ... } // or: jump to handler[opcode] } ``` For every bytecode the machine must: load the opcode byte, decode any inline operands, compute the address of the handler, take an **indirect branch** to it, perform the actual semantic work, update the program counter, and branch back. The semantic work for `iadd` is a single machine add. Everything around it is overhead — and it is overhead that recurs on *every* execution of that instruction, including on the millionth iteration of a loop. Three distinct costs stack up: **1. Dispatch overhead.** The indirect branch at the end of each handler is hard for branch predictors: the next bytecode is data-dependent. A mispredict costs on the order of a dozen-plus cycles on a modern out-of-order core — vastly more than the instruction being interpreted. Interpreter designs fight this with threaded dispatch (each handler ends with its own indirect jump, giving the predictor per-opcode history) rather than a single shared switch. **2. Data movement through the operand stack.** Bytecode semantics say values live on the operand stack. A faithful interpreter therefore stores intermediate results to memory and reloads them, where compiled code would keep them in registers. `x = x + 1` becomes load-local, push constant, add, store-local: four dispatches and several memory touches for what compiled code does in one increment. (Real interpreters cache the top of stack in a register to claw some of this back.) **3. Absence of optimization across instructions.** This is the big one and it is structural, not an implementation weakness. An interpreter sees one instruction at a time, so it cannot: - inline a call — every `invokevirtual` performs a real dispatch through the receiver's method table, and the callee's body can never be merged with the caller's; - fold constants or eliminate redundant loads; - hoist loop-invariant computation out of a loop; - eliminate a bounds check it could prove redundant; - keep a hot object's fields in registers, or scalar-replace an object that never escapes. Compiled code gets all of these because the compiler sees a whole method (and, after inlining, effectively several). ## Why not compile everything up front, then? Because compilation costs time and memory, and most code does not repay it: - **Startup latency.** Compiling every method before running it delays the first useful work; interpreting begins immediately. For a short-lived process, or the first seconds of a server's life, the interpreter is genuinely the faster choice. - **Cold code dominates.** In a large application the overwhelming majority of loaded methods execute a handful of times — configuration parsing, one-shot initialization, error paths. Native code for them consumes CPU to produce and code cache to hold, and buys nothing. Runtimes therefore compile only what proves hot, using counters the interpreter maintains. - **The profile has to come from somewhere.** Aggressive optimizations are speculative: "this call site has only ever seen one receiver type, so inline it", "this branch is never taken, so do not emit it", "this exception path is dead". Those facts are observations about the running program. The interpreter is the cheap instrument that gathers them — invocation counts, loop back-edge counts, per-call-site receiver types, branch taken/not-taken data. - **Speculation needs a safety net.** When a speculative assumption is invalidated — a new subclass is loaded and the monomorphic call site becomes polymorphic, or a "never taken" branch is finally taken — the runtime must abandon the compiled code mid-execution and continue somewhere correct. That somewhere is the interpreter: the runtime reconstructs an interpreter frame (locals and operand stack) at the corresponding bytecode index and resumes interpreting. Without an interpreter, deoptimization would have no destination, and speculative optimization would be unsafe. - **Some code is not compilable in practice.** Very large methods, and certain rarely executed constructs, are simply left to the interpreter. ## How to observe the effect Running with interpretation only (`-Xint`) versus the default typically slows CPU-bound benchmarks by roughly an order of magnitude — a crude but convincing demonstration of the gap. Conversely, this is why benchmark harnesses insist on **warmup**: measurements taken while code is still interpreted describe the interpreter, not the program. ## The takeaway framing Interpretation and compilation are not competitors; they are a two-speed system. Interpretation gives portability and instant start with no speculation; compilation gives speed but needs evidence, time, and a way back. A production JVM runs both simultaneously, on different methods, all the time.
- If the interpreter is a profiler, what exactly does it record?Per-method invocation counters and per-loop back-edge counters, which together decide when a method is hot enough to compile; and per-call-site type profiles recording which receiver classes actually appeared, plus branch taken/not-taken data. That evidence is what lets the compiler inline speculatively and prune branches it believes are dead.
- Why does a benchmark need a warmup phase?Because early iterations run interpreted or in lightly optimized compiled code, and class loading plus profiling are still in progress. Measuring then reports the runtime's transient state rather than steady-state performance. Harnesses run untimed warmup iterations until compilation has settled before recording results.
Interpreting is translating a speech sentence by sentence in real time; JIT compiling is producing a polished written translation of the paragraphs that get read over and over — but you keep the live interpreter for the parts said only once, and for when the prepared text turns out to be wrong.
saying these in an interview costs you the question
- Saying interpretation is slow only because 'it reads text' — bytecode is binary, not source
- Claiming the JVM stops interpreting once JIT compilation kicks in
- Proposing to compile every method eagerly with no mention of startup, code cache, or cold code
- Believing the JIT can speculate without a profile or without a deoptimization fallback
- Treating first-run timings as representative performance