skip to content

JIT & Adaptive Optimization

The adaptive compiler that makes a bytecode interpreter fast: hot-code detection, tiered compilation, inlining, speculation, and deoptimization when a guess turns out wrong. Interviewers bring it up to see whether you can explain why the same code gets faster after a few thousand iterations — and why your benchmark probably lied.

on this pageshow

explore

questions

25

Java is often described as both compiled and interpreted. After a class is loaded, what actually executes its methods, and how does that change while the program keeps running?

level: juniorimportance: must knowfreq 68%

answer

  1. javac -> bytecode; JIT -> machine code
  2. interpreter starts instantly, runs slowly
  3. invocation + back-edge counters mark hotness
  4. background compiler threads, no blocking
  5. JIT knows loaded classes, real profiles, real CPU

basics

~20 s

Source is compiled ahead of time into bytecode. At runtime the JVM first interprets that bytecode, counts how often each method runs, and once a method is hot a just-in-time compiler turns it into native machine code used from then on. Both modes run side by side.

solid answer

~50 s

There are two compilations. `javac` compiles source to **bytecode** - a portable instruction set, not machine code. At runtime the JVM executes bytecode in its **interpreter**, which starts instantly but is slow per instruction. While interpreting, the JVM increments per-method invocation counters and loop back-edge counters. When a method crosses a threshold it is queued for a **JIT compiler**, which translates it into native code for the actual CPU and installs it; later calls enter the compiled version instead of the interpreter. This is **mixed-mode execution**: at any instant some methods are interpreted and some compiled, and the same method can exist in both forms. The payoff is that the JVM pays compilation cost only for code that matters, and the compiler can exploit runtime facts an ahead-of-time compiler cannot know - observed receiver types, real branch frequencies, which classes are actually loaded.

code

text · 3 lines
text
java -Xint  MyApp    # interpreter only: no JIT at all, dramatically slower steady state
java -Xcomp MyApp    # compile every method on first invocation: awful startup, no profile quality
java        MyApp    # default: mixed mode - interpret, profile, compile the hot parts

go deeper

for a junior

Recall the two-step story clearly: javac makes bytecode, the JVM interprets it, hot methods get JIT-compiled to native code. Name the counters as 'the JVM counts how often a method runs'.

for a middle

Add why selective compilation is the right economics (skewed hotness distribution), that compilation is asynchronous on background threads, and that both forms of a method can coexist.

for a senior

Connect it to observable behavior: warmup curves, short-lived processes never reaching peak, and why first-request latency in a freshly deployed pod is not a code defect.

for a principal

Frame it as a runtime-versus-startup tradeoff the platform makes on your behalf, and note where you would override it - AOT/archived-profile approaches for short-lived or latency-critical-at-start workloads.

## Two different compilations "Compiled" and "interpreted" are both true of Java because the word covers two separate steps. The **first compilation** is `javac`: it turns source into `.class` files containing **bytecode**, a compact instruction set for an abstract stack machine (`iload`, `iadd`, `invokevirtual`). No real CPU executes bytecode. This step is almost entirely non-optimizing - `javac` does little beyond straightforward translation, deliberately leaving optimization to the runtime. The **second compilation** happens inside the running JVM and produces actual x86-64 or AArch64 instructions. ## The interpreter runs first When a method is first called, its bytecode is executed by the **interpreter**: a loop that fetches a bytecode, dispatches to a small native routine implementing it, and moves on. (HotSpot's is a template interpreter - machine-code stubs generated at startup - not a naive switch loop, but the cost model is the same.) Interpreting is roughly an order of magnitude slower than good native code, because each bytecode costs dispatch overhead and values move through a stack rather than living in registers. Its virtue is that it starts *immediately*: zero compile latency, and it can begin executing a method the instant the class is linked. ## Profiling decides what gets compiled While interpreting, the JVM maintains counters per method: how many times the method was entered, and how many times a loop inside it branched backwards (a **back-edge counter**, which catches a method called once but looping a million times). When the counters cross a threshold the method is judged **hot** and submitted to a compilation queue. Compilation happens on **background compiler threads**. The application thread does not block waiting for machine code; it keeps interpreting until the compiled version is installed, at which point the next call enters native code. This is why the same method can be interpreted and compiled at the same moment on different threads. ## Why not compile everything up front? Most code in a large application executes a handful of times - startup wiring, configuration parsing, error paths. Compiling all of it would spend CPU and memory on code that never runs again and would push startup latency up badly. The distribution is extremely skewed: a small fraction of methods accounts for the overwhelming majority of executed instructions. Selective compilation targets exactly that fraction. There is a second, subtler reason. A just-in-time compiler knows things a static compiler cannot: - **Which classes are actually loaded.** If only one implementation of an interface has ever been loaded, a virtual call can be compiled as a direct call. - **Observed types at each call site.** Profiles record which concrete receiver types actually showed up. - **Real branch frequencies.** A null check that has never fired can be compiled away and replaced by a trap. - **Actual CPU features.** The code is generated for the machine it is running on, not a lowest common denominator. These enable *speculative* optimizations: the compiler bets on what it observed and installs a guard so that, if the bet is later invalidated (a new subclass is loaded, an untaken branch is finally taken), execution can fall back to the interpreter. ## Consequences you can see - **Warmup.** A freshly started JVM is slow, then gets faster over seconds to minutes as hot code is compiled. Steady-state throughput is not observable in the first requests. - **Short-lived processes lose out.** A CLI tool that runs for 300 ms mostly interprets and may never recover the compilation investment. - **Performance is a moving target.** Measurements taken before warmup are not comparable to steady-state numbers. `-Xint` forces interpretation only and `-Xcomp` forces compilation on first call; both are diagnostic tools, not production settings, and both are much slower than the default mixed mode.

  • Why does HotSpot let a method keep running in the interpreter while its compilation is queued, instead of waiting for the compiled code?
    Compilation happens asynchronously on background compiler threads so that application threads never stall on the optimizer. The interpreted version is already correct, just slower, so the safe move is to keep executing it and switch on the next entry after the native code is installed. Blocking would turn every compilation into a latency spike on a request thread.
  • How does the JVM handle a method that is entered only once but loops a hundred million times?
    The invocation counter would never trigger, so HotSpot also counts loop back-edges. Once the back-edge count crosses its threshold the loop body is compiled and execution is transferred into the compiled code mid-loop via on-stack replacement, so the currently running invocation benefits rather than waiting for a next call that never comes.

Like a simultaneous interpreter at a conference who translates sentence by sentence at first, but when a speaker keeps repeating the same passage, someone hands out a printed translation and everyone reads that instead.

saying these in an interview costs you the question

  • Saying javac produces machine code, or that bytecode is CPU-specific
  • Claiming the JVM interprets everything forever and is therefore inherently slow
  • Believing a method is either interpreted or compiled permanently, never both
  • Thinking compilation blocks the application thread that triggered it
  • Assuming ahead-of-time compilation is strictly better - it cannot use runtime profiles or class-hierarchy facts

context

open as a page

A Java service handles the same kind of request over and over. The first few hundred requests are visibly slower than requests handled a minute later, with no change to code, data or load. What is the JVM doing under the hood, and how does it decide which code deserves the speed-up?

level: juniorimportance: must knowfreq 52%

basics

~20 s

The JVM first interprets bytecode while counting execution. Each method has counters for how often it is called and how often its loops iterate; when a counter crosses a threshold the JVM compiles that method to optimized native code, so hot paths speed up after warm-up.

open as a page

HotSpot ships two just-in-time compilers, historically called the client (C1) and server (C2) compilers. How do they differ in design goals and output, and why does a single JVM contain both?

level: middleimportance: must knowfreq 58%

basics

~20 s

C1 compiles fast and optimizes lightly, so code becomes native quickly with modest peak speed. C2 optimizes aggressively and compiles slowly, giving the best steady-state throughput. A single JVM uses both: C1 for quick warmup and profiling, C2 for the hottest methods.

open as a page

What does it mean for a JIT compiler to compile a method speculatively, and what does the JVM do at runtime when one of those speculative assumptions turns out to be wrong?

level: middleimportance: must knowfreq 50%

basics

~20 s

The JIT compiles using the profile gathered so far — e.g. only one receiver type seen, this branch never taken — and replaces the unproven path with an uncommon trap. If the assumption breaks, the trap fires, the runtime deoptimizes that frame back to the interpreter at the same bytecode, and the method is later recompiled without the failed assumption.

open as a page

HotSpot decides which bytecode to compile using two per-method counters. Name them, explain why one is not enough, and describe what the runtime does when a counter crosses its threshold.

level: middleimportance: must knowfreq 45%

basics

~20 s

An invocation counter (incremented on method entry) and a back-edge counter (incremented on every backward branch, i.e. loop iteration). Invocations alone miss long-running loops in rarely called methods. Crossing a threshold files an asynchronous compilation request; back-edge overflow requests on-stack replacement.

open as a page

What is on-stack replacement (OSR) in the JVM, and which specific problem does it exist to solve? Give a code shape where a runtime without OSR would perform badly.

level: middleimportance: must knowfreq 40%

basics

~20 s

OSR compiles a method's long-running loop and transfers the currently executing invocation into the compiled code mid-flight, instead of waiting for the next call. Without it, a method entered once that loops for minutes would stay interpreted for its whole run.

open as a page

When a JIT compiler inlines a method call, what happens to the generated machine code, and why is inlining called the optimization that unlocks most of the others?

level: middleimportance: must knowfreq 60%

basics

~20 s

Inlining replaces a call with a copy of the callee's body inside the caller. Removing call overhead is the small win; the real win is that both bodies become one compilation unit, so constant folding, null and bounds-check removal, escape analysis and loop optimizations can work across the boundary that disappeared.

open as a page

Java calls through interfaces and overridable methods are virtual. How can a JIT compiler still inline such a call, and what must be proved or guarded for that to be legal?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Two routes. Proof: class-hierarchy analysis shows only one loaded implementation, so the compiler emits a direct inlined call and records a dependency — loading a second implementation invalidates that compiled code. Speculation: the receiver-type profile shows a dominant type, so the compiler emits a class check, inlines behind it, and traps on mismatch.

open as a page

What do constant folding, constant propagation, and dead-code elimination do inside a just-in-time compiler, and how does each pass create new work for the others?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Constant folding evaluates expressions whose operands are already known at compile time. Constant propagation substitutes those known values into later uses. Dead-code elimination deletes computations nobody reads and branches that can never be taken. Each pass exposes fresh opportunities for the others.

open as a page

Does wrapping a Java field in a getter method make code slower at run time? Explain what the JVM does with a one-line accessor in hot code.

level: juniorimportance: should knowfreq 42%

basics

~20 s

In the interpreter the getter really is a call. Once the code is hot and JIT-compiled, a tiny accessor is inlined and the call disappears, leaving just the field load. Caveats: it only applies after warm-up, and only if the JIT can pin the call to one target.

open as a page

The Java language requires an ArrayIndexOutOfBoundsException on any out-of-range array access, so every array read carries a bounds check. How does an optimizing JIT compiler avoid paying for that check on each iteration of a hot loop, without breaking the exception semantics?

level: middleimportance: should knowfreq 46%

basics

~20 s

It proves the whole loop is in range once. The compiler hoists a single guard — the index starts at or above zero and the highest index is below the array length — before the loop, then compiles the loop body with no per-element check. If the guard fails, control goes to a slow path that runs checked code and throws at the right iteration.

open as a page

HotSpot does not inline every call. Which properties of the call site and of the target method decide the outcome, and what practical size and depth limits apply?

level: middleimportance: should knowfreq 45%

basics

~20 s

Three families of limits: callee size (tiny methods almost always, small ones if the site is hot, big ones never), call-site hotness from the collected invocation profile, and inlining depth plus total expansion budget. On top of that, the target must be statically bindable and compilable.

open as a page

What is loop-invariant code motion in an optimizing compiler, and which conditions prevent it from hoisting a field load or a computation out of a loop?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Loop-invariant code motion hoists a computation whose inputs never change inside the loop to a single evaluation before the loop. It is blocked when an input might change per iteration: a write inside the loop, an unanalysable call, possible aliasing, acquire semantics on a volatile read, or an operation whose exception or side effect must stay inside the loop.

open as a page

HotSpot's tiered compilation is described as having levels 0 through 4. Walk through what executes at each level and how a method is promoted between them.

level: seniorimportance: should knowfreq 42%

basics

~20 s

Level 0 is the interpreter. Levels 1-3 are the fast compiler at different instrumentation settings: 3 is fully profiled, 2 is lightly counted, 1 is fully optimized with no profiling. Level 4 is the optimizing compiler. The usual path is 0 to 3 to 4.

open as a page

A JIT-compiled method inlined a virtual call because only one implementation of the interface had been loaded at compile time. What happens to that compiled code when the application later loads a second implementation?

level: seniorimportance: should knowfreq 32%

basics

~20 s

The compiler recorded a dependency on "only one implementation exists". Loading a second one violates it, so during class loading the JVM invalidates the compiled code: it is made not entrant, frames on the stack are deoptimized, and the method is recompiled with a real virtual call or a type-guarded inline.

open as a page

Why can a code path that never ran during warmup be dramatically slow the first few times it executes, even though the method containing it has been JIT-compiled for a long time?

level: seniorimportance: should knowfreq 30%

basics

~20 s

The compiler saw the branch profile say "never taken" and did not compile that path at all — it put an uncommon trap there. The first execution hits the trap, deoptimizes the frame to the interpreter, reprofiles, and waits for a recompilation that now includes the path.

open as a page

Walk through, step by step, what the JVM actually does when it performs on-stack replacement on a frame that is currently executing an interpreted loop. Include how the running frame's state survives the transfer.

level: seniorimportance: should knowfreq 26%

basics

~20 s

Back-edge counter overflows at a loop bytecode index; a compile task keyed to (method, bci) produces an OSR nmethod whose entry expects the interpreter's live locals and stack. At the next back-edge the runtime copies that state into an OSR buffer, builds a compiled frame from it, and jumps to the OSR entry; the interpreted frame is discarded.

open as a page

A hot Java call site dispatches on many different receiver classes. What does HotSpot do at such a megamorphic site, why is it much slower than a site that sees one class, and what can you change?

level: seniorimportance: should knowfreq 40%

basics

~20 s

With many receiver types the inline cache gives up and the call becomes a table dispatch — a virtual-table index, or a slower interface-table search for interface calls. Nothing is inlined, so all downstream optimization is lost, and the indirect branch mispredicts. Fix by restoring one type per hot call site.

open as a page

A microbenchmark shows library A three times faster than library B, but replacing B with A in production shows no measurable improvement. How do you decide whether the benchmark result was real and what, if anything, to do next?

level: principalimportance: should knowfreq 30%

basics

~20 s

Both numbers can be correct. A microbenchmark measures an isolated operation under an artificially uniform profile; production runs it inside a mix where it may be a tiny share of total time, share call sites with other implementations, and be bounded by memory or I/O. Check the operation's share of production time first, then re-measure at the level you actually care about.

open as a page

You want to confirm that the JIT compiler actually compiled and inlined the method your benchmark claims to measure. Which HotSpot diagnostic flags would you enable, and what would you look for in their output?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Use -XX:+PrintCompilation to see which methods were compiled, at which tier, and which were made not-entrant (deoptimized). Add -XX:+UnlockDiagnosticVMOptions -XX:+PrintInlining to see which callees were inlined or why they were refused. Confirm the method under test reaches the top tier and its hot callee is inlined.

open as a page

HotSpot's C2 compiler has recognised a hot int-counted loop over an array and decided to unroll it. Describe the shape the compiled code actually takes — how the iteration space gets split up, what property the surviving hot loop's trip count must have, where SIMD (vector) instructions come from, and how safepoint polling is kept bounded inside such a loop. Then name the factors that cap how large an unroll factor the compiler will pick.

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

C2 splits the iteration space into pre-loop, main loop and post-loop. The main loop is the unrolled one: its trip count is a multiple of the unroll factor, its bounds checks are gone, and the superword pass fuses its copies into SIMD instructions. Strip mining wraps it so safepoints still get polled.

open as a page

Mechanically, how does the JVM turn an optimized compiled stack frame back into interpreter frames, given that the compiler may have inlined several callees into it, kept values only in registers, and eliminated object allocations entirely?

level: seniorimportance: nice to knowfreq 18%

basics

~20 s

Compiled code carries debug metadata mapping each trap or safepoint to a chain of virtual frames — method, bytecode index, and where every local and stack value physically lives. Deoptimization reads it, rebuilds one interpreter frame per inlined method, rematerializes scalar-replaced objects and re-locks eliminated locks, then resumes interpreting.

open as a page

You run one service image both as a long-lived server and as short-lived batch jobs and serverless invocations. How would you approach just-in-time warmup as a platform-level concern across that fleet?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Match the compiler policy to process lifetime. Long-lived processes keep full tiering and get traffic-shaped warmup before receiving load. Short-lived ones stop at the fast compiler tier and lean on class-data sharing or cached profiles, because they will never repay optimizing-compiler cost.

open as a page

A latency-sensitive JVM service shows recurring throughput dips, and its compilation output shows the same few methods being invalidated and recompiled over and over in steady state. How would you investigate whether repeated deoptimization is the cause, and what would you change?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Confirm with evidence: JFR deoptimization events or -XX:+PrintCompilation "made not entrant" lines correlated with the dips, grouped by method, bytecode index and reason. Then treat the reason — megamorphic or unstable call sites, exceptions as control flow, late class loading — rather than tuning trap limits. Rule out code-cache exhaustion, which looks identical.

open as a page

Code that the JVM optimizes only through on-stack replacement of a running loop often ends up slower than the same loop reached through ordinary method compilation. Why does that happen, and how would you structure a long-running compute loop in production code given that?

level: principalimportance: nice to knowfreq 20%

basics

~20 s

An OSR entry fixes the program state at a loop head mid-flight, so the compiler has less freedom: limited loop peeling and unrolling, values pinned to the entry contract, and no profile for the pre-loop code. Structuring the hot work as a method called many times reaches the ordinary compilation path instead.

open as a page