skip to content

A Java service handles the same kind of request over and over. The first few hundred requests are visibly slower than requests handled a minute later, with no change to code, data or load. What is the JVM doing under the hood, and how does it decide which code deserves the speed-up?

level: juniorimportance: must knowfreq 52%

answer

  1. Interpret first, compile the hot minority
  2. Invocation counter = calls; back-edge counter = loop iterations
  3. Threshold crossed -> async compile request
  4. Profile collected while interpreting
  5. Warm-up curve, early numbers are not steady state

basics

~20 s

The JVM first interprets bytecode while counting execution. Each method has counters for how often it is called and how often its loops iterate; when a counter crosses a threshold the JVM compiles that method to optimized native code, so hot paths speed up after warm-up.

solid answer

~50 s

Java code does not start as machine code. The JVM **interprets** bytecode first: slow per instruction, but it starts instantly and lets the runtime watch real behaviour. While interpreting, the JVM keeps per-method counters - an **invocation counter** bumped on method entry, and a **back-edge counter** bumped whenever control jumps backwards, i.e. once per loop iteration. When counters cross a compilation threshold the method is marked *hot* and queued to a background compiler thread, which emits optimized native code; subsequent executions enter the compiled version instead of the interpreter. That is why JVM processes show a **warm-up curve**: early requests pay interpretation plus profiling overhead, later ones run compiled code, often several times faster. Two practical consequences: numbers taken in the first seconds of a process are not representative, and code that is executed rarely may never be compiled at all - deliberately, because compiling everything would cost CPU and memory for no gain.

go deeper

for a junior

Be able to say: bytecode is interpreted first, the JVM counts how often code runs, hot code gets compiled to native, hence the warm-up curve.

for a middle

Name both counters, explain that compilation is asynchronous and profile-guided, and note that rarely used code is deliberately left interpreted.

for a senior

Tie it to operations: warm-up affects deploy and autoscaling behaviour, first-request latency and the validity of measurements taken right after start.

for a principal

Frame it as a runtime economics decision - spend a bounded CPU and code-cache budget where execution concentrates, and accept a warm-up transient in exchange for profile-guided steady-state code quality.

## Bytecode is not machine code A `.class` file holds **bytecode**: a portable instruction set for an abstract stack machine, not instructions your CPU can run. Something must bridge that gap. HotSpot uses two mechanisms at once - an **interpreter** that decodes and executes one bytecode at a time, and a **just-in-time (JIT) compiler** that translates whole methods into native code. Everything starts in the interpreter, which is why start-up is immediate but the first executions are slow. ## Why not compile everything up front Compilation is not free: it burns CPU on compiler threads and memory in the code cache. In a typical program the overwhelming majority of methods run a handful of times (start-up wiring, configuration parsing, one-off setup) and compiling them would never repay the cost. The runtime therefore spends its compilation budget only where execution actually concentrates. This is the **hot spot** idea the VM is named after: find the small fraction of code responsible for most of the running time and optimize that. ## How the runtime finds hot code Detection is counter-based sampling built into the interpreter (and into compiled code at lower optimization levels): - The **invocation counter** increments on every entry to a method. A method called hundreds of thousands of times is obviously worth compiling. - The **back-edge counter** increments on every backward branch - each iteration of a `for`, `while` or `do` loop. This catches the opposite shape: a method entered only once that spends minutes inside a loop. Without it, `main()` with one giant loop would never qualify as hot. When a counter crosses its threshold, the interpreter raises a compilation request. The request is **asynchronous**: it goes on a queue served by compiler threads, and the interpreter keeps executing the method meanwhile. When the compiled code (an *nmethod*) is installed, the method entry point is patched so later calls jump straight into native code. If the trigger was the back-edge counter, the runtime can also swap the already-running loop over to compiled code without waiting for the method to be re-entered - **on-stack replacement**. ## Profiling, not just counting While interpreting, the JVM also records *what* the code did: which types actually showed up at a virtual call site, which branches were taken, whether a null was ever seen. That profile is what lets the compiler produce code far better than a static compiler could - it can specialize for the types and paths this run actually uses. Counters decide *when* to compile; profiles decide *how*. ## What the warm-up curve looks like A freshly started service typically shows: (1) an initial phase where nearly everything is interpreted and latencies are high and noisy; (2) a middle phase where hot methods are compiling and being installed - throughput climbs, occasionally with hiccups as compiler threads compete for CPU; (3) a steady state where the hot path is fully compiled and latency flattens. Reaching steady state takes anywhere from a few hundred to many thousands of executions of the hot path, depending on thresholds and how much of the code is hot. ## Common misreadings The speed-up is not caching of results, not the OS page cache, and not the garbage collector settling down (though those can contribute separately). It is code being replaced by better code. Equally, it is not permanent in a rigid sense: if the runtime's assumptions about the program are later violated, compiled code can be discarded and execution can fall back to the interpreter, and the counters start the cycle again.

  • Why does the JVM need a back-edge counter at all, when it already counts invocations?
    Because some methods are hot without being called often. A method entered once that runs a loop for minutes has an invocation count of one, so an invocation-only policy would leave it interpreted forever. The back-edge counter measures work done inside the method, which lets the runtime notice long-running loops and compile them.
  • Is compilation synchronous - does the calling thread wait for the compiler?
    No, by default it is asynchronous. The interpreter files a request onto a compilation queue served by dedicated compiler threads and carries on executing the method in the interpreter. When the compiled code is ready it is installed and later executions use it. Blocking compilation exists as a diagnostic option but is not the normal mode.

Like a road authority that does not pave every track at once: it counts traffic first, and only the routes with heavy repeated use get resurfaced.

saying these in an interview costs you the question

  • Saying Java is compiled to native code ahead of time, so no runtime compilation happens
  • Claiming the speed-up is result caching or JVM 'caching the method'
  • Believing every method eventually gets compiled if the process runs long enough
  • Confusing javac (source to bytecode) with the JIT (bytecode to machine code)
  • Assuming compilation blocks the application thread until native code is ready

context