skip to content

HotSpot's tiered compilation is described as having levels 0 through 4. Walk through what executes at each level and how a method is promoted between them.

level: seniorimportance: should knowfreq 42%

answer

  1. 0 interpreter, 4 = C2, 1/2/3 = C1 variants
  2. 3 = full profiling, 2 = limited counters, 1 = no profiling terminal
  3. Standard path 0 -> 3 -> 4
  4. Thresholds scale with compiler queue length
  5. Level 3 is slower than level 1 - instrumentation costs

basics

~20 s

Level 0 is the interpreter. Levels 1-3 are the fast compiler at different instrumentation settings: 3 is fully profiled, 2 is lightly counted, 1 is fully optimized with no profiling. Level 4 is the optimizing compiler. The usual path is 0 to 3 to 4.

solid answer

~60 s

- **Level 0** - interpreter, collecting basic counters. - **Level 3** - C1 code **with full profiling** (invocation, back-edge, branch and receiver-type counters). This is the normal second stop and where the profile C2 needs is gathered. - **Level 4** - C2 code, fully optimized, no profiling. - **Level 2** - C1 with *limited* counters only, used when the C2 queue is long: get decent code now, keep enough counting to promote later. - **Level 1** - C1 fully optimized with **no** profiling, the terminal state for methods C2 will never want (trivial getters, or when C2 refuses to compile). The common path is **0 -> 3 -> 4**. Promotion is driven by counter thresholds scaled by queue length, so when compiler threads are backed up, thresholds rise and paths like 0 -> 2 -> 3 -> 4 or 0 -> 3 -> 1 appear. Level-3 code is meaningfully slower than level 1 because instrumentation is not free - which is why the JVM tries not to leave methods sitting there.

code

text · 5 lines
text
113   45       3       java.lang.String::hashCode (55 bytes)
    118   52       3       com.example.Order::total (31 bytes)
    241   52 %     4       com.example.Order::total @ 8 (31 bytes)
    242   52       3       com.example.Order::total (31 bytes)   made not entrant
    244   61       1       com.example.Order::getId (5 bytes)

go deeper

for a junior

It is enough to know level 0 is the interpreter, level 4 is the optimizing compiler, and the levels in between are the fast compiler with varying amounts of profiling.

for a middle

Get the standard 0 -> 3 -> 4 path and the role of each C1 variant right, and know that level 3 pays a real instrumentation cost.

for a senior

Reason about the operational consequences: adaptive thresholds under queue pressure, methods stranded at level 3, compiler-thread starvation on small containers, and code-cache footprint.

for a principal

Discuss it as a feedback control system - queue length feeding thresholds - and what that implies for capacity planning, container CPU limits and warmup SLOs across a fleet.

## Why five levels and not two With two compilers you might expect two states. Tiering has five because *profiling is a separate axis from optimization*: C1 can emit code with full instrumentation, partial instrumentation, or none, and each combination is useful in a different situation. ## The levels **Level 0 - interpreted.** Every method starts here. The interpreter maintains method-entry and back-edge counters. It profiles cheaply but executes slowly. **Level 1 - C1, full optimization, no profiling.** The fastest C1 code, but it collects nothing, so a method here can never accumulate the evidence to be promoted. It is a *terminal* state, used for methods deemed trivial (a getter that C2 would gain nothing from) or when C2 declines to compile the method. **Level 2 - C1 with limited profiling.** Only invocation and back-edge counters, not the expensive branch and type profiles. Used as a stopgap when the C2 queue is long: the method gets reasonable code immediately, and enough counting to be promoted later. **Level 3 - C1 with full profiling.** Counters plus branch profiles and receiver-type profiles at virtual call sites. This is where the data C2 needs comes from. Because instrumentation adds work on hot paths, level-3 code is commonly around 30% slower than level 1. **Level 4 - C2.** Fully optimized, speculative, no instrumentation. The destination for genuinely hot code. ## Typical transitions - **0 -> 3 -> 4** is the standard path. Interpret briefly, run profiled C1 code while a good profile accumulates, then compile with C2. - **0 -> 2 -> 3 -> 4** appears under **C2 queue pressure**: rather than keep a hot method in the slow, instrumented level 3, the JVM parks it at level 2 (or moves it there) so it runs faster while waiting. - **0 -> 3 -> 1** or **0 -> 1**: the method is trivial, or C2 refused it (for example the method is too large to inline or compile under current limits), so it settles in unprofiled C1 code. - **4 -> 0**: **deoptimization**. A speculative assumption failed, the compiled frame is discarded and execution resumes in the interpreter; the method may later be recompiled, sometimes at level 3 again to re-profile. ## Thresholds are adaptive, not fixed The trigger is not a single constant. HotSpot combines the invocation counter and the back-edge counter against thresholds (`Tier3InvocationThreshold`, `Tier4InvocationThreshold` and their compound variants) and then **scales those thresholds by the length of the compiler queues**. When compiler threads fall behind, the bar rises so only genuinely hot methods are queued; when the queues are empty, the bar drops and code is promoted sooner. The queues are also priority-ordered, so a method that is heating up fast can overtake one that crossed the threshold earlier. This adaptivity is the reason two runs of the same application can compile the same methods at slightly different times, and why a CPU-starved container behaves differently from a roomy one. ## Why this matters operationally - **A method stuck at level 3 is a performance smell.** It is running instrumented code indefinitely, paying profiling cost with no C2 payoff, typically because it is hot enough to profile but never crosses the level-4 bar, or because C2 keeps refusing it. - **Compiler-thread starvation is real.** On small CPU allocations, C2 compilations queue up, thresholds inflate and warmup stretches out. `-XX:CICompilerCount` and the CPU limit are the levers. - **Code cache holds several versions.** A method may occupy level-3 and level-4 code simultaneously before the old version is flushed, so tiering roughly multiplies code-cache footprint compared with a single-compiler setup. Exhausting the cache (`ReservedCodeCacheSize`) disables the compiler and silently drops the process toward interpreted speed. `-XX:+PrintCompilation` shows the level in each line and is the direct way to observe all of this.

  • You see a hot method that stays at level 3 for the entire run and never reaches level 4. What are the plausible causes?
    Either it never crosses the level-4 threshold - which inflates when compiler queues are long, common on CPU-constrained containers - or the optimizing compiler keeps bailing out on it, for instance because the method is very large or hits a compilation limit. Repeated deoptimization can also cycle it back to a profiled tier. PrintCompilation plus the compiler-queue behaviour usually distinguishes them; the fix is often to split the method or give the process more CPU.
  • Why does level 2 exist at all if level 3 collects strictly more information?
    Because full profiling is expensive. When the optimizing-compiler queue is backed up, a hot method would otherwise sit in slow instrumented code for a long time. Level 2 gives it faster C1 code while retaining just the invocation and back-edge counters needed to promote it later, trading profile richness for throughput during the wait.

saying these in an interview costs you the question

  • Assuming the levels are a strict 0-1-2-3-4 sequence every method walks
  • Thinking level 3 code is faster than level 1 because the number is higher
  • Believing compilation thresholds are fixed constants independent of load
  • Ignoring that a deoptimized method drops back to the interpreter and may re-profile
  • Forgetting that tiering multiplies code-cache usage

context