Why can a code path that never ran during warmup be dramatically slow the first few times it executes, even though the method containing it has been JIT-compiled for a long time?
answer
- zero branch count → uncommon trap, not code
- reasons: unreached / unstable_if / null_check
- first execution: trap → interpret → reprofile → recompile
- first exception is doubly expensive: throw + deopt
- fix = warm the cold paths, not disable speculation
basics
~20 sThe compiler saw the branch profile say "never taken" and did not compile that path at all — it put an uncommon trap there. The first execution hits the trap, deoptimizes the frame to the interpreter, reprofiles, and waits for a recompilation that now includes the path.
solid answer
~60 sBranch profiles are part of the speculation surface. If a branch has a zero count when the method is compiled, the optimizing compiler does not emit code for it: it emits an **uncommon trap** with reason `unreached` or `unstable_if`. That is not laziness — deleting the cold side lets the compiler keep registers, drop merge points, and optimize the hot path as straight-line code. The first time that branch is actually taken, the trap fires. The frame is deoptimized to the interpreter and continues there; the method reprofiles and is eventually recompiled with the branch present. Until that recompilation lands, the code runs interpreted or at a lower tier — orders of magnitude slower than steady state. The same effect covers implicit null checks (`null_check` when the first null arrives), never-seen `instanceof`/cast outcomes, and exception paths, which is why the *first* failure in an error handler can be far slower than the successful path it replaced. The practical answer is warmup coverage: exercise error paths, alternate types and boundary branches before real traffic, and expect a latency outlier on genuinely rare paths.
code
text · 12 linesjdk.Deoptimization {
method = com.acme.OrderService.price(Order)
bci = 47
reason = unstable_if
action = reinterpret
}
jdk.Deoptimization {
method = com.acme.OrderService.price(Order)
bci = 61
reason = null_check
action = make_not_entrant
}go deeper
Recall that the JIT skips compiling paths that never ran, so the first time such a path executes the method falls back to the interpreter for a while.
Name the mechanism — zero branch count becomes an uncommon trap — and the sequence trap → deopt → reprofile → recompile, plus the null-check and cast variants.
Connect it to production symptoms, confirm it from JFR deoptimization events or PrintCompilation, and prescribe warmup coverage for error and alternate paths.
Treat it as a tail-latency budget question: which rare paths must be pre-warmed, what SLO impact a one-time cliff has, and why global de-speculation is the wrong trade.
## What the profile actually records for branches The interpreter and the profiling compilation tier count both directions of every conditional. When the optimizing compiler runs, a branch can be: - **balanced** — compile both sides normally; - **strongly biased** — compile both, but lay out the hot side to fall through and predict accordingly; - **never taken (count zero)** — treat as unreachable and emit an **uncommon trap** in its place. The third case is the interesting one. Replacing an entire branch with a trap is not just skipping work: it removes a control-flow merge, which lets the compiler keep values in registers across the region, propagate constants and types that the merge would have destroyed, avoid spilling, and reduce code size (important for the instruction cache and for further inlining decisions). Deleting the cold half is often what makes the hot half fast. ## What happens on first execution of the cold path 1. Control reaches the trap stub with reason `unreached` (never compiled) or `unstable_if` (the profile said this side was dead). 2. The runtime deoptimizes: the compiled frame is expanded into interpreter frame(s) with locals and expression stack restored, and execution continues in the interpreter at that exact bytecode index. 3. The compiled nmethod is normally made **not entrant** and the method reprofiles. 4. After enough invocations, the method is recompiled — now with a real branch, because the counter is no longer zero. Between steps 2 and 4 the method runs interpreted or at a lower tier. For a method that was hot, that can be a 10–50× slowdown for the affected calls, plus the compilation work itself competing for CPU. ## The other members of the same family Branch pruning is the clearest case, but the pattern repeats wherever the profile encodes "never seen": - **Implicit null checks.** If a dereference has never seen null, the compiler omits the test entirely and relies on the hardware fault plus a signal handler; the first null converts a would-be `NullPointerException` path into a `null_check` deopt. - **Casts and `instanceof`.** A never-failing cast compiles without the slow path; the first failure traps. - **Exception paths.** Handlers whose bytecodes have never executed are prime candidates for pruning, so the first thrown exception in a long-running service is disproportionately expensive — the throw itself, stack-trace capture, *and* a deopt plus recompile. - **Array store checks, division by zero, class initialization checks** on classes not yet initialized at compile time. ## Why it shows up as a production symptom Typical shapes: - A service is fine for hours; the first downstream timeout produces a latency spike far larger than the timeout budget suggests, because the error path deoptimizes several hot methods at once. - A code path guarded by a rarely-true condition — month-end, leap day, a feature flag, an unusual currency — runs slowly on exactly the day it matters. - A canary deploy that shifts traffic mix makes previously-cold branches warm, producing a burst of deopts and recompiles that looks like a memory or GC problem but is compilation. `-XX:+PrintCompilation` shows `made not entrant` clustered at the moment of the event; JFR's `jdk.Deoptimization` events name the method, bytecode index and reason (`unstable_if`, `unreached`, `null_check`) directly, which is the quickest confirmation. ## What to do about it - **Warm the paths you care about.** If tail latency on error handling matters, exercise the error paths during warmup, not just the happy path. The same applies to alternate types at polymorphic sites. - **Don't use exceptions for control flow** — beyond the direct cost, every rare-path throw risks a deopt. - **Expect one-time costs and measure steady state separately** from first-occurrence cost; conflating them makes a healthy system look pathological. - **Do not attempt to disable speculation globally.** Turning off branch pruning or trap-based optimization would slow the hot path everywhere to protect a rare one; the correct lever is warmup coverage, and only in the rarest cases per-method compiler directives. The underlying bargain is worth stating plainly: the JVM buys hot-path speed by refusing to compile what has not happened, and pays for it with a one-time cliff when the unexpected finally happens.
- Why does the compiler delete a never-taken branch instead of just compiling it as cold code?Removing the branch removes a control-flow merge, so values stay in registers, types and constants propagate through the region, and no spill code is needed. It also shrinks the compiled method, which helps the instruction cache and leaves inlining budget for other calls. Compiling the cold side would cost all of that to serve a path the profile says never happens.
- Why is the first exception thrown in a long-running service unusually expensive?Three costs land together: constructing the exception with a stack trace walk, executing a handler path that was likely pruned as unreached, and the resulting deoptimization plus recompilation of a previously hot method. Subsequent throws are cheaper because the path is now compiled and profiled.
- How would you confirm that a latency spike came from deoptimization rather than GC?Correlate timestamps across sources: JFR jdk.Deoptimization events or -XX:+PrintCompilation 'made not entrant' lines should line up with the spike, while GC logs should show no pause of comparable length. Deopt spikes also concentrate on the first occurrence of an event and shrink on repetition, whereas GC pauses recur with allocation.
saying these in an interview costs you the question
- Attributing every first-occurrence latency spike to garbage collection without checking compilation events
- Believing a compiled method contains machine code for every branch in the source
- Thinking a single deoptimization permanently degrades the method
- Proposing to disable JIT speculation or drop to a lower tier to avoid rare-path deopts, penalizing the hot path everywhere
- Confusing this with class-loading invalidation — here nothing is loaded, the data simply took a path the profile called impossible