What is on-stack replacement (OSR) in the JVM, and which specific problem does it exist to solve? Give a code shape where a runtime without OSR would perform badly.
answer
- Normal compile takes effect on the NEXT call
- Loop in main: invocation count stays 1
- OSR entry = loop head bytecode index
- Interpreter state packaged, compiled frame built, jump
- PrintCompilation shows % and @bci for OSR
basics
~20 sOSR compiles a method's long-running loop and transfers the currently executing invocation into the compiled code mid-flight, instead of waiting for the next call. Without it, a method entered once that loops for minutes would stay interpreted for its whole run.
solid answer
~50 sNormal JIT compilation replaces a method's **entry point**: the optimized code takes effect from the *next* invocation onwards. That is useless when the invocation you care about is the one already running. The classic shape is a method called once containing a long loop: ```java public static void main(String[] args) { long sum = 0; for (int i = 0; i < 2_000_000_000; i++) sum += work(i); } ``` Here the invocation counter for `main` is 1 forever. The back-edge counter, however, climbs fast. When it overflows, HotSpot requests an **OSR compilation**: a compiled version of the method whose *entry point is the loop head*, keyed to that bytecode index. At the next back-edge the running frame's state - locals and expression stack - is packaged up, a compiled frame is built from it, and execution jumps into the optimized loop and continues from the current iteration. So OSR is 'change the engine while the car is moving': hot loops get optimized without waiting for method re-entry.
code
java · 9 linespublic static void main(String[] args) {
long sum = 0;
// invocation counter for main() = 1, forever
// back-edge counter overflows within milliseconds
for (int i = 0; i < 2_000_000_000; i++) {
sum += i ^ (i >>> 3);
}
System.out.println(sum);
}go deeper
Say what problem it solves: a method called once whose loop runs a long time would otherwise stay interpreted; OSR moves the running loop into compiled code.
Add the mechanism at a high level - back-edge counter triggers a compile keyed to the loop's bytecode index, running frame state is transferred, iteration continues where it was.
Distinguish OSR nmethods from standard nmethods, mention that a method can hold several, and use PrintCompilation's % marker to confirm what path warm-up took.
Treat it as a completeness property of adaptive compilation - without OSR the optimizer cannot reach an entire class of workloads - and use it to argue for structuring hot work as repeatedly called methods.
## The gap normal compilation leaves When HotSpot compiles a method the usual way, it installs an nmethod and patches the method's entry so that **future calls** land in native code. Any invocation already on the stack keeps running in whatever code it started in - the interpreter cannot simply teleport into a compiled method, because the two use completely different frame layouts and value locations (interpreter locals array and expression stack, versus machine registers and a compiled stack frame laid out by the compiler). For most code that is fine: hot methods are called constantly, so 'next call' is microseconds away. It breaks down for a method whose single invocation *is* the workload: - `main` with a top-level processing loop - a batch job or ETL step that iterates until finished - a benchmark harness loop - a worker thread's `run()` that loops for the lifetime of the process Without OSR, all of these would execute interpreted from start to finish, potentially an order of magnitude slower, no matter how hot the loop body is. ## What OSR does Back-edge counting detects those loops (see hot-code detection generally). When a method's back-edge counter overflows at a particular backward branch, HotSpot requests a compilation that is **specific to that loop entry point**: the compile task is keyed on `(method, bytecode index)`, and the resulting nmethod has an unusual entry contract. Instead of expecting arguments in the standard calling convention, an OSR nmethod expects the *complete interpreter state at that bytecode index* - the values of all live locals and the expression stack contents. When the compilation is ready, the next time execution reaches that back-edge the runtime performs the transfer: it collects the current frame's state into a buffer, sets up a compiled frame initialized from it, and jumps to the OSR entry. The interpreted frame is replaced by a compiled one *in place*, in the middle of the method, with the loop counter at whatever value it had reached. Nothing is recomputed and no iteration is repeated. ## What OSR is not - It is **not** deoptimization. Deoptimization is the opposite direction (compiled back to interpreter) and is triggered by broken assumptions, not by hotness. - It is **not** a general-purpose 'hot swap' of running code at arbitrary points; the entry contract is defined at specific loop bytecode indices, not anywhere. - It does **not** make the method's normal entry point compiled. An OSR nmethod is usable only for that mid-method entry; if the method is also called often, a separate standard compilation is produced for normal calls. A method can therefore have several nmethods alive at once: one standard, plus one per OSR entry bci that was hot. ## Practical signals In `-XX:+PrintCompilation` output, OSR compilations are marked with `%` and annotated with the bytecode index of the loop, e.g. `Batch::crunch @ 12`. Seeing a lot of `%` lines tells you your workload is loop-shaped and warming up through OSR rather than through repeated calls. `-XX:-UseOnStackReplacement` disables the mechanism, which is occasionally used to demonstrate the effect: a loop-in-main microbenchmark can become dramatically slower with OSR turned off, which is the cleanest experimental proof that OSR is doing the work. ## Why an interviewer asks this It separates people who have only heard 'the JIT compiles hot methods' from people who understand that compilation must interact with *frames on the stack*. It also underpins a real engineering habit: because OSR-compiled loops are entered under constraints, experienced engineers tend to structure long computations so the hot part is a method called many times, which reaches the normal compilation path.
- Does an OSR compilation also speed up later normal calls to the same method?No. An OSR nmethod has a special entry contract tied to one loop bytecode index and is only usable for transferring a running frame in at that point. If the method is also invoked frequently, the runtime produces a separate standard compilation for the normal entry point. One method can have a standard nmethod plus one OSR nmethod per hot loop entry.
- How is OSR different from deoptimization?OSR moves execution from less optimized code into more optimized code because the loop proved hot. Deoptimization moves execution the other way, from compiled code back to the interpreter, because an assumption the compiler relied on turned out to be wrong. They share machinery for mapping frame state between representations, but the trigger and the direction are opposite.
Normal compilation is repaving a road for tomorrow's traffic; OSR is lifting the car that is already driving onto the new surface without stopping it.
saying these in an interview costs you the question
- Describing OSR as swapping compiled code back to the interpreter (that is deoptimization)
- Claiming the loop restarts from the beginning in compiled form
- Saying OSR replaces the method entry so subsequent calls use the OSR code
- Believing OSR can happen at any bytecode, not at loop entry points
- Assuming OSR is needed for ordinary frequently-called methods