You want to confirm that the JIT compiler actually compiled and inlined the method your benchmark claims to measure. Which HotSpot diagnostic flags would you enable, and what would you look for in their output?
answer
- PrintCompilation: what compiled, which tier, when it went quiet
- made not entrant = deoptimized and being recompiled
- UnlockDiagnosticVMOptions + PrintInlining: inlined or refused, with reason
- PrintAssembly = ground truth, verbose
- Run diagnostics in a separate run from the reported numbers
basics
~20 sUse -XX:+PrintCompilation to see which methods were compiled, at which tier, and which were made not-entrant (deoptimized). Add -XX:+UnlockDiagnosticVMOptions -XX:+PrintInlining to see which callees were inlined or why they were refused. Confirm the method under test reaches the top tier and its hot callee is inlined.
solid answer
~50 sTwo flags cover most of it. `-XX:+PrintCompilation` prints a line per compilation event: a timestamp, the compilation identifier, tier level, flags, and the method. Reading it, you check that the method under test appears, that it climbs to the top tier, and that the stream goes quiet by the time measurement begins. Entries marked *made not entrant* or *made zombie* indicate deoptimization and recompilation — if those keep appearing during measurement, you are not in steady state. `-XX:+UnlockDiagnosticVMOptions -XX:+PrintInlining` prints the inlining tree for each compilation with the decision per callee, including reasons for refusal such as the callee being too big, the call site being megamorphic, or the callee being cold. This is how you confirm the small method you assumed was inlined actually was. Beyond that, `-XX:+PrintAssembly` (with a disassembler plugin) shows the final machine code, and GC logging separates collection pauses from compilation effects. Run these on the benchmark JVM, not production.
code
text · 8 lines# what got compiled, at which tier, and when
java -XX:+PrintCompilation -jar benchmarks.jar
# add per-callee inlining decisions (diagnostic flag needs unlocking)
java -XX:+UnlockDiagnosticVMOptions -XX:+PrintInlining -jar benchmarks.jar
# control run: interpreted only, should be dramatically slower
java -Xint -jar benchmarks.jargo deeper
Know that -XX:+PrintCompilation shows which methods the JIT compiled, and that a quiet log during measurement means warmup finished.
Add PrintInlining behind UnlockDiagnosticVMOptions, and be able to say why inlining matters — it is the optimization that enables the others.
Read the output critically: tier climb, not-entrant markers, OSR entries, refusal reasons, and separate compilation effects from GC pauses using GC logs.
Insist on hypothesis-then-verification as team practice, and be explicit that a benchmark whose compilation shape nobody can describe is not evidence for a decision.
## Why look at compiler output at all A benchmark reports a number; it does not report whether the number describes the code you think it does. The two failure modes the flags below catch are: the method under test never reached the optimizing tier during measurement, and an assumption about inlining ("this tiny accessor is free") is simply false in this configuration. Both quietly invalidate conclusions. ## `-XX:+PrintCompilation` This is the workhorse. It emits one line per compilation event with, roughly: elapsed milliseconds since start, a compilation id, a set of single-character flags, the tier level, and the method signature with its compiled size. What to read out of it: - **Did the method under test get compiled at all?** If it never appears, it never crossed the thresholds; your benchmark is timing interpreted code. - **What tier did it reach?** Under tiered compilation, methods climb from interpreted through fast-compiler tiers to the top optimizing tier. Measuring while the method is still at an intermediate tier measures profiled, less-optimized code. - **Is the log quiet during measurement?** A stream of compilation events during the timed phase means the JVM was still warming up. This is the single most useful sanity check the flag provides. - **Are methods being made not entrant?** That marker means an existing compiled version has been invalidated, typically because a speculative assumption failed; the method will be recompiled. Occasional occurrences during warmup are normal; repeated occurrences during measurement mean the workload keeps invalidating the compiler's assumptions and the steady state you are reporting is not stable. - **OSR compilations** appear separately and indicate a long-running loop was swapped to compiled code mid-execution — usually a sign that a benchmark's own loop, rather than the method under test, is the hot region. ## `-XX:+PrintInlining` This flag is diagnostic, so it needs `-XX:+UnlockDiagnosticVMOptions` first. For each compilation it prints the inlining tree: which callees were inlined into the compiled method, at what depth, and for those that were not, a stated reason. Common refusal reasons include the callee being too large to inline at its call frequency, the call site being megamorphic so no single target can be chosen, the callee being considered cold, the inlining depth limit being reached, or the total compiled size limit being hit. Why it matters for benchmarking: - Inlining is the enabling optimization. Escape analysis, constant folding across call boundaries and dead-code elimination mostly become possible only after inlining exposes the callee's body. If your callee was not inlined, the optimizations you assumed applied did not. - Conversely, if the benchmark harness's own wrapper code got inlined into your measured region, the boundary between harness and measurement blurred. - A benchmark that inlines everything because only one implementation is loaded is telling you about a monomorphic world. If production has several implementations, the inlining decision is different and the benchmark's advantage may not survive. ## Supporting flags - **`-XX:+PrintAssembly`** (with `-XX:+UnlockDiagnosticVMOptions` and a disassembler library installed) prints the generated machine code. This is the ground truth — it will show you the vector instructions, the eliminated allocation, or the loop that was unrolled. It is verbose and requires comfort reading assembly, so it is a last resort rather than a routine step. - **`-XX:+PrintCompilation` combined with GC logging** separates two confounds: a pause during measurement may be a collection, not a compilation. Enabling unified GC logging alongside lets you attribute outliers correctly. - **`-Xint`** forces interpreted-only execution and **`-XX:-TieredCompilation`** or tier-limiting options change the compilation regime. These are useful as *controls*: running a benchmark interpreted should be dramatically slower, and if it is not, the measured work is probably not happening at all. - **Compiler control and log-compilation output** provide structured, machine-readable detail when the summary flags are not enough, and tooling exists to visualise compilation and inlining decisions from that output. ## Practical discipline 1. Enable the flags on a run *separate* from the one you report numbers from — the printing itself has cost and can perturb timing. 2. Look first for silence during the measurement window; that alone validates warmup. 3. Then confirm the method under test is at the top tier. 4. Then confirm the inlining shape matches your mental model, especially for small hot callees. 5. Treat repeated not-entrant events during measurement as a red flag that the workload is bimodal or that the benchmark's profile is unstable. ## The bigger point These flags convert a benchmark from an opaque number into a claim you can check. The most valuable habit is not memorising the flag names but forming a hypothesis — "this should be compiled at the top tier and this callee should be inlined" — and then confirming it, because a benchmark whose compilation shape you cannot describe is a benchmark whose result you cannot defend.
- What does 'made not entrant' in PrintCompilation output tell you about a benchmark run?It means an existing compiled version of that method was invalidated and will no longer be entered; execution falls back and the method is recompiled. It normally follows a failed speculation — a new receiver type appeared, a previously untaken branch was taken, or a class was loaded that broke an assumption. Seeing it during warmup is expected; seeing it repeatedly during the measurement window means the reported steady state is not stable.
- PrintInlining says a hot callee was 'too big' to inline. What does that change about your benchmark's conclusions?Inlining is what enables most downstream optimization, so a callee that stayed out-of-line did not get its body folded, scalarized or constant-propagated into the caller. Any conclusion that assumed those optimizations applied is unsupported. It also means the measured cost includes real call overhead, which may or may not match production depending on whether the same size and frequency conditions hold there.
- Why run the diagnostic flags in a separate run rather than the one you report?The printing itself is not free — it writes to standard output on the compiler threads and can perturb timing and even scheduling. The diagnostic run answers 'is the compilation shape what I assumed', and the clean run produces the numbers. Mixing them risks reporting a figure influenced by the instrumentation.
saying these in an interview costs you the question
- Assuming a small method is always inlined without checking the inlining decision.
- Reading PrintCompilation output as a performance profile rather than a record of compilation events.
- Ignoring repeated not-entrant entries during the measurement window.
- Enabling PrintInlining without -XX:+UnlockDiagnosticVMOptions and concluding the flag does not exist.
- Reporting timings from the same run that had verbose diagnostics enabled.