HotSpot ships two just-in-time compilers, historically called the client (C1) and server (C2) compilers. How do they differ in design goals and output, and why does a single JVM contain both?
answer
- C1 = fast compile, light optimization, linear-scan registers
- C2 = sea of nodes, speculation, graph colouring, slow
- C1 also collects the profile C2 consumes
- C2-only = slow warmup; C1-only = low ceiling
- Tiering: C1 buys time and data, C2 spends them
basics
~20 sC1 compiles fast and optimizes lightly, so code becomes native quickly with modest peak speed. C2 optimizes aggressively and compiles slowly, giving the best steady-state throughput. A single JVM uses both: C1 for quick warmup and profiling, C2 for the hottest methods.
solid answer
~50 sThey sit at opposite ends of the same tradeoff. **C1** is a fast, low-latency compiler. It does a single-pass-style translation with cheap local optimizations - constant folding, simple inlining, register allocation via a linear-scan allocator. Compile time is short, so hot code becomes native within milliseconds, but the generated code is maybe two to three times slower than C2's. **C2** is an aggressive optimizing compiler built on a sea-of-nodes intermediate representation with global value numbering, loop transformations, deep profile-guided inlining, escape analysis and graph-colouring register allocation. It can take orders of magnitude longer per method and consumes far more memory, but produces the fastest code. The JVM contains both because neither alone is good enough: C2-only means a long slow warmup while nothing is compiled; C1-only means you never reach peak throughput. Tiered compilation runs C1 first with instrumentation, then re-compiles the survivors with C2 using the profile C1 collected.
code
text · 4 lines-XX:+TieredCompilation # default: C1 then C2, driven by profiles
-XX:-TieredCompilation # C2 only - long interpreted warmup, rarely a good idea
-XX:TieredStopAtLevel=1 # C1 only, no profiling counters - fast start, low ceiling
-XX:CICompilerCount=4 # total background compiler threads (split across C1 and C2)go deeper
Know that there are two JIT compilers, one fast and rough, one slow and highly optimizing, and that modern JVMs use both.
Name concrete differences - compile time, optimization sets, register allocation - and explain that C1 collects the profile that C2 needs.
Discuss what tiering costs: compiler-thread CPU, code-cache footprint, and when restricting to C1 is defensible for short-lived or memory-constrained processes.
Treat the pair as an adaptive-optimization policy: where speculation pays off, what a bad profile costs, and how fleet-wide warmup strategy interacts with the compiler configuration you standardize on.
## The same tradeoff, two answers Every optimizing compiler trades compile time for code quality. In an ahead-of-time toolchain you can afford minutes. Inside a running JVM the compiler competes with the application for CPU, so the tradeoff must be made explicitly - and HotSpot makes it twice, with two compilers. ## C1 - the fast compiler C1 (historically the "client" compiler, because desktop applications care about startup) is built for short compile times. It parses bytecode into a straightforward high-level IR, applies cheap and mostly local optimizations - constant folding, dead-code elimination, null-check elimination, method inlining of small methods, simple loop handling - lowers to a low-level IR, and allocates registers with **linear scan**, an allocator chosen precisely because it is near-linear in time rather than optimal. The result is compiled within microseconds to a few milliseconds per method. Its code is far better than interpretation (commonly several times faster) yet noticeably behind C2. Crucially, C1 can also emit **instrumented** code: versions that keep counting invocations, back-edges, branch outcomes and observed receiver types. That instrumentation costs some throughput but produces the profile data the next tier needs. ## C2 - the optimizing compiler C2 (the "server" compiler) is a classic heavyweight optimizer. It builds a **sea-of-nodes** graph in which control and data dependencies are unified, then runs global analyses: global value numbering, loop unrolling and peeling, range-check elimination, escape analysis (which can remove allocations or locks entirely), aggressive profile-driven inlining across many levels, and graph-colouring register allocation. Most of C2's power comes from **speculation**. Because it is fed a profile, it can assume the observed behaviour is the real behaviour: a call site that only ever saw one receiver type is compiled as a direct, inlinable call guarded by a type check; a branch never taken is replaced by an uncommon trap. If a guard fails at runtime, the frame is deoptimized back to the interpreter and the method is eventually recompiled. That leverage is exactly why C2 needs a profile and therefore why it wants to run *after* something else has been executing the method. The cost is real: C2 compilations can be milliseconds to hundreds of milliseconds each, and use significantly more compiler-thread CPU and memory. ## Why one JVM needs both Consider the two single-compiler worlds: - **C2 only** (old `-server` mode). Nothing is compiled until a method is very hot, so the application interprets for a long time, and the first compilations arrive late. Warmup is slow, and the profile is gathered by the interpreter, which is a comparatively coarse profiler and slow while doing it. - **C1 only** (old `-client` mode, and what `-XX:TieredStopAtLevel=1` gives you today). Startup is excellent and peak throughput plateaus well below what the machine can do. **Tiered compilation** composes them. Execution starts interpreted, moves quickly into C1 code (fast to produce, immediately much faster than interpreting), gathers a high-quality profile there far more cheaply than the interpreter could, and promotes only the genuinely hottest methods to C2 using that profile. You get near-C1 warmup with C2 steady state, and the profiling burden is carried by code that is already fast. A useful mental model: **C1 buys time and buys data; C2 spends both.** ## Practical notes - Tiered compilation is the default and is normally the right choice. `-XX:-TieredCompilation` reverts to C2-only behaviour and is rarely justified. - Both compilers run on background threads, sized by `-XX:CICompilerCount`; with tiering the pool is split between C1 and C2 threads, so on tiny containers you may see compilation starved for CPU. - Because both compilers emit into the same code cache and the same method can hold C1 and later C2 versions, tiering increases code-cache pressure relative to a single-compiler configuration. - "Client" and "server" are historical names; modern JVMs choose configuration by ergonomics and tiering, not by a `-client`/`-server` switch (the client VM is gone on mainstream 64-bit builds).
- If C2 produces better code, why not send every method straight to C2 and skip C1 entirely?Two reasons. C2 compilation is expensive in time and memory, so compiling cold methods wastes CPU that the application needs, and the queue delay would leave hot code interpreted for longer. Just as important, C2's biggest wins are speculative and require a profile; without one it must compile conservatively and produces markedly worse code.
- What is lost when you run with -XX:TieredStopAtLevel=1?You get C1-quality code only: no escape analysis, no deep speculative inlining, no aggressive loop optimization, so steady-state throughput is typically a few times lower than full tiering. In exchange you get faster warmup, lower compiler CPU usage and a much smaller code-cache footprint, which can be the right trade for short-lived processes.
C1 is a quick sketch artist who gives you a usable drawing in a minute and takes notes about what matters; C2 is the oil painter who works from those notes for hours and produces the piece you hang on the wall.
saying these in an interview costs you the question
- Thinking -client and -server select different JVMs on modern 64-bit builds
- Claiming C1 exists only for GUI or desktop applications
- Saying C2 is always faster overall, ignoring its compile-time and warmup cost
- Believing C2 can optimize aggressively without a profile
- Assuming disabling tiered compilation improves performance