Why does a tiered runtime start a method in an interpreter and compile it only after invocation and loop counters cross a threshold?
answer
- most methods stay cold
- spend compiler time where it repays
- counters, not guesses
- entries plus loop back-edges
- the low tier profiles for the high one
basics
~20 sCompilation costs time, CPU and memory, and most methods run only a handful of times. Counting invocations and loop iterations finds the small hot fraction where optimisation will be repaid, and the interpreted phase supplies the profile the compiler then optimises against.
solid answer
~50 sCompiling is not free: an optimising compiler runs inside the same process, competing for the same cores, and its output occupies memory. Meanwhile the execution profile of a real program is extremely skewed - a large share of methods run once or twice, and a small set of loops and call paths carry most of the work. So the runtime starts everything in an interpreter, which begins instantly, and attaches two counters to each method: how often it is entered, and how often a loop inside it goes round. When the sum crosses a threshold, the method is queued for compilation, usually first at a cheap tier that produces decent code quickly and continues collecting observations, and later at an expensive tier for the few methods that stay hot. Effort is thereby spent in proportion to how much execution is left to accelerate.
go deeper
Recall that a runtime which compiles while it runs does not compile everything: it starts by interpreting and only compiles the parts that turn out to run a lot.
Explain the two counters - method entries and loop back-edges - what each is a proxy for, why both are needed, and why the lowest tier is also the place where profile data is gathered.
Show what the tiering means operationally: compilation competing for CPU during a traffic ramp, hot-but-unimportant code being optimised ahead of latency-critical code, and compilation decisions differing between two runs of the same workload.
Own the trade-off in the threshold itself - reaching peak sooner against wasting compiler budget on code that cools - and decide when a workload's shape justifies moving away from profile-driven compilation altogether.
## The economics the design is answering A runtime that compiles while it runs is spending a budget it must earn back. Every method it compiles costs compiler CPU taken from the application, memory for the produced code, and a delay before that code is available. The payoff is the difference in speed between the compiled form and the interpreted form, multiplied by **how much of that method's execution is still ahead of it**. That last quantity is the one nobody knows in advance. A program's execution is heavily skewed: initialisation, configuration and error paths often run once, while a handful of loops and call paths run continuously. Compiling everything therefore spends most of the budget on code that will never repay it, and spends it at the worst possible moment - startup, when the application is already doing its most work-per-second. ## What the counters actually estimate Each method carries counters that are cheap to bump from interpreted code: - **An invocation counter**, incremented when the method is entered. It tracks repetition across calls. - **A loop back-edge counter**, incremented each time control jumps backwards to the top of a loop inside the method. It tracks repetition *within* one call. Neither counter measures size, nesting or complexity. They are a proxy for one prediction: *code that has run a lot is likely to keep running*. The threshold is where the runtime judges that prediction strong enough to bet compiler time on it. Both counters are needed because the two kinds of repetition are independent - a method called ten thousand times with no loop, and a method called once containing a loop of ten million iterations, are both hot, and only one of them is visible to an invocation counter. ## Why the interpreter is not merely a fallback The lowest tier does a second job that the compiler cannot do for itself: it **observes**. While code runs there, the runtime can record which branches are taken, which types actually appear at a call site, which array indices stay in range, and which paths are never reached at all. That record is the raw material for the aggressive optimisation the top tier performs. A compiler invoked at startup would have none of it and could only optimise for what is provable from the program text. This is why running interpreted first is a deliberate investment rather than a delay to be eliminated. The interpreted phase buys knowledge that no amount of static analysis can supply. ## The tiers as a ladder | Tier | Starts | Code quality | Cost to produce | Its other job | |---|---|---|---|---| | Interpreter | immediately | lowest | none | collect counters and profile | | Quick compiler | after a low threshold | good | small | keep profiling while running faster | | Optimising compiler | after a high threshold | highest | large | speculate on the profile collected below | Two details follow from the ladder. First, a method may exist in several forms at once, and calls made before compilation finishes keep using the older form. Second, promotion is not one-way: if an assumption the top tier made stops holding, execution can fall back down the ladder. ## Consequences you can observe 1. **Two runs of the same workload need not compile the same set of methods.** Thresholds are crossed in a particular order that depends on timing, on the input, and on how compiler threads were scheduled, so compilation decisions are not perfectly reproducible. 2. **Hot code is not the same as important code.** A method on the latency-critical path that runs rarely may stay interpreted indefinitely, while a busy background loop gets the expensive treatment. 3. **Compilation competes with the application.** Compiler work runs on the same machine; during heavy compilation a service has fewer cores available for requests, which is part of why early throughput is depressed rather than merely un-optimised. 4. **Thresholds are a tuning surface, not a constant.** Lower them and you reach peak sooner but waste effort on code that cools; raise them and you waste less but stay slow longer. Runtimes differ in where they set them and whether they adapt them, and the trade-off is genuinely two-sided. ## What this is not It is not lazy compilation for its own sake, and it is not an admission that the compiler is slow. It is a scheduling decision: given a fixed budget of compiler time inside a live process, spend it where the remaining execution is greatest, and use the time before spending it to learn something worth optimising for.
- A method containing one long-running loop is entered exactly once, so its invocation counter never rises - how does it still get compiled?The loop's back-edge counter rises instead, and when it crosses the threshold the runtime compiles the method and then swaps the still-running activation over to the compiled version at the loop head, carrying the live local values across. Replacing the executing frame in place is what lets an already-started loop benefit, rather than only the next call.
- Besides somewhere to run, what does the interpreted phase give the optimising compiler?A profile: which branches are taken and how often, which types actually reach each call site, which paths were never executed. The top tier optimises against observed behaviour rather than only what the program text proves, which is how it can specialise code that static analysis would have to keep general.
- What goes wrong if the thresholds are set very low?The runtime compiles code that was merely warm, so compiler threads take CPU from the application and the code cache fills with methods that soon stop running. Startup gets slower and less predictable, and the expensive tier may be reached on the basis of a thin profile, producing speculation that is invalidated shortly afterwards.
saying these in an interview costs you the question
- Says the runtime compiles every method before execution begins
- Thinks the interpreter only handles code the compiler cannot translate
- Believes compiling everything up front would strictly be faster
- Assumes a threshold measures method size rather than execution count
- Forgets that compilation consumes CPU the application also needs
- Thinks only invocation counts matter, missing loop repetition