skip to content

You run one service image both as a long-lived server and as short-lived batch jobs and serverless invocations. How would you approach just-in-time warmup as a platform-level concern across that fleet?

level: principalimportance: nice to knowfreq 26%

answer

  1. Does the process live long enough to repay compilation?
  2. Servers: full tiering + traffic ramp + compiler CPU headroom
  3. Short-lived: TieredStopAtLevel=1 + class-data sharing
  4. Readiness check must mean warm, not just alive
  5. Synthetic warmup can poison type profiles

basics

~20 s

Match the compiler policy to process lifetime. Long-lived processes keep full tiering and get traffic-shaped warmup before receiving load. Short-lived ones stop at the fast compiler tier and lean on class-data sharing or cached profiles, because they will never repay optimizing-compiler cost.

solid answer

~50 s

Treat warmup as a per-workload decision with a shared default. **Long-lived servers**: keep full tiered compilation. The engineering work is in the deployment path, not flags - health checks that do not admit traffic until the process has executed representative work, gradual traffic ramping or slow-start load balancing, and capacity for a first-minute throughput dip during rolling deploys. Also give compiler threads CPU headroom; tight container limits stretch warmup badly. **Short-lived processes** (CLI, batch steps, function invocations): full tiering is a bad trade - you pay optimizing-compiler CPU for code that stops running. `-XX:TieredStopAtLevel=1` gets native code fast with a low ceiling you never reach anyway. Combine with class-data sharing to cut class-loading cost, which usually dominates in such processes. **Cross-cutting**: shrink the startup work itself, since interpreted startup code is where the time goes. Measure the actual distribution of process lifetimes before choosing, and validate with real latency percentiles per deploy rather than steady-state microbenchmarks.

code

text · 5 lines
text
# long-lived server: default tiering, headroom for compiler threads, watched code cache
-XX:ReservedCodeCacheSize=256m -XX:+UnlockDiagnosticVMOptions -XX:+PrintCompilation   # diagnosis only

# short-lived job or function invocation
-XX:TieredStopAtLevel=1 -XX:SharedArchiveFile=app.jsa -Xshare:auto

go deeper

for a junior

Know that JVMs are slow at first and fast later, and that very short-lived programs never reach peak speed.

for a middle

Be able to name the lever for short-lived processes (stop at the fast compiler tier) and explain why it is the right trade there.

for a senior

Cover the operational side: readiness that means warm, traffic ramping, compiler CPU contention in small containers, and code-cache monitoring.

for a principal

Present it as a fleet policy problem - workload classes with reviewed option sets, evidence required to deviate, capacity planning for rolling deploys, and the risk that synthetic warmup poisons profiles.

## Frame the problem correctly Warmup is not a flag question, it is an economics question: **does this process live long enough to repay the compiler investment?** The optimizing compiler spends CPU now to save CPU later. If "later" never arrives, the spend is pure loss - and worse, it competes with the application during exactly the window where latency is most visible. So the first step is data: what is the distribution of process lifetimes and request volumes across your fleet? A single answer for "a service image" is almost always wrong when the same image runs a 200 ms function and an eight-hour server. ## Long-lived servers Keep full tiered compilation. It is the default because it wins here. The real work is operational: - **Do not admit traffic to a cold process.** A readiness check that only pings a health endpoint declares a JVM ready while its request path is still interpreted. Better: exercise representative code paths before signalling ready, or accept traffic behind a **slow start / gradual ramp** at the load balancer so a cold instance receives a fraction of its steady share initially. - **Budget for the deploy dip.** During a rolling restart, part of the fleet is warming. If capacity is planned at steady-state throughput, deploys become brownouts. Either over-provision during rollouts or slow the rollout. - **Give compiler threads CPU.** On a container with one or two CPUs, compiler threads contend directly with request handling and compiler-queue backlog inflates thresholds, so warmup takes far longer than on a developer laptop. This is one of the most common causes of "it is fast in staging, slow in prod". - **Watch code cache.** Large applications with tiering can exhaust the reserved code cache; when that happens compilation shuts off and the process silently degrades toward interpreted speed. Monitor it rather than discovering it from a latency graph. ## Short-lived processes Here the objective flips to time-to-finish, and most of it is not JIT at all - class loading, verification and initialization usually dominate. - **`-XX:TieredStopAtLevel=1`** keeps the fast compiler and drops profiling and the optimizing tier. Code becomes native quickly, compiler CPU drops sharply, code-cache footprint shrinks. The peak-throughput ceiling you give up is one you were never going to reach. - **Class-data sharing** (an archive of parsed class metadata, including application classes) removes a large fraction of startup cost and is usually a bigger win than any compiler flag for very short runs. - **Cut the startup path.** Lazy initialization, fewer classes loaded, less reflection and less classpath scanning shorten the interpreted phase directly. Framework-level startup cost tends to dwarf compiler policy here. - Newer platform work in this space - archived training profiles and ahead-of-time caches of loaded classes and compilation decisions - moves in the same direction: reuse work from a previous run instead of rediscovering it. Evaluate against your own JDK version rather than assuming availability. ## Choosing a default and letting workloads override A workable platform stance: 1. **Default to full tiering** in the base image, since getting it wrong on a server is more expensive than getting it wrong on a batch job. 2. **Expose a workload class** (`server` / `short-lived`) that maps to a small, reviewed set of JVM options, rather than letting every team hand-tune flags. Flag sprawl is a long-term maintenance cost and most hand-tuned flag sets are cargo-culted. 3. **Require evidence to deviate.** The measurement that matters is end-to-end: for servers, latency percentiles over the first N minutes after deploy; for short-lived jobs, wall-clock time to completion across many real invocations. Steady-state microbenchmarks answer neither question. ## Anti-patterns to name - **Synthetic warmup loops** that hammer a fake payload. They can compile the wrong shapes, poison the type profile with receivers the real traffic never sends, and trigger deoptimization when real traffic arrives. Warm with representative work or not at all. - **`-Xcomp`** as a "pre-warm" trick. It compiles everything on first call without a good profile: slower startup and worse code. - **Disabling tiered compilation** to "get to C2 faster" - it does the opposite. - **Treating warmup as a JVM problem only**, when the dominant term is often framework initialization and connection-pool or cache priming.

  • A team proposes a warmup loop that replays synthetic requests at startup before the instance is marked ready. What would you check before approving it?
    Whether the synthetic traffic is representative in shape, not just volume. If it exercises different receiver types, branch directions or payload sizes than production, the profile it builds is misleading - the optimizing compiler will speculate on the wrong facts, and real traffic will trigger deoptimization and recompilation, which can be worse than starting cold. I would prefer shadowing a slice of real traffic, or a load-balancer slow start, over synthetic replay.
  • How would you tell whether a post-deploy latency spike is warmup or something else?
    Warmup has a characteristic signature: the spike is worst on the first requests, decays over seconds to minutes on a curve, appears on every new instance, and correlates with compilation activity and compiler-thread CPU. Connection-pool and cache priming look similar but resolve differently, so I would separate them by observing whether the decay tracks compilation counts, and by testing an instance that is warm but has a cold cache.

It is the difference between a delivery van and a race car: sinking money into a tuned engine only pays if you are going to drive it hard for hours, and a vehicle doing five-minute errands is better off just starting reliably.

saying these in an interview costs you the question

  • Applying one flag set to every workload regardless of process lifetime
  • Using -Xcomp to 'pre-warm' a service
  • Marking a JVM ready as soon as the port is open and calling warmup solved
  • Assuming compiler behaviour on a laptop predicts behaviour in a one-CPU container
  • Blaming the JIT for startup cost that is really class loading and framework initialization

context