skip to content

What is the difference between a sampling profiler and an instrumentation (tracing) profiler in Java?

level: juniorimportance: must knowfreq 70%

answer

  1. Sampling = periodic stack snapshots, statistical, cheap
  2. Instrumentation = bytecode probes per call, exact, heavy + distorting
  3. Hot tiny methods get inflated by instrumentation overhead
  4. Sampling can't give exact call counts
  5. async-profiler / JFR = low-overhead, safepoint-free sampler

basics

~20 s

A sampling profiler peeks at what the program is doing every so often and adds up where it spends most time. An instrumentation profiler adds extra code to every method to count and time every single call exactly, which is more accurate but much slower.

solid answer

~50 s

Sampling and instrumentation are the two ways a profiler figures out where a program spends its time. A sampling profiler periodically (say every 10 ms) interrupts the running threads and records each one's call stack; over many samples, methods that appear most often are the hot ones. It is statistical and approximate but has very low, roughly constant overhead. An instrumentation (tracing) profiler instead injects extra bytecode at method entry/exit so it records an exact event for every call: precise call counts and per-method timings, but with large and uneven overhead because cheap, frequently-called methods get heavily inflated. Rule of thumb: use sampling to find where wall-clock or CPU time goes in a realistic run; use instrumentation when you need exact invocation counts or to trace a specific code path. async-profiler is the modern low-overhead sampler on the JVM.

go deeper

for a junior

Can state the one-line distinction: sampling takes periodic snapshots and is approximate but cheap; instrumentation times every call and is exact but slow.

for a middle

Explains the overhead model (per-sample vs per-call), why instrumentation distorts hot tiny methods, and that sampling can't give exact counts.

for a senior

Chooses the right tool per goal, names async-profiler/JFR as the low-overhead default, and explains accuracy-vs-overhead tradeoffs concretely.

for a principal

Discusses safepoint bias, how instrumentation can defeat JIT inlining, AsyncGetCallTrace, and sets team-wide profiling defaults / production profiling strategy.

## The problem profilers solve When a program is slow, you need to know *where* the time goes — which methods, which call paths. A **profiler** is a tool that measures this. The two fundamental techniques are **sampling** and **instrumentation**. Understanding both, and their tradeoffs, is essential to reading a profile correctly. ## Key terms (defined from scratch) - **Call stack:** the chain of methods currently in progress on a thread. If `main` called `a()` which called `b()`, the stack is `main → a → b`. The method actively running is on top. - **Hot method / hotspot:** code where the program spends a large fraction of its time — the worthwhile thing to optimize. - **Overhead:** the extra cost the profiler itself adds. High overhead distorts the very measurement you are trying to take. - **Bytecode:** the portable instructions a Java `.class` file contains, run by the JVM. It can be rewritten ("instrumented") at load time. ## Sampling profilers A **sampling profiler** does not watch every event. Instead, on a fixed interval (e.g. every 1–10 milliseconds) it **interrupts** the threads and records each thread's current call stack — one *sample*. It does this thousands of times. Then it counts: if method `b()` is on top of the stack in 40% of samples, the profiler estimates `b()` consumed roughly 40% of the CPU/wall time. Properties: - **Statistical / approximate:** it never sees every call; it infers time from the proportion of samples. More samples = better accuracy. - **Low, roughly constant overhead:** the cost is per-sample, not per-method-call, so a billion fast calls cost the same to observe as one slow call. Typical overhead is a few percent. - **Does not distort relative timings** much, because it doesn't add cost inside the methods being measured. - **Limitation — it cannot give exact call counts.** It tells you where time goes, not how many times something was called. ## Instrumentation (tracing) profilers An **instrumentation profiler** (also called a **tracing** profiler) rewrites the program so that *every* method records an event on entry and exit. On the JVM this is usually done by **bytecode injection** — inserting timing/counting probes at method boundaries when classes load. From these events it reconstructs exact call counts and per-method elapsed time. Properties: - **Exact:** precise invocation counts and which method called which. - **High and *uneven* overhead:** every call pays a fixed probe cost. A tiny getter called a million times has its measured time massively inflated, while a method called once is barely affected. This is **measurement distortion** — the profile can promote a method to "hot" purely because the probe overhead dominates its real work. - **Can defeat the JIT:** the inserted code can block inlining and other optimizations the compiler would normally apply, so the instrumented program runs differently from production. ## The safepoint-bias problem (JVM-specific) Many older JVM samplers (e.g. those built on `JVMTI GetAllStackTraces`) can only capture a stack when the thread is at a **safepoint** — a special spot the JIT inserts where it is safe to pause a thread (method returns, loop back-edges, allocations). The JIT may *remove* safepoints from tight, optimized loops. So the sampler can only "see" threads at safepoints, and the recorded stack is whatever safepoint the thread next reached — **not** where it actually was when the timer fired. The result is **safepoint bias**: time gets misattributed to safepoint-bearing methods and the truly hot loop is hidden. This makes naive JVM sampling untrustworthy for exactly the hot code you care about. ## The modern middle ground: async-profiler **async-profiler** is the de-facto modern JVM profiler. It is still a *sampler* (low overhead) but **safepoint-free**: it uses OS signals (`SIGPROF` / `perf_events`) and the JVM's `AsyncGetCallTrace` API to grab the *real* native + Java stack at the instant the timer fires, regardless of safepoints. So it gives you sampling's cheap, undistorted overhead **without** the safepoint-bias error, and can profile CPU, allocations, locks, and wall-clock. **JDK Flight Recorder (JFR)** + JDK Mission Control is the built-in, always-low-overhead alternative. These are why "just use a sampler" is good default advice today. ## When to use which - **Sampling (async-profiler / JFR):** the default. Use to find where CPU/wall-clock time goes in a realistic, production-like run with negligible distortion. - **Instrumentation:** use when you need *exact* call counts, to trace a specific narrow path, or to profile code too short-lived to sample well — accepting that absolute timings are unreliable. ## Deriving your own answer From the above: sampling = periodic stack snapshots, cheap, statistical, no exact counts, watch out for safepoint bias on the JVM; instrumentation = per-call probes via bytecode, exact counts, expensive and distorting especially for hot tiny methods. The JVM-savvy answer adds *async-profiler/JFR as the safepoint-free sampling sweet spot.*

  • Which technique would you use to find out exactly how many times a method was called?
    Instrumentation — sampling is statistical and only estimates where time goes, so it cannot give exact invocation counts.
  • Why does instrumentation tend to make small, frequently-called methods look disproportionately hot?
    Each call pays a fixed probe cost on entry/exit. For a tiny method that does almost no real work, that fixed overhead dwarfs the actual cost and inflates its measured time.

saying these in an interview costs you the question

  • Claiming instrumentation is 'more accurate' without noting it distorts timings of hot methods
  • Believing sampling gives exact invocation counts
  • Thinking sampling overhead grows with the number of method calls (it's per-sample, not per-call)
  • Assuming all JVM samplers are safepoint-free (older JVMTI-based ones are not)

context