skip to content

Sampling vs Instrumentation Profilers

Sampling profilers snapshot stacks periodically for low overhead but suffer safepoint bias; instrumenting profilers count every call exactly but distort what they measure. Interviewers ask which you would attach to production, and async-profiler is the modern middle ground.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

5

What is the difference between a sampling profiler and an instrumentation (tracing) profiler in Java?

level: juniorimportance: must knowfreq 70%

answer

  1. Sampling = periodic stack snapshots, statistical, cheap
  2. Instrumentation = bytecode probes per call, exact, heavy + distorting
  3. Hot tiny methods get inflated by instrumentation overhead
  4. Sampling can't give exact call counts
  5. async-profiler / JFR = low-overhead, safepoint-free sampler

basics

~20 s

A sampling profiler peeks at what the program is doing every so often and adds up where it spends most time. An instrumentation profiler adds extra code to every method to count and time every single call exactly, which is more accurate but much slower.

solid answer

~50 s

Sampling and instrumentation are the two ways a profiler figures out where a program spends its time. A sampling profiler periodically (say every 10 ms) interrupts the running threads and records each one's call stack; over many samples, methods that appear most often are the hot ones. It is statistical and approximate but has very low, roughly constant overhead. An instrumentation (tracing) profiler instead injects extra bytecode at method entry/exit so it records an exact event for every call: precise call counts and per-method timings, but with large and uneven overhead because cheap, frequently-called methods get heavily inflated. Rule of thumb: use sampling to find where wall-clock or CPU time goes in a realistic run; use instrumentation when you need exact invocation counts or to trace a specific code path. async-profiler is the modern low-overhead sampler on the JVM.

go deeper

for a junior

Can state the one-line distinction: sampling takes periodic snapshots and is approximate but cheap; instrumentation times every call and is exact but slow.

for a middle

Explains the overhead model (per-sample vs per-call), why instrumentation distorts hot tiny methods, and that sampling can't give exact counts.

for a senior

Chooses the right tool per goal, names async-profiler/JFR as the low-overhead default, and explains accuracy-vs-overhead tradeoffs concretely.

for a principal

Discusses safepoint bias, how instrumentation can defeat JIT inlining, AsyncGetCallTrace, and sets team-wide profiling defaults / production profiling strategy.

## The problem profilers solve When a program is slow, you need to know *where* the time goes — which methods, which call paths. A **profiler** is a tool that measures this. The two fundamental techniques are **sampling** and **instrumentation**. Understanding both, and their tradeoffs, is essential to reading a profile correctly. ## Key terms (defined from scratch) - **Call stack:** the chain of methods currently in progress on a thread. If `main` called `a()` which called `b()`, the stack is `main → a → b`. The method actively running is on top. - **Hot method / hotspot:** code where the program spends a large fraction of its time — the worthwhile thing to optimize. - **Overhead:** the extra cost the profiler itself adds. High overhead distorts the very measurement you are trying to take. - **Bytecode:** the portable instructions a Java `.class` file contains, run by the JVM. It can be rewritten ("instrumented") at load time. ## Sampling profilers A **sampling profiler** does not watch every event. Instead, on a fixed interval (e.g. every 1–10 milliseconds) it **interrupts** the threads and records each thread's current call stack — one *sample*. It does this thousands of times. Then it counts: if method `b()` is on top of the stack in 40% of samples, the profiler estimates `b()` consumed roughly 40% of the CPU/wall time. Properties: - **Statistical / approximate:** it never sees every call; it infers time from the proportion of samples. More samples = better accuracy. - **Low, roughly constant overhead:** the cost is per-sample, not per-method-call, so a billion fast calls cost the same to observe as one slow call. Typical overhead is a few percent. - **Does not distort relative timings** much, because it doesn't add cost inside the methods being measured. - **Limitation — it cannot give exact call counts.** It tells you where time goes, not how many times something was called. ## Instrumentation (tracing) profilers An **instrumentation profiler** (also called a **tracing** profiler) rewrites the program so that *every* method records an event on entry and exit. On the JVM this is usually done by **bytecode injection** — inserting timing/counting probes at method boundaries when classes load. From these events it reconstructs exact call counts and per-method elapsed time. Properties: - **Exact:** precise invocation counts and which method called which. - **High and *uneven* overhead:** every call pays a fixed probe cost. A tiny getter called a million times has its measured time massively inflated, while a method called once is barely affected. This is **measurement distortion** — the profile can promote a method to "hot" purely because the probe overhead dominates its real work. - **Can defeat the JIT:** the inserted code can block inlining and other optimizations the compiler would normally apply, so the instrumented program runs differently from production. ## The safepoint-bias problem (JVM-specific) Many older JVM samplers (e.g. those built on `JVMTI GetAllStackTraces`) can only capture a stack when the thread is at a **safepoint** — a special spot the JIT inserts where it is safe to pause a thread (method returns, loop back-edges, allocations). The JIT may *remove* safepoints from tight, optimized loops. So the sampler can only "see" threads at safepoints, and the recorded stack is whatever safepoint the thread next reached — **not** where it actually was when the timer fired. The result is **safepoint bias**: time gets misattributed to safepoint-bearing methods and the truly hot loop is hidden. This makes naive JVM sampling untrustworthy for exactly the hot code you care about. ## The modern middle ground: async-profiler **async-profiler** is the de-facto modern JVM profiler. It is still a *sampler* (low overhead) but **safepoint-free**: it uses OS signals (`SIGPROF` / `perf_events`) and the JVM's `AsyncGetCallTrace` API to grab the *real* native + Java stack at the instant the timer fires, regardless of safepoints. So it gives you sampling's cheap, undistorted overhead **without** the safepoint-bias error, and can profile CPU, allocations, locks, and wall-clock. **JDK Flight Recorder (JFR)** + JDK Mission Control is the built-in, always-low-overhead alternative. These are why "just use a sampler" is good default advice today. ## When to use which - **Sampling (async-profiler / JFR):** the default. Use to find where CPU/wall-clock time goes in a realistic, production-like run with negligible distortion. - **Instrumentation:** use when you need *exact* call counts, to trace a specific narrow path, or to profile code too short-lived to sample well — accepting that absolute timings are unreliable. ## Deriving your own answer From the above: sampling = periodic stack snapshots, cheap, statistical, no exact counts, watch out for safepoint bias on the JVM; instrumentation = per-call probes via bytecode, exact counts, expensive and distorting especially for hot tiny methods. The JVM-savvy answer adds *async-profiler/JFR as the safepoint-free sampling sweet spot.*

  • Which technique would you use to find out exactly how many times a method was called?
    Instrumentation — sampling is statistical and only estimates where time goes, so it cannot give exact invocation counts.
  • Why does instrumentation tend to make small, frequently-called methods look disproportionately hot?
    Each call pays a fixed probe cost on entry/exit. For a tiny method that does almost no real work, that fixed overhead dwarfs the actual cost and inflates its measured time.

saying these in an interview costs you the question

  • Claiming instrumentation is 'more accurate' without noting it distorts timings of hot methods
  • Believing sampling gives exact invocation counts
  • Thinking sampling overhead grows with the number of method calls (it's per-sample, not per-call)
  • Assuming all JVM samplers are safepoint-free (older JVMTI-based ones are not)

context

open as a page

You suspect a Java service is slow under production-like load. Walk through how you'd choose between a sampling and an instrumentation profiler and which tools you'd reach for.

level: middleimportance: must knowfreq 55%

basics

~20 s

Start with a low-overhead sampling profiler (like async-profiler or JFR) under realistic load to see where time actually goes, because it barely slows the app. Only switch to an instrumentation profiler if you need exact call counts or to trace one specific path, knowing it's slower and can distort results.

open as a page

Why does a sampling profiler have roughly constant overhead while an instrumentation profiler's overhead scales with how often methods are called?

level: middleimportance: should knowfreq 50%

basics

~20 s

Sampling only does work once per timer tick (e.g. every 10 ms), no matter how busy the code is, so its cost is fixed. Instrumentation does extra work on every single method call, so the more calls there are, the more it costs.

open as a page

What is safepoint bias in JVM sampling profilers, and how do modern profilers like async-profiler avoid it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Older JVM samplers can only grab a thread's stack when the thread reaches a 'safepoint' — a safe pause spot. The JIT removes safepoints from tight loops, so the sampler records the wrong place and blames safepoint-friendly methods. async-profiler avoids this by capturing the stack instantly via OS signals, no safepoint needed.

open as a page

Beyond raw overhead, how can an instrumentation profiler change a program's behavior so much that its results misdirect optimization, and what would you do about it at scale?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Adding probes to every method can stop the JIT from optimizing (especially inlining) small methods, so the profiled program runs differently from production and points you at the wrong 'hot' code. To avoid being misled, prefer low-overhead sampling, scope any instrumentation narrowly, and always validate findings against a sampled, optimized run.

open as a page