skip to content

In VisualVM, what is the difference between sampling and instrumenting (profiling) CPU, and when do you use each?

level: middleimportance: should knowfreq 50%

answer

  1. Sampling = periodic stack snapshots, cheap, approximate
  2. Instrumenting = bytecode probes per call, precise, heavy, distorting
  3. Sample first, instrument a narrow set second
  4. Instrumenting skews cheap-but-hot methods and can disable inlining

basics

~20 s

Sampling periodically takes a snapshot of what each thread is running and adds up where time is spent — cheap but approximate. Instrumenting modifies the code to record every method call — precise but slow and it can skew the numbers.

solid answer

~50 s

Both measure where CPU time goes, but differently. Sampling pauses briefly at a fixed interval (say every 10–20 ms), grabs each thread's current stack, and statistically attributes time to whatever methods appear most often. It has low, roughly constant overhead and is safe even on busy systems, but it can miss short-lived methods and is approximate. Instrumenting (full profiling) injects measurement code into method entry and exit so every call and its exact duration is counted. That gives precise call counts and timings, but the injected code adds significant overhead, can distort relative timings (cheap methods called millions of times balloon), and slows the app substantially. Use sampling first to find hot areas with little disruption, then optionally instrument a narrowed set of classes for exact call counts. Avoid broad instrumenting on production or latency-sensitive systems.

go deeper

for a junior

Knows sampling is cheap/approximate and instrumenting is precise/heavy, and that you'd start with sampling.

for a middle

Explains the mechanism (stack snapshots vs bytecode probes), the overhead trade-off, and the sample-then-instrument workflow.

for a senior

Discusses distortion of cheap-but-hot methods, JIT/inlining interference, and why broad instrumenting is unsafe in prod; narrows instrumenting to specific classes.

for a principal

Knows safepoint bias and its limits, when to abandon VisualVM for async-profiler/JFR, and sets guidance on profiling methodology and acceptable production overhead.

## What CPU profiling means **Profiling** means measuring where a program spends its time so you can optimise the parts that matter. CPU profiling specifically answers: *which methods consume the most processor time?* Optimising anything else is wasted effort, so this is the foundation of performance work. VisualVM offers two fundamentally different techniques. ## Sampling A **sampler** wakes up on a timer (e.g. every 10 or 20 milliseconds), and for each thread it records the **call stack** — the chain of method calls currently executing (method A called B called C…). It does this thousands of times and then tallies: if method C appears in 30% of the samples, it's estimated to be using ~30% of CPU time. - **Overhead:** low and roughly fixed (just periodic stack snapshots), independent of how many calls your code makes. - **Accuracy:** statistical/approximate. A method that runs for less than one sampling interval may never be captured. The longer you sample, the more accurate the picture. - **Distortion:** minimal — the program runs almost normally, so the *relative* timings are realistic. - **Bias caveat:** classic samplers can only snapshot threads at JVM **safepoints** (points where the JVM can safely stop a thread). This can bias results toward methods that sit near safepoints. (Lower-level tools like async-profiler avoid this; VisualVM's sampler is safepoint-based.) ## Instrumenting (a.k.a. profiling / tracing) **Instrumenting** rewrites the **bytecode** (the compiled form the JVM runs) of the classes you profile, inserting tiny timing/counting code at every method **entry** and **exit**. Now every single call is recorded exactly. - **Overhead:** high and proportional to call frequency — a trivial getter called ten million times now also runs the injected probe ten million times. - **Accuracy:** exact **call counts** and measured durations. - **Distortion:** significant. The probe cost is large relative to cheap methods, so a method that's actually trivial can look expensive (its measured time is dominated by measurement). Inlining and JIT optimisations the compiler would normally apply may also be suppressed, changing behaviour. - **Footprint:** you usually restrict it to specific packages/classes to keep it bearable; instrumenting everything can make the app crawl. ## The JIT wrinkle The JVM **JIT (Just-In-Time) compiler** optimises hot code at runtime, including **inlining** (folding a small method's body into its caller so there's no call at all). Instrumenting can prevent or change inlining decisions, so the profiled run may not behave like the real run. Sampling, by touching nothing, preserves the real optimisation behaviour better. ## Practical workflow 1. **Start with sampling.** It's cheap, safe, and tells you the hot region (the 'where', roughly). 2. **If you need exact call counts** for a narrow set of methods (e.g. 'is this called once or a thousand times?'), instrument *just those classes*. 3. **Never broadly instrument production** or a latency-SLA service — the overhead and distortion make conclusions unreliable and can harm users. Prefer sampling, or move to dedicated low-overhead tooling like JFR. ## Summary table | | Sampling | Instrumenting | |---|---|---| | Mechanism | Periodic stack snapshots | Bytecode probes on entry/exit | | Overhead | Low, ~constant | High, scales with call count | | Accuracy | Approximate (% of time) | Exact call counts/durations | | Distortion | Minimal | Can be large (cheap methods skew) | | Use when | First pass, busy systems | Narrow, need exact counts |

  • Why can a trivial method look expensive under instrumenting profiling?
    The injected entry/exit probe runs on every call. For a method that's intrinsically cheap but called millions of times, the probe overhead dominates the measured time, inflating its apparent cost relative to reality.
  • What is a 'safepoint' and why does it matter for VisualVM's sampler?
    A safepoint is a point where the JVM can safely pause a thread. VisualVM's sampler only snapshots at safepoints, so time can be misattributed to methods near safepoints (safepoint bias). Tools like async-profiler sample outside safepoints to avoid this.

saying these in an interview costs you the question

  • Saying instrumenting is 'more accurate so always better' — its overhead distorts timings
  • Believing sampling captures every method call (it can miss short ones)
  • Recommending broad instrumenting on a production/latency-sensitive service
  • Not knowing instrumenting works by modifying bytecode

context