You suspect a Java service is slow under production-like load. Walk through how you'd choose between a sampling and an instrumentation profiler and which tools you'd reach for.
answer
- Sample first (async-profiler / JFR), under realistic load
- Match dimension: CPU vs wall-clock vs lock vs allocation
- Flame graph: read by width
- Instrument only for exact counts / narrow trace, in non-prod
- Sampling = production-safe; instrumentation = controlled env
basics
~20 sStart with a low-overhead sampling profiler (like async-profiler or JFR) under realistic load to see where time actually goes, because it barely slows the app. Only switch to an instrumentation profiler if you need exact call counts or to trace one specific path, knowing it's slower and can distort results.
solid answer
~50 sMy default for a production-like investigation is **sampling**, because it adds only a few percent overhead and doesn't distort hot methods — so the profile reflects reality. I'd attach **async-profiler** (or enable **JFR**, which is built-in and production-safe) to a representative load and produce a **flame graph** to find the dominant call paths. I'd profile the right resource: CPU sampling for compute-bound code, but wall-clock/off-CPU or lock/allocation profiling if the latency is actually I/O, lock contention, or GC pressure rather than CPU. Only if I then need *exact invocation counts* — say to confirm a method is called far more than expected, or to trace a precise narrow path — would I reach for **instrumentation**, and I'd do it in a controlled environment, not production, because of its heavy, uneven overhead and the risk it defeats JIT inlining. In short: sample first to localize, instrument selectively to confirm counts.
go deeper
Knows to use a profiler and that sampling is the lighter-weight default.
Picks sampling first for production-like load, names async-profiler/JFR, and knows instrumentation is for exact counts.
Matches the profiling dimension (CPU/wall-clock/lock/allocation) to the symptom, reads flame graphs correctly, and reasons about overhead/fidelity tradeoffs.
Designs the org's profiling strategy: always-on JFR in prod, escalation playbooks, avoiding safepoint-biased tools, and tying profiling into perf regression gates.
## The decision framework The right profiler depends on the **question you're asking** and the **environment**. Here's the reasoning a strong engineer applies. ### Step 1 — Reproduce under realistic load Profiling an idle or toy workload is misleading; the JIT compiles different code, caches behave differently, and contention only appears under concurrency. So first get a **production-like load** (real traffic shadow, load test, or profile in prod with a safe tool). ### Step 2 — Default to sampling, because of overhead and fidelity - **Overhead:** sampling adds a few percent and is *time-based*, so it won't change which methods look hot. Instrumentation's per-call cost can be many×, and it lands unevenly, creating **false hotspots**. In production you essentially *must* use sampling. - **Fidelity to production:** instrumentation can block **JIT inlining**, so you'd profile code that doesn't exist in the optimized build. Sampling observes the real, optimized program. Reach for: - **async-profiler** — low-overhead, safepoint-free sampler; produces **flame graphs**; can sample CPU, allocations, locks, and wall-clock; sees native/kernel frames. - **JDK Flight Recorder (JFR)** + **JDK Mission Control** — built into the JDK, designed for always-on sub-percent overhead, rich event stream (GC, allocation, locks, I/O). Safe to leave running in production. ### Step 3 — Profile the *right* resource A service can be slow for reasons CPU sampling won't show: - **CPU-bound** → on-CPU sampling; the flame graph's widest frames are your targets. - **Blocked on I/O / locks** (latency high but CPU low) → **wall-clock / off-CPU** profiling or **lock** profiling; on-CPU sampling would show an almost-idle app and tell you nothing. - **GC pressure** → **allocation** profiling (who allocates the most) plus GC logs/JFR GC events. Choosing the wrong dimension is the most common profiling mistake. ### Step 4 — Use instrumentation only for what sampling can't answer Sampling tells you *where time goes* but not *exact counts*. Switch to **instrumentation** when: - you need exact **invocation counts** (e.g. confirm an N+1 query calls a method 10,000× per request), - you must **trace a specific, narrow path** end-to-end, - the code of interest is too short-lived to accumulate samples. Do this in a **controlled (non-prod) environment**, scope it to the suspect packages (don't instrument everything), and treat absolute timings as unreliable — read the *counts*, not the times. ### Step 5 — Interpret carefully - Read flame graphs by **width** (share of samples), not depth. - Watch for **safepoint bias** if using a legacy JVMTI sampler — prefer async-profiler/JFR. - Validate a hypothesis with a targeted change + re-profile (the scientific loop), rather than optimizing the first wide frame blindly. ## Term definitions - **Flame graph:** a visualization where each box is a stack frame, width = fraction of samples; stacked to show call hierarchy. - **On-CPU vs off-CPU / wall-clock:** on-CPU counts only time threads are running on a core; wall-clock/off-CPU also counts time blocked (I/O, locks, sleep). - **JIT inlining:** the compiler substituting a small method's body into its caller to remove call overhead. ## Deriving your own answer Default to a low-overhead **sampler** (async-profiler/JFR) under **realistic load**; **match the profiling dimension** to the suspected bottleneck (CPU/wall-clock/lock/allocation); only escalate to **instrumentation** in a controlled run when you specifically need exact counts or a precise trace. That sequence is the senior-grade answer.
- Your CPU profile shows the app is mostly idle yet latency is high. What do you do?Switch to wall-clock/off-CPU profiling (or lock profiling). The threads are blocked on I/O or contention, which on-CPU sampling can't see; you need a dimension that counts blocked time.
- When is instrumentation clearly the better choice over sampling?When you need exact invocation counts or to trace one specific narrow path — e.g. confirming a method is called thousands of times per request — accepting heavier overhead in a controlled environment.
saying these in an interview costs you the question
- Defaulting to instrumentation in production
- Only doing CPU sampling when the bottleneck is I/O, locks, or GC
- Profiling an idle/toy workload and trusting the result
- Reading flame graphs by depth instead of width
- Optimizing the first wide frame without forming and testing a hypothesis