Describe JFR's event model: what kinds of events does it capture, and how does the sampling-and-buffering design keep the cost low enough for production?
answer
- Event = timestamp + duration + thread + stack + typed fields
- Kinds: allocation, GC, lock/monitor, exception, I/O, method-sample
- Three shapes: duration/instant, sampled, periodic
- Cheap via: native + per-thread buffers + sampling + thresholds
- Correlate event types to find root cause
basics
~20 sJFR records typed 'events' for things like object allocation, garbage collection, lock contention, thrown exceptions, file/socket I/O, and periodic method (CPU) samples. Events go into per-thread buffers and CPU profiling uses sampling, so the overhead stays around 1%.
solid answer
~50 sJFR's data model is a stream of **events**: each event has a start time, optional duration, the originating thread, and typed fields. There are broadly three categories. **Instant/duration events** record discrete happenings — object allocation (inside/outside TLAB), GC start/end and pause times, monitor/lock contention (a thread blocked entering a monitor), thrown exceptions and errors, and I/O (file and socket read/write durations). **Sampled events** answer 'where is the CPU' and 'who is allocating' by periodically sampling thread stacks rather than instrumenting every call — this is the execution-sample (method profiling) event. **Periodic events** snapshot JVM and OS state at intervals: heap usage, GC configuration, thread counts, CPU load, environment. The low cost comes from: native instrumentation, **per-thread buffers** (threads write locally and rarely contend), **sampling** instead of exhaustive instrumentation for the hot paths, and configurable **thresholds** so cheap-but-noisy events (e.g. only record locks held > 10 ms) are filtered. Events flush from thread buffers to a global buffer to the disk repository and ultimately the .jfr file.
go deeper
Can name the main event kinds (allocation, GC, lock, exception, I/O, CPU/method samples) and say JFR records them as timestamped events.
Explains the event structure (timestamp/duration/thread/fields), distinguishes sampled CPU profiling from exhaustive tracing, and knows events buffer per-thread.
Articulates the three event shapes (duration/sampled/periodic), thresholds and per-event enablement, the TLAB allocation model, and how per-thread buffering plus sampling keep overhead near 1%.
Designs event configurations against an overhead budget, defines custom application JFR events for domain telemetry, and correlates event types across a recording to drive root-cause analysis at scale.
## What an 'event' is In JFR, **everything recorded is an event**. An *event* is a typed record with, at minimum: a **timestamp**, optionally a **duration** (for things that take time), the **thread** that produced it, a **stack trace** (often), and a set of **typed fields** specific to that event kind. Events are described by *event types* (each has a name like `jdk.ObjectAllocationInNewTLAB`), so a recording is essentially a typed log of what the JVM and application did. ## The kinds of events The coverageArea lists the headline categories — here is what each tells you: - **Allocation events** — record object allocations. The JVM gives each thread a *TLAB* (Thread-Local Allocation Buffer, a private slab of the heap it can allocate from without locking); JFR emits an event when a thread allocates a new TLAB or allocates an object *outside* a TLAB. These reveal **allocation pressure** — the root cause of frequent GC. - **GC (garbage collection) events** — when collections start/stop, how long they **paused** application threads, which generation, and heap occupancy before/after. This is how you diagnose GC-induced latency. - **Lock / monitor events** — when a thread **contends** for a monitor (waits to enter a `synchronized` block) or parks on a `java.util.concurrent` lock. These expose **contention** that serializes your application. - **Exception / error events** — exceptions and errors thrown. A flood of thrown-and-caught exceptions is a common hidden cost. - **I/O events** — file and socket reads/writes and how long they took, exposing slow disk or network calls. - **Method sampling (execution-sample) events** — periodically, JFR captures the stack of running threads. Aggregating these stacks reconstructs a **CPU profile** (a flame graph of where time is spent) without instrumenting every method. Beyond these, JFR emits hundreds of event types (class loading, compilation, thread state, safepoints, JVM flags, container limits, etc.). ## Three structural categories 1. **Duration/instant events** — fired when a specific thing happens (a GC, a lock wait, an allocation). Often gated by a **threshold** so only the significant ones are kept (e.g. record only locks held longer than 10 ms, or I/O longer than 20 ms). 2. **Sampled events** — JFR *samples* rather than records every occurrence. Method profiling samples stacks at a configured rate; allocation profiling can also be sampled. Sampling is statistical: you get an accurate *distribution* at a tiny fraction of the cost of full instrumentation. 3. **Periodic events** — emitted on a timer to snapshot state: heap summary, GC config, thread/CPU/OS metrics, JVM flags. ## Why this keeps overhead near 1% - **Native instrumentation inside the JVM** — events are emitted by JVM internals, not by bytecode rewriting of your classes. - **Per-thread buffers** — each thread writes events into its own thread-local buffer, so threads almost never contend with each other to record. Full buffers are flushed to a global buffer and then to an on-disk *repository*. - **Sampling, not exhaustive tracing** — the costly question ('where is the CPU/where do allocations come from') is answered by sampling, so cost is bounded by the sample rate, not the application's call volume. - **Thresholds and per-event enablement** — a `.jfc` settings template enables/disables each event type and sets thresholds, letting you trade coverage for cost. The `default` template is tuned for < 1%; `profile` raises sampling density. ## Putting it together for diagnosis Because events are typed and timestamped, JMC (or a custom consumer) can correlate them: a latency spike (duration on a request event) lining up with a GC pause and a burst of allocation events points straight at allocation pressure; a throughput cliff lining up with monitor-contention events points at a lock. That correlation across event types is the analytical power of the model.
- Why does JFR sample stacks for method profiling instead of instrumenting every method entry and exit?Full instrumentation cost scales with call volume and can add tens of percent of overhead, distorting timing. Sampling caps the cost at the sample rate and still yields an accurate statistical distribution of where time is spent — keeping JFR under its ~1% budget.
- What is a TLAB and why does it appear in JFR allocation events?A TLAB (Thread-Local Allocation Buffer) is a private slab of the heap each thread allocates from without locking. JFR emits events when a thread allocates a new TLAB or allocates outside one, which is how it tracks allocation pressure cheaply.
saying these in an interview costs you the question
- Claiming JFR instruments every method call — CPU profiling is sampling-based, which is precisely why it is cheap.
- Saying allocation events record every single object — allocation is sampled / TLAB-based, not one event per object.
- Treating JFR as 'just a CPU profiler' — it captures GC, locks, exceptions, I/O and many JVM events too.
- Assuming all event types are always on — they are configurable, often gated by duration thresholds.