Walk through the full Observation lifecycle and explain how handlers, scopes, predicates, and conventions interact when you build custom instrumentation.
answer
- create(+predicate check) -> start(onStart) -> openScope(onScopeOpened) -> error(onError) -> stop(onStop)
- scope = thread-bound context for children + MDC
- observe()/observeChecked() wrap it safely; else finally-close
- FirstMatching (tracing) vs AllMatching composite handlers
- predicate drops all; sampling samples spans only; scope leaks in async
basics
~20 sCreate an Observation with a convention and context, start it (onStart fires), open a scope so nested work sees trace context (onScopeOpened), run the work, record errors (onError), then stop it (onStop records timing/finishes the span). Predicates decide up front whether it runs at all; conventions decide naming/tags; handlers do the actual recording.
solid answer
~50 sThe Observation API unifies metrics and tracing behind one recording. Lifecycle: Observation.createNotStarted(convention, contextSupplier, registry) builds an observation whose Context carries all data; predicates are consulted so unwanted observations become NoopObservations. On start(), every handler whose supportsContext is true gets onStart (note start time, open span). openScope() binds the context to the current thread via a ThreadLocal-backed scope so child observations and logging (MDC trace/span ids) see the parent — onScopeOpened; closing gives onScopeClosed. Failures call error(throwable), setting the error on the context and firing onError. stop() fires onStop where handlers record duration, apply ObservationFilters to finalize KeyValues, and finish the span. observe(...) wraps start/openScope/run/error/close/stop correctly, which matters because forgetting to close a scope leaks context across threads. Conventions govern names/tags; handlers (often composite: FirstMatching for tracing) govern sinks. Design decisions: correct cardinality, cheap predicates, scope discipline for async, and not annotating hot paths.
code
java · 33 linesimport io.micrometer.observation.Observation;
import io.micrometer.observation.ObservationRegistry;
class OrderService {
private final ObservationRegistry registry;
OrderService(ObservationRegistry registry) { this.registry = registry; }
Order place(OrderRequest req) {
Observation obs = Observation
.createNotStarted("order.place", registry)
.lowCardinalityKeyValue("channel", req.channel()) // bounded -> metric tag
.highCardinalityKeyValue("order.id", req.id()) // unbounded -> span only
.contextualName("place order");
// observe() runs the full lifecycle: start -> openScope -> run -> error -> closeScope -> stop
return obs.observe(() -> doPlace(req));
}
// Manual form makes the scope discipline explicit:
Order placeManual(OrderRequest req) {
Observation obs = Observation.start("order.place", registry);
try (Observation.Scope scope = obs.openScope()) { // binds context to this thread (MDC, children)
return doPlace(req);
} catch (RuntimeException ex) {
obs.error(ex); // fires onError, marks span errored
throw ex;
} finally {
obs.stop(); // fires onStop: record timer, finish span, apply filters
}
}
private Order doPlace(OrderRequest req) { return /* ... */ null; }
}go deeper
Know the rough order: start, run, stop, and that handlers do the recording.
Explain scopes for context propagation and that observe() wraps the lifecycle safely.
Detail predicate/convention/filter/handler roles, composite handlers, and error handling in the lifecycle.
Own cardinality budgets, noise/sampling policy, async context propagation, naming standards, and testing strategy across services.
## One recording, two signals Micrometer's **Observation** exists so a single instrumentation point can emit **both** a metric (timer) and a **trace span** without double-instrumenting. The moving parts: - **`Observation.Context`** — the mutable per-observation data bag (name, tags, error, arbitrary entries). - **`ObservationConvention`** — names the observation and derives low/high-cardinality `KeyValues`. - **`ObservationPredicate`** — decides *whether* to record (returns a `NoopObservation` if not). - **`ObservationHandler`** — the sinks: `onStart/onStop/onError/onEvent/onScope*` produce metrics/spans/etc. - **`ObservationFilter`** — last-chance mutation of the context's `KeyValues` before stop. ## The lifecycle, step by step ```java Observation obs = Observation.createNotStarted(convention, () -> new MyContext(...), registry); ``` 1. **Creation + predicate check.** The registry consults all `ObservationPredicate`s. If any returns `false`, you get a **`NoopObservation`** and everything below is a no-op — nothing recorded. 2. **start()** -> for each handler with `supportsContext == true`, **`onStart(context)`** runs: metric handler captures a `Timer.Sample`; tracing handler starts a span. Convention's low-cardinality KeyValues are attached here. 3. **openScope()** -> **`onScopeOpened`**. A **Scope** binds the observation to the **current thread** (ThreadLocal). This is what lets nested/child observations attach as children and lets logging frameworks put trace/span ids into **MDC**. The scope is an `AutoCloseable`; closing it -> **`onScopeClosed`** (and there's `onScopeReset` for propagation resets). 4. **error(throwable)** (on failure) -> stores the error on the context and fires **`onError`**; the span is marked errored, an error tag is added. 5. **stop()** -> **`onStop`**. Handlers finalize: metric handler records the timer duration with the final tags; tracing handler ends the span. **`ObservationFilter`s** run to finalize the `KeyValues` (add common tags, redact) just before recording. The safe way to do all of this is: ```java obs.observe(() -> doWork()); // start, openScope, run, error, closeScope, stop // or, checked: obs.observeChecked(() -> doWork()); ``` Doing it manually requires try/finally around scope close and stop. ## Handler composition Multiple handlers can support the same context. Micrometer provides composites: - **`FirstMatchingCompositeObservationHandler`** — delegates to the *first* child that supports the context. Boot uses this for **tracing** so only one tracer handler runs. - **`AllMatchingCompositeObservationHandler`** — delegates to *all* supporting children. Handlers run per registered order; state that must survive start->stop lives on the **Context**, never on handler fields (concurrency). ## Scope discipline and async The subtle production issue is **scope leakage**. A scope is thread-bound; if you open a scope and hand the work to another thread (executor, reactive pipeline) without proper context propagation, the child sees the wrong parent or the ThreadLocal is never cleared. For imperative async, use Micrometer's context-propagation library (`ContextSnapshot`) or `ObservationThreadLocalAccessor`; for reactive, propagate via the Reactor context. Always close scopes in `finally` (or use `observe(...)`). ## Where @Observed fits `@Observed` + `ObservedAspect` is just a convenient front door that runs exactly this lifecycle around a method using a default convention (overridable). It does not change any of the above semantics. ## Principal-level design decisions - **Cardinality budget:** enforce that ids/URLs go to high-cardinality (spans), bounded values to low-cardinality (metrics). A single bad tag can multiply time series and take down the metrics backend. - **Noise control:** cheap `ObservationPredicate`s to drop health/actuator/static; combine with tracer **sampling** (predicate drops the whole observation; sampling keeps metrics but samples spans — choose deliberately). - **Naming standards:** central `GlobalObservationConvention`s or shared library conventions so names/tags are consistent across services. - **Performance:** handlers and predicates run on the request thread — keep them non-blocking; never do I/O in a callback. - **Where to instrument:** rely on framework instrumentation (web, JDBC, messaging) first; add `@Observed`/manual observations for business-meaningful boundaries, not hot inner loops. - **Testing:** `TestObservationRegistry` lets you assert an observation was started/stopped with expected tags. ## Gotchas recap - Forgetting to close a scope -> ThreadLocal/MDC leak across pooled threads. - Predicate is all-or-nothing (both metric+span); use handler `supportsContext` for per-sink control. - Convention names/tags; Filter mutates; Predicate drops; Handler records — keep the four roles distinct. - `@Observed` self-invocation isn't intercepted (proxy).
- What breaks if you open a scope but hand the work to another thread without context propagation?The child thread doesn't see the parent observation, so child spans attach to the wrong parent or none, and MDC trace ids are missing/wrong. The original thread's ThreadLocal may also leak. You need context-propagation (ContextSnapshot/ObservationThreadLocalAccessor) or Reactor context.
- You need actuator metrics but no actuator spans. Predicate or sampling?Not a predicate — it drops the whole observation (metric and span). Keep the observation and control spans at the tracer via sampling (or a tracing-specific handler's supportsContext), so metrics are recorded while spans are sampled/skipped.
- How do you unit-test that an observation was recorded with the right tags?Use TestObservationRegistry; run the code against it and assert the observation was started/stopped and carried the expected low/high-cardinality KeyValues.
saying these in an interview costs you the question
- Not opening/closing a scope and expecting child spans/MDC to work
- Using an ObservationPredicate when you actually want to keep metrics but sample spans
- Storing per-observation state on handler fields instead of the Context
- Blocking I/O inside a handler or predicate on the request thread
- Annotating hot inner-loop methods with @Observed and adding overhead everywhere