skip to content

Walk through the full lifecycle of a manually managed Observation (createNotStarted → start → openScope → error → stop). What breaks context propagation and how do you avoid it?

level: principalimportance: should knowfreq 30%

answer

  1. start→openScope→(error)→close→stop, stop in finally
  2. openScope = thread-local binding = correlation + MDC
  3. no scope → orphan child spans, uncorrelated logs
  4. thread hop loses context → use ContextSnapshot/ContextExecutorService
  5. observe() does the whole dance safely; TestObservationRegistry to assert

basics

~20 s

Create with createNotStarted, start() it, openScope() to bind it to the thread (so child spans/logs correlate), run work, call error(e) on failure, close the scope, then stop(). Forgetting the scope breaks correlation; forgetting stop() leaks. observe() does all this safely; async work needs context propagation.

solid answer

~40 s

Manual lifecycle: `createNotStarted(name, registry)` builds it; `start()` fires handlers' onStart (Timer.Sample begins, span opens); `openScope()` binds the Observation to the current thread via thread-locals so nested Observations and logs (MDC traceId/spanId) correlate as children; run the work; on failure call `observation.error(throwable)` so the span is marked ERROR and metrics get an exception tag; then `scope.close()` and finally `stop()` in a finally block. Two classic breakages: (1) **not opening a scope** — child spans attach to the wrong parent and logs lose correlation; (2) **crossing threads** (async/reactive/executor) — thread-locals don't follow, so context is lost. Fixes: prefer `observe(...)`; for async use Micrometer's **ContextPropagation** (`ContextSnapshot`, `ContextExecutorService`, or reactor's `ObservationThreadLocalAccessor`) to capture and restore context on the worker thread.

code

java · 36 lines
java
import io.micrometer.context.ContextSnapshot;
import io.micrometer.context.ContextSnapshotFactory;
import io.micrometer.observation.Observation;
import io.micrometer.observation.ObservationRegistry;
import java.util.concurrent.ExecutorService;

class AsyncOrderProcessor {
    private final ObservationRegistry registry;
    private final ExecutorService pool;

    AsyncOrderProcessor(ObservationRegistry registry, ExecutorService pool) {
        this.registry = registry;
        this.pool = pool;
    }

    void process(Order order) {
        Observation observation = Observation.start("order.process", registry);
        // Capture current thread-local context (current Observation + MDC)
        ContextSnapshot snapshot = ContextSnapshotFactory.builder().build().captureAll();

        pool.submit(() -> {
            // Restore context on the worker thread so the child span links correctly
            try (ContextSnapshot.Scope ignored = snapshot.setThreadLocals();
                 Observation.Scope scope = observation.openScope()) {
                doWork(order); // nested Observations here become children
            } catch (Exception e) {
                observation.error(e);
                throw e;
            } finally {
                observation.stop(); // always stop, even on failure
            }
        });
    }

    private void doWork(Order order) { /* ... */ }
}

go deeper

for a junior

Should just prefer observe() and know start/stop must be balanced.

for a middle

Explains openScope and that stop must be in finally.

for a senior

Knows scope drives correlation and that thread hops lose context, naming a propagation fix.

for a principal

Designs async/reactive instrumentation with ContextSnapshot/ObservationThreadLocalAccessor and enforces it via TestObservationRegistry and executor decorators org-wide.

## The full manual lifecycle ```java Observation observation = Observation.createNotStarted("order.process", registry); observation.start(); // (1) onStart: Timer.Sample, span opened try (Observation.Scope scope = observation.openScope()) { // (2) bind to thread-locals doWork(); // (3) nested Observations + logs correlate here } catch (Exception e) { observation.error(e); // (4) mark span ERROR, add exception KeyValue throw e; } finally { observation.stop(); // (5) onStop: record Timer, end span } ``` Step by step: 1. **createNotStarted** — builds the Observation + Context, resolves convention; **emits nothing yet**. 2. **start()** — invokes every handler's `onStart(context)`: the meter handler begins a `Timer.Sample`; the tracing handler creates/continues a **span**. The clock is now running. 3. **openScope()** — returns an `Observation.Scope` and **binds the Observation to the current thread** through thread-locals. This is what makes it the **current** Observation: `registry.getCurrentObservation()` returns it, **child Observations become child spans**, and the tracing handler pushes **traceId/spanId into the MDC** so logs on this thread are correlated. **No scope ⇒ no parent/child linkage and no log correlation**, even though the timer and a (parentless) span still record. 4. **error(throwable)** — sets `context.setError(t)`; handlers mark the span status ERROR and add an `exception`/`error` KeyValue to metrics. If you swallow the exception without calling error(), your signals show success. 5. **scope.close()** then **stop()** — close unbinds the thread-local (restores the previous current Observation); stop() runs `onStop`: records the Timer duration and ends the span. **Order matters**: close the scope before/at stop; leaving a scope open leaks thread-local state onto pooled threads. `observe(Runnable)` / `observeChecked(...)` does **all** of 2–5 for you in a correct try/catch/finally — which is why it's preferred and why manual management is reserved for cases like async where you must split start and stop across boundaries. ## What breaks context propagation **A. Missing scope.** Recording a Timer works without a scope, but correlation doesn't: child spans orphan and logs lack IDs. Symptom: flat traces, uncorrelated logs. **B. Crossing threads.** Thread-locals are **per-thread**. The moment work hops to another thread — `@Async`, an `ExecutorService`, `CompletableFuture.supplyAsync`, a reactive scheduler, a messaging listener — the current Observation/MDC **does not follow**. Child spans on the worker attach to nothing (or the wrong root). This is the #1 real-world tracing bug. **Fixes for async:** - **Micrometer Context Propagation** library: `ContextSnapshot.captureAll()` on the caller thread, `snapshot.wrap(runnable)` / `setThreadLocals()` on the worker to restore the Observation + MDC. `ObservationThreadLocalAccessor` teaches the propagation library how to move the current Observation. - **ContextExecutorService / ContextScheduledExecutorService** — wrap your executor so submitted tasks auto-restore context. - **Reactor**: with `Hooks.enableAutomaticContextPropagation()` (Reactor 3.5+) and the propagation library, the current Observation flows through the reactive chain; put the Observation in the Reactor `Context` via `.contextWrite(...)` / the tap-and-observe operators, not thread-locals. - **Spring** wires much of this: `@Async` uses `ContextPropagatingTaskDecorator`/`TaskDecorator`; the framework's own executors are context-aware when propagation is on the classpath. ## Other gotchas - **Double stop / never stop**: stopping twice is a no-op-ish misuse; never stopping **leaks** the Timer.Sample and an open span. Always `finally { stop(); }`. - **error() after stop()** is too late — record the error before stopping. - **Reusing an Observation** instance for multiple runs is invalid; create a new one per operation. - **Scope not closed on a pooled thread** poisons the next task with a stale current Observation. - **NOOP registry** in tests: manual lifecycle calls are safe no-ops, so tests won't NPE but also won't assert real behavior — use a `TestObservationRegistry` to assert. ## When to go manual Use the manual form when start and stop are **inherently separated** — request received on one thread, completed on another (async servers, messaging, reactive) — and pair it with context propagation. For straight-line synchronous work, always use `observe(...)`.

  • A Timer is recorded correctly but your traces show child operations as separate root spans. What did you forget?
    openScope(). The Timer only needs start/stop, but parent/child span linkage and MDC log correlation require binding the Observation to the thread via a Scope, so children see it as the current Observation.
  • How do you keep the Observation context intact when work is handed to a thread pool or @Async method?
    Thread-locals don't cross threads. Capture a ContextSnapshot on the caller and restore it on the worker (or wrap the executor with ContextExecutorService / use Spring's context-propagating TaskDecorator), enabling Micrometer Context Propagation with ObservationThreadLocalAccessor.
  • How would you unit-test that your instrumentation emitted the right Observation with the right KeyValues?
    Use TestObservationRegistry and assert via TestObservationRegistryAssert — that an Observation with the expected name started, stopped, had the expected low/high KeyValues, and recorded (or didn't record) an error.

saying these in an interview costs you the question

  • Believing thread-locals propagate automatically across executors/reactive schedulers.
  • Skipping openScope() and expecting correlated child spans/logs.
  • Not stopping the Observation in a finally block (leaks span + Timer.Sample).
  • Calling error() after stop(), or swallowing exceptions without error().
  • Reusing one Observation instance for multiple operations.

context