How does trace context survive across thread boundaries (async, thread pools, reactive) in Micrometer Tracing, and how does sampling affect what you can observe?
answer
- current span = thread-local -> doesn't cross threads
- context-propagation: ContextSnapshot capture/restore
- wrap executors / TaskDecorator; @Async not automatic
- Reactor: Hooks.enableAutomaticContextPropagation()
- sampling head-based at root, propagates; default 0.1
basics
~20 sTrace context is stored per-thread, so it does not automatically follow work handed to another thread. You must propagate it — wrap executors, use context-propagation instrumentation, or capture/restore the context. Sampling decides upfront whether a whole trace is recorded, so unsampled traces produce no spans.
solid answer
~40 sThe current span lives in a **thread-local** (via the tracer's scope). When work moves to another thread — an `@Async` method, a raw `ExecutorService`, a `CompletableFuture`, or a reactive scheduler hop — the thread-local does not travel, so the child work either starts a new root or loses correlation. Micrometer solves this with the **context-propagation** library (`ContextRegistry`, `ContextSnapshot`): you capture a snapshot on the caller thread and restore it on the worker. In practice you wrap executors (`ContextExecutorService` / a `TaskDecorator`), and Boot's Observation instrumentation already handles `WebClient`/Reactor when `Hooks.enableAutomaticContextPropagation()` is on. **Sampling** is a per-trace head decision: `management.tracing.sampling.probability` (default 0.1) decides at the root whether the trace is recorded; that decision propagates, so an unsampled trace emits no spans anywhere — you still get trace IDs in logs but nothing in Zipkin.
code
java · 33 lines// Make a thread pool trace-aware via a TaskDecorator that
// captures the caller's context and restores it on the worker thread.
@Configuration
class AsyncTracingConfig {
@Bean
ThreadPoolTaskExecutor tracedExecutor() {
ThreadPoolTaskExecutor executor = new ThreadPoolTaskExecutor();
executor.setTaskDecorator(runnable -> {
// capture on the submitting thread
ContextSnapshot snapshot = ContextSnapshotFactory.builder().build().captureAll();
return () -> {
// restore on the worker thread; scope closed after run
try (ContextSnapshot.Scope scope = snapshot.setThreadLocals()) {
runnable.run();
}
};
});
executor.initialize();
return executor;
}
}
// Alternatively, wrap any ExecutorService:
// ExecutorService traced = ContextExecutorService.wrap(
// Executors.newFixedThreadPool(4),
// () -> ContextSnapshotFactory.builder().build().captureAll());
// application.yml
// management:
// tracing:
// sampling:
// probability: 0.1 # head-based: 10% of traces recorded & exportedgo deeper
Know context is per-thread and sampling limits what's recorded.
Explain why @Async needs a trace-aware executor and what probability does.
Use context-propagation snapshots / Reactor hooks correctly and debug scope leaks.
Architect sampling strategy (head + collector tail-based), govern cost/overhead, and prevent thread-pool context contamination fleet-wide.
## Why threads matter A tracer tracks the **current span** in a **thread-local** (Brave's `CurrentTraceContext`, OTel's `Context`). MDC (which holds `traceId`/`spanId` for logs) is also thread-local. So anything that moves execution to a *different* thread than the one that opened the span risks losing the context: the worker thread's thread-local is empty, so new spans have no parent and logs lose the IDs. Common boundary crossings: - **`@Async` methods** running on a `TaskExecutor`. - Manual `ExecutorService.submit(...)` / `CompletableFuture.supplyAsync(...)`. - **Reactive** pipelines where operators run on different **Reactor Schedulers** (`subscribeOn`/`publishOn`). - Messaging listeners, scheduled tasks, and callbacks. ## The context-propagation library Micrometer ships `io.micrometer:context-propagation`, a small library that standardizes moving thread-local context across boundaries: - **`ContextRegistry`** knows how to read/write registered thread-locals (tracing registers its own accessors). - **`ContextSnapshot`** captures the current thread-locals into an immutable snapshot on the source thread; you then **restore** it on the target thread inside a try-with-resources scope. - Helpers: `ContextSnapshot.captureAll()`, `snapshot.wrap(Runnable)`, `ContextSnapshotFactory`, and `ContextExecutorService` / `ContextScheduledExecutorService` to auto-wrap tasks submitted to an executor. ### Making `@Async` / executors trace-aware Wrap the executor so every submitted task restores the snapshot. Spring also lets you register a `TaskDecorator` on the `ThreadPoolTaskExecutor` that captures a snapshot at submit time and opens its scope at run time. ### Reactive For Reactor/WebFlux you enable **automatic** propagation once at startup: `Hooks.enableAutomaticContextPropagation()`. Boot 3 turns this on for you when Micrometer Tracing + Reactor context-propagation are present, so `WebClient` and reactive controllers keep the trace across scheduler hops. Reactive context flows via the **Reactor `Context`**, not thread-locals, which is why explicit bridging is needed. ## Sampling — head-based decision **Sampling** answers 'should this trace be recorded and exported?' The default strategy is **head-based**: the decision is made once at the **root span** (trace entry) and encoded in the sampling flag, which then **propagates** to every downstream service (in the `traceparent` flags or `X-B3-Sampled`). Consequences: - `management.tracing.sampling.probability` (default **0.1** = 10%) sets the fraction recorded. `1.0` records all; `0.0` records none. - Because the decision propagates, sampling is **consistent** across the whole trace — you never get half a trace. Either every service records it or none do. - An **unsampled** trace still generates trace/span IDs (so logs stay correlated) but emits **no spans to Zipkin** — a frequent 'tracing looks broken' confusion. - Head-based sampling can't decide based on the *outcome* (you don't yet know it will error/be slow). **Tail-based sampling** (keep the interesting traces) is done outside the app, typically in an **OpenTelemetry Collector**, not by Micrometer itself. ## Cost / design trade-offs (principal lens) - **Overhead**: span creation and context propagation add small CPU/allocation cost; the bigger cost is export volume and backend storage. Sampling is the primary lever. - **Consistency vs completeness**: low probability saves money but you'll miss most traces; combine low head sampling with tail sampling in a collector to keep errors/slow traces. - **Baggage size**: propagated context adds header bytes on every hop. - **Thread-pool leaks**: if you open a scope on a pooled thread and don't close it, the next task on that thread inherits a stale span — subtle cross-request contamination. Always use try-with-resources / `observe()`. - **Virtual threads / structured concurrency**: still thread-local based; context-propagation wrapping remains necessary. ## Gotchas - Setting `probability` low in prod then wondering why Zipkin is sparse — it's working as designed. - Assuming `@Async` propagates context automatically — it does not without a trace-aware executor/decorator. - Manually creating a `Thread` and never restoring a snapshot -> orphaned spans. - Expecting Micrometer to do tail-based sampling — it does head-based; tail-based lives in the collector.
- Why can Micrometer Tracing not implement tail-based sampling (keep only error/slow traces) on its own, and where is it done?Head-based sampling decides at the root span before the outcome is known, and that decision propagates to keep the trace consistent — you can't retroactively 'un-drop' spans other services never recorded. Tail-based sampling needs to buffer all spans of a trace and decide after completion, which requires a central component that sees the whole trace: typically an OpenTelemetry Collector, not the application.
- A pooled worker thread starts logging with a stale trace ID from a previous request. What went wrong?A span scope was opened on that pooled thread and never closed (no try-with-resources / not using observe()). The thread-local retained the old span, so the next task reusing the thread inherited it. The fix is to always scope spans with try-with-resources or the Observation observe() helper so the thread-local is cleared.
saying these in an interview costs you the question
- Assuming @Async or a raw ExecutorService propagates trace context automatically.
- Thinking sampling is decided per-span or per-service rather than once at the root and propagated.
- Believing an unsampled trace produces no trace IDs at all (it still has IDs in logs; it just isn't exported).
- Expecting Micrometer to do tail-based sampling of errors/slow traces.
- Leaving a span scope open on a pooled thread, contaminating later requests.