Why do traceId/spanId disappear from logs on @Async, thread pools, or reactive chains, and how do you fix it?
answer
- MDC is ThreadLocal -> lost across threads
- @Async / executor / publishOn = empty IDs
- TaskDecorator + ContextSnapshot.setThreadLocals()
- Reactor: Hooks.enableAutomaticContextPropagation()
- close scope or you leak stale IDs into pooled threads
basics
~20 straceId/spanId live in MDC, which is thread-local. When work moves to another thread (executor, @Async, reactor), the new thread has no MDC, so logs show empty IDs. Fix it by propagating context: wrap executors, use a TaskDecorator, or enable Reactor context propagation.
solid answer
~40 sMDC is a thread-local map, and Micrometer Tracing stamps traceId/spanId onto whichever thread the span is *scoped* on. Hand the work to a different thread — a raw `ExecutorService`, `@Async`, `CompletableFuture.supplyAsync`, or a reactive operator switching schedulers — and that thread starts with an empty MDC, so its logs print `-`. Fix by propagating context across the boundary. For imperative code: use Micrometer's context-propagation library (`ContextSnapshot`/`ContextRegistry`), wrap the executor with `ContextExecutorService`, or configure a `TaskDecorator` on the `@Async`/`ThreadPoolTaskExecutor` that captures and restores the snapshot. For Reactor: add the context-propagation dependency and call `Hooks.enableAutomaticContextPropagation()` (Reactor 3.5+), so trace context flows through the reactive chain and into MDC on each signal. Never rely on manually copying MDC unless you also restore trace scope.
code
java · 25 lines@Configuration
class AsyncConfig {
@Bean
ThreadPoolTaskExecutor appExecutor() {
var ex = new ThreadPoolTaskExecutor();
ex.setTaskDecorator(new ContextPropagatingTaskDecorator());
ex.initialize();
return ex;
}
static class ContextPropagatingTaskDecorator implements TaskDecorator {
private final ContextSnapshotFactory factory =
ContextSnapshotFactory.builder().build();
@Override public Runnable decorate(Runnable runnable) {
ContextSnapshot snapshot = factory.captureAll();
return () -> {
try (ContextSnapshot.Scope scope = snapshot.setThreadLocals()) {
runnable.run(); // logs here now show the caller's traceId/spanId
}
};
}
}
}go deeper
May just know 'async loses the trace'.
Knows MDC is thread-local and that a TaskDecorator or wrapped executor is needed.
Can implement propagation for both imperative and reactive paths and explain the MDC-vs-scope distinction.
Anticipates leaks, third-party pools, and standardizes context propagation as a platform concern across the codebase.
## Root cause: thread-locality **MDC** is backed by a `ThreadLocal`. Micrometer Tracing writes `traceId`/`spanId` into MDC only for the thread on which the span is currently *scoped*. This is fine for a normal request handled start-to-finish on one Servlet thread. It breaks the moment execution crosses a thread boundary, because the destination thread has its own (empty) MDC and no active span scope. Boundaries that lose it: - `@Async` methods (run on a `TaskExecutor`) - Manually submitted tasks to an `ExecutorService` / `CompletableFuture.supplyAsync(...)` - Reactive pipelines that `publishOn`/`subscribeOn` a different `Scheduler` - `@Scheduled` jobs (there's simply no incoming request/span — often expected) - Parallel streams Symptom: those log lines print `[app,-,-]`. ## The unifying fix: context propagation Micrometer ships the **context-propagation** library (`io.micrometer:context-propagation`) with a `ContextRegistry`, `ContextSnapshot`, and `ThreadLocalAccessor`s. A `ThreadLocalAccessor` teaches the library how to capture/restore a given thread-local — including the tracing context that ultimately drives MDC. You capture a snapshot on the source thread and restore it on the destination thread. ### Imperative options 1. **Wrap the executor**: `ContextSnapshotFactory.builder().build().captureAll()` then `ContextExecutorService` / `ContextScheduledExecutorService`, or Micrometer's `ContextExecutorService.wrap(delegate, snapshotSupplier)`. 2. **TaskDecorator** on your `ThreadPoolTaskExecutor` (used by `@Async`): capture a snapshot when the task is submitted, open it inside the run: ```java public Runnable decorate(Runnable runnable) { ContextSnapshot snapshot = ContextSnapshotFactory.builder().build().captureAll(); return () -> { try (ContextSnapshot.Scope scope = snapshot.setThreadLocals()) { runnable.run(); } }; } ``` This restores the tracing scope on the worker thread, which re-populates MDC there. ### Reactive option Reactor stores context in the `Context`/`ContextView`, not thread-locals. Bridge the two: - add `io.micrometer:context-propagation`, - call `Hooks.enableAutomaticContextPropagation()` once at startup (Reactor 3.5+, pulled in by recent Spring Boot). After that, trace context is captured into the Reactor context and restored to thread-locals (and thus MDC) around each operator invocation, so logs inside `map`/`flatMap` show the right IDs. Older code used `.contextWrite(...)` and manual MDC bridging; automatic propagation supersedes most of that. ## Gotchas - **Copying MDC alone is not enough** for correctness: if you only `MDC.setContextMap(...)` on the worker but don't restore the *span scope*, then any *new* child spans created there won't be parented correctly. Propagate the trace context, and MDC follows. - **Leaks**: always close the scope (try-with-resources). Reusing pooled threads without clearing restored thread-locals leaks IDs into the next unrelated task — logs then show a *stale, wrong* traceId, which is worse than an empty one. - **`@Async` without a decorator** is the single most common cause of "my async logs lost the traceId." - **Third-party thread pools** (Kafka listeners, gRPC, custom libraries) need the same treatment; instrumentation may or may not do it for you. ## When to use Any time you deliberately move request-scoped work to another thread and still want correlated logs/traces. If a background job legitimately has no originating request, an empty traceId is correct — don't fabricate one; start a fresh span instead if you want its own trace.
- You copied MDC into the worker thread manually and logs show the traceId, but new child spans in the worker aren't nested under the parent. Why?Copying MDC only fixes the printed strings; it doesn't restore the actual trace *scope*. New spans read the current trace context (not MDC) for their parent, so without restoring the tracing context they start a detached trace. Propagate the context (ContextSnapshot) instead of just MDC.
- After enabling propagation, a pooled thread sometimes logs a traceId from a previous request. What happened?The scope wasn't closed, so restored thread-locals leaked into the next task reusing that pooled thread. Always restore inside try-with-resources (ContextSnapshot.Scope) so the thread-locals are cleared when the task ends.
saying these in an interview costs you the question
- Saying MDC automatically follows work onto other threads
- Fixing async correlation by only copying MDC without restoring trace scope
- Forgetting to close the scope, causing stale traceIds leaking across pooled tasks
- Believing @Async preserves tracing context out of the box without a TaskDecorator