You need a `WebFilter` that measures total request latency and propagates a correlation ID to downstream reactive/logging code. What are the threading and context-propagation pitfalls, and how do you do it correctly?
answer
- threads hop -> ThreadLocal/MDC unreliable + leaks
- latency: start nanoTime + doFinally on chain.filter
- correlation id: Reactor Context via contextWrite
- read with deferContextual
- MDC bridge = Micrometer context-propagation / enableAutomaticContextPropagation
basics
~20 sNever use a ThreadLocal/MDC set imperatively — reactive work hops threads, so the value won't be there downstream. Measure latency by timing around the returned Mono (record start, then .doFinally on chain.filter). Propagate the correlation ID via the Reactor Context using .contextWrite(...), not a thread-local.
solid answer
~40 sIn WebFlux a single request is processed across many event-loop threads as operators execute, so anything stored in a `ThreadLocal` (including SLF4J MDC) at the top of the filter is *not* reliably visible where the actual work runs — thread-locals don't follow the reactive chain. For **latency**, capture `start = System.nanoTime()` and attach `chain.filter(exchange).doFinally(sig -> record(System.nanoTime()-start))`; `doFinally` fires on complete/error/cancel regardless of thread. For **correlation ID propagation**, put it in the **Reactor `Context`** via `chain.filter(exchange).contextWrite(ctx -> ctx.put(KEY, id))`, and read it downstream with `Mono.deferContextual`. To bridge Reactor Context to MDC for logging, use the **Micrometer Context Propagation** library (`ContextRegistry` / `ThreadLocalAccessor`, auto-wired by `Hooks.enableAutomaticContextPropagation()` in recent Reactor/Boot), which restores thread-locals around operator execution. Also register response headers via `beforeCommit` and never block.
code
java · 45 linesimport io.micrometer.core.instrument.MeterRegistry;
import org.springframework.core.Ordered;
import org.springframework.core.annotation.Order;
import org.springframework.stereotype.Component;
import org.springframework.web.server.*;
import reactor.core.publisher.Mono;
import java.util.UUID;
import java.util.concurrent.TimeUnit;
@Component
@Order(Ordered.HIGHEST_PRECEDENCE) // outermost: measures the whole chain
public class TracingLatencyFilter implements WebFilter {
static final String CORRELATION_KEY = "correlationId";
private final MeterRegistry meterRegistry;
TracingLatencyFilter(MeterRegistry meterRegistry) {
this.meterRegistry = meterRegistry;
}
@Override
public Mono<Void> filter(ServerWebExchange exchange, WebFilterChain chain) {
String id = header(exchange, "X-Correlation-Id", UUID.randomUUID().toString());
long start = System.nanoTime();
exchange.getResponse().beforeCommit(() -> {
exchange.getResponse().getHeaders().set("X-Correlation-Id", id);
return Mono.empty();
});
return chain.filter(exchange)
// fires on complete / error / cancel, on whatever thread
.doFinally(signal -> meterRegistry
.timer("http.server.latency", "outcome", signal.name())
.record(System.nanoTime() - start, TimeUnit.NANOSECONDS))
// propagate the id through the reactive pipeline (NOT a ThreadLocal)
.contextWrite(ctx -> ctx.put(CORRELATION_KEY, id));
}
private static String header(ServerWebExchange ex, String name, String fallback) {
String v = ex.getRequest().getHeaders().getFirst(name);
return (v != null && !v.isBlank()) ? v : fallback;
}
}go deeper
Know that thread-locals are unreliable in reactive code and latency is measured around the returned Mono.
Use doFinally for latency and understand Reactor Context exists for propagation instead of ThreadLocal.
Explain contextWrite/deferContextual, beforeCommit timing, and outermost ordering for end-to-end latency.
Discuss the Micrometer context-propagation bridge (ThreadLocalAccessor, enableAutomaticContextPropagation), cross-request thread-local leakage, cancellation semantics, and singleton-filter state hazards.
## The core problem: threads hop WebFlux runs on a small pool of **event-loop threads** (Reactor Netty). A request's operators may execute on different threads over time (and `publishOn`/`subscribeOn` deliberately switch threads). Imperative side-effects at the *top* of `filter(...)` run on whatever thread invoked the filter, but the controller/repository code runs later, possibly elsewhere. Therefore: - `ThreadLocal.set(...)` / `MDC.put(...)` at the start of the filter is **not guaranteed** to be visible in downstream reactive operators, and if you don't clear it you can **leak** the value onto a shared event-loop thread for the *next* request. This is a classic reactive bug. ## Measuring latency correctly Timing is about the *lifetime of the returned Mono*, not wall-clock in the filter body: ```java long start = System.nanoTime(); return chain.filter(exchange) .doFinally(signal -> { long micros = (System.nanoTime() - start) / 1_000; meterRegistry.timer("http.server.latency").record(micros, TimeUnit.MICROSECONDS); }); ``` - Put this filter at a **low `@Order`** (outermost) so it measures time spent in all other filters + the handler. - `doFinally` receives a `SignalType` (`ON_COMPLETE`, `ON_ERROR`, `CANCEL`) — good for tagging outcomes and catching client disconnects (CANCEL). - Avoid measuring inside `.then(...)` only, which wouldn't fire on error. ## Propagating a correlation ID the reactive way Use the **Reactor `Context`** — an immutable key/value map that flows *up* the subscription, opposite to data flow — written with `contextWrite`: ```java return chain.filter(exchange) .contextWrite(ctx -> ctx.put(CORRELATION_KEY, id)); ``` Downstream code reads it: ```java Mono.deferContextual(ctx -> { String id = ctx.get(CORRELATION_KEY); // use id return Mono.just(...); }); ``` Because `Context` is carried by the reactive pipeline itself, it survives thread hops — unlike a `ThreadLocal`. ## Bridging Context ↔ MDC for logging Loggers use MDC, which is thread-local. To make the correlation ID appear in log lines from reactive code, use **Micrometer Context Propagation** (`io.micrometer:context-propagation`): - Register a `ThreadLocalAccessor` (or use Micrometer Tracing / Boot's auto-config) so the framework *temporarily restores* the thread-local around each operator's execution and clears it afterward. - In recent Reactor + Spring Boot you enable `reactor.core.publisher.Hooks.enableAutomaticContextPropagation()` (Boot does this when context-propagation is on the classpath), so values in the Reactor `Context` are mirrored into thread-locals (MDC) automatically during operator execution and cleaned up after — no leaks. - Do **not** hand-roll `MDC.put/remove` around `chain.filter`; timing on async boundaries is wrong and leaks across requests. ## Response headers with correct timing To emit the correlation ID / timing as a response header without racing the handler's commit: ```java exchange.getResponse().beforeCommit(() -> { exchange.getResponse().getHeaders().set("X-Correlation-Id", id); return Mono.empty(); }); ``` ## Other pitfalls - **No blocking**: metrics export, ID generation, and logging must be non-blocking; a blocking call on the event loop stalls many requests. - **Cancellation**: clients disconnecting cancel the Mono; `doFinally` sees `CANCEL` — count it, don't treat it as success. - **Ordering vs Security**: Spring Security is a `WebFilter`; if you need the authenticated principal in your correlation/log, order after security (higher order) or read the principal via `exchange.getPrincipal()` reactively. - **Don't store per-request state in filter fields** — filters are singletons shared across all requests; keep per-request data in local variables, the exchange attributes (`exchange.getAttributes()`), or the Reactor Context. ## When to use what - Reactor `Context` + `contextWrite`: propagating request-scoped values (correlation IDs, tenant, auth) to reactive code. - Micrometer context-propagation + `enableAutomaticContextPropagation`: getting those values into MDC/thread-locals for logging and tracing libraries. - `exchange.getAttributes()`: sharing values with the *same-request* downstream filters/handler when you don't need thread-local restoration. - `doFinally`/`beforeCommit`: latency + response-side effects with correct timing.
- Why is `MDC.put("cid", id)` at the top of a WebFilter unreliable, and can it cause bugs beyond missing logs?Reactive operators run on event-loop threads that differ from the one that entered the filter, so MDC (a ThreadLocal) may be empty where the work runs. Worse, since event-loop threads are shared and pooled, a value you set and never clear can leak into the next request handled by that thread — cross-request contamination. Use the Reactor Context plus Micrometer context-propagation instead.
- Why place the latency filter at HIGHEST_PRECEDENCE?Lowest order = outermost = first in, last out. It wraps all other filters and the handler, so the interval it measures includes their time too, giving true end-to-end latency.
- How does the correlation ID actually reach MDC in log lines from reactive code?Via Micrometer's context-propagation library: a registered ThreadLocalAccessor mirrors the Reactor Context value into the MDC ThreadLocal around each operator's execution and clears it afterward. With `Hooks.enableAutomaticContextPropagation()` (enabled by Boot when context-propagation is present) this happens automatically.
saying these in an interview costs you the question
- Using MDC.put/ThreadLocal set at the top of the filter and assuming it's visible downstream
- Not clearing thread-locals, causing cross-request leakage on shared event-loop threads
- Storing per-request state in the singleton filter's instance fields
- Measuring latency only in .then() so errors/cancellations are missed
- Blocking (synchronous metric export or logging I/O) on the event loop