skip to content

What are the hidden costs of going all-reactive, and how do you catch the classic mistakes?

level: seniorimportance: should knowfreq 50%

answer

  1. broken stack traces → checkpoint / reactor-tools
  2. ThreadLocal dead → Reactor Context
  3. forget to subscribe = nothing runs
  4. blocking stalls loop → BlockHound
  5. load-test, not just StepVerifier

basics

~20 s

Costs: harder debugging (broken stack traces, no ThreadLocal), a steep functional learning curve, less mature libraries, and one blocking call can stall the event loop. Catch mistakes with BlockHound, Reactor context/checkpoints, and load tests — not just unit tests.

solid answer

~50 s

Going all-reactive taxes the whole plumbing. Debugging is the biggest hit: exceptions produce fragmented stack traces that don't show the logical call chain, so you rely on Reactor's onOperatorDebug/checkpoint() or the reactor-tools agent. ThreadLocal-based tooling breaks — MDC logging, Spring Security context, transaction context — because work hops threads; you must propagate state through the Reactor Context instead. The programming model itself is a learning curve (operators, cold vs hot publishers, subscribe timing), and mistakes like forgetting to subscribe mean nothing runs. Library maturity is thinner — you need non-blocking drivers for everything. The signature failure is a stray blocking call (JDBC, RestTemplate, Thread.sleep, logging to a blocking sink) on an event-loop thread, which stalls every multiplexed request; you guard against it with BlockHound in tests. And because the benefit only appears under concurrency, correctness must be validated with realistic load tests, not just happy-path unit tests.

code

java · 20 lines
java
// Guard against blocking the event loop — install BlockHound in tests
@BeforeAll
static void setUp() {
    BlockHound.install(); // throws if a blocking call runs on a non-blocking thread
}

// Add production-safe breadcrumbs instead of relying on stack traces
Mono<Order> order(String id) {
    return repo.findById(id)
        .checkpoint("after findById " + id)   // labeled traceback point
        .flatMap(this::enrich)
        .checkpoint("after enrich");
}

// Propagate correlation id via Reactor Context (ThreadLocal/MDC won't survive thread hops)
Mono<String> handle() {
    return Mono.deferContextual(ctx ->
        Mono.just("traceId=" + ctx.get("traceId"))
    ).contextWrite(Context.of("traceId", "abc-123"));
}

go deeper

for a junior

Know the headline costs: harder to debug, ThreadLocal breaks, and a blocking call can stall everything.

for a middle

Name the tools: BlockHound, checkpoint(), Reactor Context, StepVerifier; explain lazy subscription and .block() throwing on the loop.

for a senior

Explain Context propagation for Security/MDC, production debugging with reactor-tools, and why load tests (not unit tests) validate the choice.

for a principal

Frame these as total-cost-of-ownership: developer velocity, on-call debuggability, and observability tooling maturity versus the efficiency ceiling — the basis for a build/adopt decision across teams.

## Why 'all-reactive' is a whole-stack commitment Reactive isn't a library you sprinkle on a controller — it's a property the *entire request pipeline* must preserve. Break it anywhere (data access, HTTP client, logging, a third-party SDK) and you either lose the benefit or actively harm throughput. The costs below are what teams underestimate. ## 1. Debuggability In MVC, an exception's stack trace mirrors your call graph — you see controller → service → repo. In Reactor, operators run on assorted threads asynchronously, so the stack trace shows Reactor internals, **not** the logical flow that assembled the pipeline. Mitigations: - **`Hooks.onOperatorDebug()`** — global, captures assembly-time stack traces. Heavy; dev-only. - **`.checkpoint("label")`** — lightweight, per-pipeline breadcrumbs in production. - **reactor-tools `ReactorDebugAgent`** — instruments at load time, low overhead, recommended for prod. - Breakpoints are awkward because execution is deferred and thread-hopping. ## 2. ThreadLocal-based infrastructure breaks A huge amount of Java infra hangs off `ThreadLocal`: SLF4J **MDC** (correlation IDs in logs), Spring Security's `SecurityContextHolder`, transaction synchronization, Micrometer tracing. Because reactive work migrates across threads, `ThreadLocal` is unreliable. Reactive Spring instead threads state through the **Reactor `Context`** (an immutable key-value map attached to the subscription): - Security uses `ReactiveSecurityContextHolder` (Context-based), not `SecurityContextHolder`. - Logging correlation needs Context ↔ MDC bridging (`context-propagation` library / Micrometer `ContextRegistry`). Getting this wrong yields empty log context and lost trace IDs. ## 3. The programming-model learning curve - **Nothing happens until you subscribe** — a `Mono`/`Flux` returned but never subscribed (and not returned to the framework) simply never executes. Forgetting to return the publisher from a method is a common silent bug. - **Cold vs hot** publishers, operator fusion, `flatMap` vs `concatMap` (ordering/concurrency), when to `subscribeOn`/`publishOn`. Misuse causes subtle concurrency or ordering bugs. - Blocking bridges must use `Schedulers.boundedElastic()`; calling `.block()` on an event-loop thread throws (`block()/blockFirst()/blockLast() are blocking, which is not supported in thread reactor-http-nio-*`). ## 4. The signature failure: blocking the event loop One synchronous call — JDBC, `RestTemplate`, `Thread.sleep`, `File` I/O, a blocking log appender, even some JSON or crypto libs — on a Netty event-loop thread **parks that thread**, stalling *every* connection it multiplexes. Under load this looks like sporadic, correlated latency spikes that are miserable to diagnose. - **`BlockHound`** is a Java agent that instruments known blocking JDK/native calls and throws when one runs on a thread marked non-blocking. Wire it into the test suite (`BlockHound.install()`), so a blocking call fails CI rather than production. ## 5. Library & ops maturity Every integration needs a non-blocking counterpart: R2DBC instead of JDBC, reactive Redis/Mongo/Kafka drivers, `WebClient` instead of `RestTemplate`. Some SDKs have none, forcing offload. APM/observability tooling historically had weaker reactive support (though Micrometer context-propagation has largely closed this). ## 6. Testing shifts Unit-test correctness with **`StepVerifier`** (asserts emitted signals) and **`VirtualTimeScheduler`** for time-based operators. But the *reason* you chose reactive — behavior under concurrency — is only exercised by **load tests**; passing unit tests tell you little about whether the event loop stays unblocked at scale. ## When the cost is worth it Accept these costs when the concurrency/streaming payoff is real (see the 'when it pays off' question) and the whole ecosystem is non-blocking. Otherwise MVC — increasingly with virtual threads — gives you the imperative model, intact stack traces, working ThreadLocals, and mature drivers, for most of the scalability.

  • Why does MDC-based logging (correlation IDs) often come out empty in WebFlux, and how do you fix it?
    MDC stores context in a ThreadLocal, but reactive work hops threads, so the value isn't present on the thread that logs. Fix by carrying the value in the Reactor Context and bridging it to MDC — via the Micrometer context-propagation library / ContextRegistry (or manual Context ↔ MDC copy in a hook).
  • A method returns a Mono but the side effect never happens. Most likely cause?
    Nothing subscribed to it — reactive publishers are lazy and do nothing until subscribed. Either the Mono wasn't returned up to the framework (which subscribes), or an inner publisher was created but never composed into the returned chain (a 'lost' Mono). Never call subscribe() manually in a handler; return the publisher.
  • Why aren't passing StepVerifier unit tests enough to trust a WebFlux migration?
    StepVerifier checks the logical signal correctness of a single pipeline, not behavior under concurrency. The whole point of reactive is staying non-blocking at scale, which only load tests reveal — a stray blocking call passes unit tests but stalls the loop under real traffic.

saying these in an interview costs you the question

  • Assuming ThreadLocal-based Spring Security / MDC / transactions just work in WebFlux
  • Calling .block() inside a handler to 'simplify' code
  • Manually calling subscribe() in a controller instead of returning the publisher
  • Skipping BlockHound and load tests because unit tests are green
  • Treating reactive as a per-endpoint choice rather than a whole-chain commitment

context