A team is deciding whether to build a new service's asynchronous layer around demand-driven reactive streams, around suspending sequential code (coroutine or async-await style), or around plain futures dispatched on an event loop. How would you reason about that choice?
answer
- Cardinality first: one result vs many over time
- Is a rate mismatch real, or does the request cycle pace everything?
- Blocking driver anywhere = complexity without benefit
- Operator pipelines cost stack causality and debuggability
- One default per service; futures as interop currency
basics
~20 sChoose by problem shape. One result per request means futures or suspending code. Many items over time with a real rate mismatch and a bounded-memory requirement justifies streams. Weigh debuggability, the viral spread of the model through signatures, and whether the whole path is genuinely non-blocking.
solid answer
~60 sI would ask five questions before naming a technology. 1. **Cardinality.** Request and response is a single outcome — futures or suspending code. Streams earn their complexity only where items arrive over time. 2. **Is flow control a real requirement?** If every stage is paced by an inbound request cycle, demand signalling buys nothing. If a fast source can outrun a slow sink and memory must stay bounded under overload, it buys the property that matters most in an incident. 3. **Is the whole path non-blocking?** Without non-blocking drivers end to end you hop to a thread pool anyway and keep the complexity while losing the benefit. 4. **Debuggability and onboarding.** Operator pipelines lose stack causality and step-through debugging; sequential suspending code reads like blocking code and keeps a usable causal trace. That is a running cost paid by every engineer. 5. **Virality and uniformity.** The model spreads through every signature, and mixed models need boundary bridges that are where deadlocks live. Pick one default per service. My default: suspending sequential code for request and response, streams for genuinely streaming or rate-mismatched paths, futures as the interop currency between libraries.
go deeper
Not expected to lead this call; be able to say that a stream is for many items over time and a future or suspending call is for a single result.
Compare on concrete grounds — cardinality, whether backpressure is needed, and how readable the resulting code is — rather than by naming a framework.
Bring the operational evidence: blocking audit of dependencies, incident history, observability of queues and demand, and the cost of bridging models.
Own the decision: state the axes, commit to a default with named exceptions, define the non-negotiables such as bounded queues and cancelling timeouts, and describe how you would revisit the choice with evidence.
## Decide on problem shape, not on fashion All three options solve *do not park a thread while waiting*. They differ in what else they impose. The failure mode of this decision is choosing the most powerful model for the whole service and paying its cost on the ninety percent of code that never needed it. ## The axes that actually decide it ### 1. Cardinality of the result Single outcome per operation — a lookup, a command, an RPC — is future or suspension shaped. A sequence over time with a rate the consumer does not control is stream shaped. Most services are overwhelmingly the former with a few of the latter, which argues for a default plus exceptions rather than a single model everywhere. ### 2. Whether flow control is a live requirement Demand signalling exists to keep in-flight work bounded when a producer can outpace a consumer. Ask concretely: is there a path where the source's rate is independent of your consumption rate? Ingest from a broker, change feeds, large scans, fan-out to slow sinks — yes. A handler that makes two calls and returns a response — no, because inbound concurrency limits already pace everything. Adopting streams for flow control you do not need is pure cost. ### 3. End-to-end non-blocking reality The benefit depends on the *whole* path. If a driver, an authentication library or a serializer blocks, you hop to a bounded thread pool at that point and inherit its queue and its saturation behaviour. You then carry the full complexity of the async model while your capacity is still set by that pool. Audit the dependencies before deciding; this single check has reversed many of these decisions. ### 4. Debuggability, and who pays for it This is the most underweighted axis. Operator pipelines break the correspondence between the physical stack and the causal history: a failure shows the delivery machinery rather than the call that started the work, thread-scoped context does not follow the data, and stepping through a pipeline in a debugger is genuinely hard. Sequential suspending code preserves an ordinary reading order and usually a reconstructed causal trace. Multiply the difference by every engineer, every incident, and every new hire. ### 5. Virality and mixing Asynchrony spreads through signatures: one suspending or stream-returning call forces the shape on every caller up the chain. Mixing models means bridges at the boundary, and bridges are where threads get parked, pools get exhausted and deadlocks appear. Choose one default per service, make the boundaries explicit, and put the bridges in as few places as possible. ### 6. Ecosystem and operational fit Do your libraries, tracing, metrics and profilers understand the model? Does the framework let you express a per-stage concurrency limit? Can you observe demand and queue depth, which is what you will need at three in the morning when latency climbs? A model you cannot observe is a model you cannot operate. ## What the evidence usually says Ask what has actually hurt this team in production. If incidents are *out of memory because a queue grew*, demand signalling addresses the root cause directly and is worth its price on the affected paths. If incidents are *thread pool exhausted because everything was blocked*, any non-blocking model helps and the simplest one wins. If incidents are *nobody could work out what this pipeline does*, the answer is not more operators. ## A defensible position - **Default:** sequential suspending code for request and response work. It preserves reading order and debuggability, gives structured lifetimes and cancellation, and covers the large majority of handlers. - **Streams:** for paths that are genuinely streaming, unbounded in rate, or multi-stage with a mismatch — ingest, exports, fan-out, subscriptions. Introduce them as a bounded region with an explicit boundary rather than as the service-wide idiom. - **Futures:** as the lingua franca between libraries and models. Almost every ecosystem can produce and consume them, so they are the cheapest interop currency; keep them at seams rather than as the primary programming style. - **Non-negotiables regardless of choice:** bounded queues everywhere, an explicit concurrency limit per stage, timeouts that cancel, cancellation that propagates, and end-to-end context propagation that does not depend on thread affinity. ## How to present this in an interview Do not answer with a technology. Give the axes, state which evidence you would collect — incident history, the blocking audit of dependencies, the shape of the traffic — and then commit to a defensible default with named exceptions. The signal being looked for is that you can price complexity against the property it buys, and that you know reactive streams buy bounded behaviour under overload rather than throughput.
- What evidence would change your mind toward adopting reactive streams service-wide?Incident history dominated by unbounded queue growth and memory exhaustion under load spikes, plus traffic that is genuinely streaming rather than request and response, and a dependency audit showing a non-blocking path end to end. If the same incidents were thread pool exhaustion or plain slow downstreams, a simpler non-blocking model addresses them at far lower cost.
- What is the cost of mixing two async models in one service?Every boundary between them needs a bridge, and bridges are where a thread gets parked to wait for the other model to finish. Those points are the classic sources of pool exhaustion and self-deadlock, they hide from load tests, and they fragment context propagation and error semantics. If mixing is unavoidable, isolate the bridges in a small number of reviewed places with their own dedicated, bounded resources.
saying these in an interview costs you the question
- Choosing the model by popularity or by claimed throughput rather than by the shape of the workload
- Assuming reactive streams increase throughput, when the property they buy is bounded behaviour under overload
- Overlooking a blocking driver in the path, which nullifies the benefit while keeping all the complexity
- Treating debuggability and onboarding cost as unimportant next to a technical benchmark
- Mixing models freely across a codebase without acknowledging that the bridges are where deadlocks live