skip to content

A junior engineer implements an aggregation gateway endpoint that calls a pricing service, an inventory service, and a reviews service, then merges the three JSON responses into one payload. In their code, each call uses a blocking HTTP client and is invoked one after another inside a single method. What is wrong with this implementation, and what would fix it?

level: middleimportance: must knowfreq 70%

answer

  1. Promise.all, not sequential awaits
  2. join once, not one at a time
  3. bounded pool sizing, not unbounded threads
  4. timeout per call, not per whole request
  5. salvage partial results on join failure

basics

~20 s

Calling each service one after another means the gateway waits for all of them added together, which is slow. Fixing it means starting all three calls at roughly the same time and only waiting once, so the total wait is about as long as the slowest single call.

solid answer

~40 s

The implementation defeats the purpose of aggregation: instead of a wait roughly equal to one round trip, the client now waits for the sum of three sequential calls. The fix is to issue the three downstream requests concurrently -- using async/await, CompletableFuture.allOf, a reactive Mono.zip/Flux.merge, or parallel tasks on a bounded pool -- and join on all three results before composing the response. Concurrency also needs bounding, a sized thread or connection pool rather than unbounded threads-per-request, so a burst of aggregate requests doesn't exhaust downstream connections, and each call needs its own timeout so the join doesn't hang indefinitely on one slow dependency.

go deeper

for a junior

Should recognize that calling services one after another is slower than calling them together, even without knowing the exact concurrency primitives involved.

for a middle

Should name at least one concrete concurrency mechanism, async/await, Promise.all, CompletableFuture, or a reactive zip operator, and explain that the total wait becomes roughly the slowest call, not the sum.

for a senior

Should discuss bounded pool sizing, per-call timeouts, and how to salvage partial results when one branch of the join fails rather than discarding everything.

for a principal

Should connect fan-out concurrency choices to platform-wide resource budgets, how many concurrent aggregate requests the gateway can sustain given downstream connection limits, and weigh reactive versus thread-per-call models at that scale.

## What concurrent fan-out requires Making a downstream fan-out actually concurrent requires a concrete mechanism, not just "calling three services." - **In a thread-based stack** this typically means submitting each call as a task to an executor and getting back a future (`CompletableFuture.supplyAsync` for each call, then `CompletableFuture.allOf(f1, f2, f3)` to join), or using async/await in languages that support it so each call starts before the previous one's result is needed. - **In a reactive stack** it means composing publishers with an operator like `Mono.zip` or `Flux.merge` so subscriptions to all three sources happen up front and the pipeline completes when all have emitted. Either way, the key property is that the three network calls are all in flight at the same time, and the gateway blocks or awaits exactly once, on the join, rather than three separate times. ## Why it matters This matters because the entire value proposition of aggregation is amortizing round-trip cost. If the fan-out inside the gateway is accidentally sequential, the aggregate response time becomes the sum of the three call latencies instead of roughly the maximum of them. A gateway meant to turn three 100ms calls into one 100ms-ish response instead turns them into a 300ms response, which can be slower than letting the client call the three services directly if they happen to be reachable in parallel from the client too. The whole architectural investment in an aggregation layer is wasted by this one implementation mistake. ## What correct concurrency costs Building correct concurrency has real costs beyond just calling the right API. - The gateway needs a **bounded resource pool** -- threads, connections, or both -- sized for its expected concurrent load, because unbounded per-request thread or connection creation degrades badly under traffic spikes. - **Error handling** also becomes more delicate: when three calls run concurrently and one throws, the code must decide whether to salvage the other two results or fail everything, and naive use of a joining primitive like `allOf().join()` will propagate the first exception and discard the rest unless each future's result is inspected individually with something like `exceptionally()` or `handle()`. - **Tests** get harder too, since concurrent code can have race conditions and non-deterministic ordering that simple sequential code never exhibits. ## Failure modes 1. **The sequential-await bug.** In production, a common failure mode from get-it-working-quickly code is exactly the sequential-await bug described in this question: it passes functional tests, since the merged output is correct, and only shows up as a latency regression once load testing or real users notice the endpoint is slower than expected. 2. **Forgetting per-call timeouts.** A second failure mode is forgetting per-call timeouts even after fixing the concurrency: if the join waits on all three futures with no individual timeout, one hung downstream call can stall the entire gateway response indefinitely, converting a should-be-fast concurrent fan-out into a request that never returns. 3. **Unbounded thread creation.** A third, more insidious failure is unbounded thread creation per request; under a traffic spike this can exhaust memory or hit OS thread limits, crashing the gateway process instead of gracefully shedding load. ## Where it shows up A well-known real-world illustration is Netflix's historical use of Hystrix, where each downstream call in an aggregation flow was wrapped as a Hystrix command executing on its own dedicated thread pool, with results composed via futures or reactive Observables once all commands resolved or timed out -- explicit infrastructure built specifically to make concurrent fan-out safe and boundable at scale. In simpler modern stacks, a Node.js gateway using `Promise.all([callA(), callB(), callC()])` or a Spring WebFlux gateway using `Mono.zip(callA(), callB(), callC())` achieve the same concurrent-fan-out-then-join shape with far less code, but the underlying requirement -- start all calls before awaiting any of them, bound the resources involved, and time each call out independently -- is identical regardless of the specific framework.

  • If an aggregation gateway uses CompletableFuture.allOf(future1, future2, future3).join() and future2 throws an exception, what typically happens to future1 and future3's results?
    allOf itself completes exceptionally once any of the futures fails, so join() on it will throw -- but future1 and future3 keep running independently in the background unless explicitly cancelled. The calling code must inspect each individual future separately, for example with exceptionally() or handle(), to still retrieve the successful results from future1 and future3 rather than discarding everything just because future2 failed.
  • Why does unbounded thread creation for downstream calls become dangerous under load?
    Unbounded thread creation means each incoming aggregate request spawns several new threads, so a traffic spike multiplies threads faster than the OS or JVM can efficiently schedule them, causing context-switch overhead and memory pressure from thread stacks that can crash the process outright rather than degrading gracefully. A bounded pool with backpressure or request shedding fails more predictably under the same load.
  • How should each downstream call's timeout relate to the gateway's own overall response-time budget?
    Each downstream call's timeout must be strictly less than the gateway's overall SLA budget, with some margin left over for merging and serialization work. If one call's timeout equals or exceeds the whole request's budget, that single dependency can single-handedly blow the gateway's own response-time target even when the other calls return promptly.

Like ordering three dishes at once from three different kitchen stations instead of waiting for the appetizer to fully finish before even telling the kitchen you want a main course -- you place all three orders immediately and eat once everything's plated.

saying these in an interview costs you the question

  • writes the fan-out as sequential awaits and calls it aggregation
  • no timeout on individual downstream calls, only an overall request timeout if any
  • one failed future aborts the whole join with no attempt to salvage the others' results
  • proposes unbounded thread-per-call with no discussion of pool sizing

context