skip to content

A mobile app screen needs data from five different backend microservices to render fully. Instead of having the app call all five services directly over the network, what does putting a gateway aggregation pattern in front of them do, and why does that help?

level: juniorimportance: must knowfreq 75%

answer

  1. one request in, many out
  2. fan-out then fan-in
  3. high-latency network problem
  4. composite response, single hop for the client
  5. BFF is the per-client cousin

basics

~20 s

It puts one gateway in front of many services. The gateway makes all the backend calls itself and sends the client one combined response, so the client needs only one request instead of many slow round trips.

solid answer

~40 s

Gateway aggregation collapses several backend calls into a single client-facing endpoint. The client sends one request; the gateway fans it out to the needed services, ideally in parallel, waits for their responses, merges the results into one payload matching what the client actually needs, and returns it. This matters most over high-latency links like mobile networks, where each round trip is expensive: five sequential client-to-service calls at 200ms each is roughly a second, while one client-to-gateway call plus parallel fan-out inside a low-latency datacenter can finish well under 250ms. It also shrinks how much backend topology the client has to know about and reduces data/battery use on constrained devices.

go deeper

for a junior

Should be able to state that the client makes one call instead of many and that the gateway is the thing making the multiple backend calls on the client's behalf.

for a middle

Should additionally know that the fan-out calls need to run in parallel to get the latency benefit, and give a rough sense of why network round-trip cost matters more on mobile.

for a senior

Should reason about partial-failure handling, per-call timeouts, and when composition logic belongs in a shared gateway versus a per-client BFF.

for a principal

Should discuss where aggregation sits in the wider system architecture, its coupling and ownership costs across teams, and when a different approach (client-side composition, GraphQL, event-driven materialized views) is a better fit than a hand-rolled aggregator.

## What the gateway actually does A gateway aggregation pattern places one component -- often called an **aggregating gateway**, or in its client-specific form a **Backend-For-Frontend (BFF)** -- between a caller and a set of backend services. The client issues a single request describing what it needs, for example "everything for the product detail screen." The gateway translates that into a set of downstream calls: - one to a pricing service - one to inventory - one to reviews - one to recommendations - one to shipping estimates It issues these concurrently rather than one after another, collects each response as it arrives, merges the pieces into a single composite payload shaped the way the client actually wants it, and only then replies. From the client's point of view the multi-service fan-out is invisible: it made one call and got back one document. ## Why the pattern exists This pattern exists because a naive "smart client, dumb backend" design makes the client responsible for knowing every service it needs and calling each one directly. That is fine on a fast local network, but mobile networks, satellite links, and cross-region calls each carry real round-trip latency, often 100-300ms. A screen needing five sequential calls can burn well over a second of pure network wait before any processing happens. Aggregation moves the fan-out to a location -- typically the same datacenter or availability zone as the backend services -- where inter-service latency is a few milliseconds, so the client pays for one expensive round trip instead of five. ## The trade-off The **upside** is fewer round trips and a simpler client: - less orchestration code - fewer loading states to manage - no client-side merging logic - less data/battery burned on constrained devices The **downside** is that the gateway becomes a mediator with its own deployment, on-call burden, and failure surface. It must be kept in sync with every backend service it composes, it adds one more network hop and one more piece of infrastructure to operate, and it can become a coupling point where several teams' changes must be coordinated through its composition logic. Overused across too many screens, it sprawls into a maze of bespoke composition endpoints that are expensive to maintain. ## Failure modes 1. **Partial failure.** In production, the sharpest failure mode is partial failure: four of five downstream calls succeed and one times out or errors. A naive aggregator that fails the whole response whenever any single call fails means an otherwise-healthy response gets torpedoed by one flaky dependency. Teams typically mitigate this by returning partial data -- omitting or flagging the failed section, or substituting a cached value -- and by putting a strict per-call timeout on each downstream request so one slow service cannot stall the entire response. 2. **A latency floor.** A second failure mode is a latency floor set by the slowest call if fan-out is truly parallel, or by the sum of all calls if a bug turns the fan-out sequential -- a common real mistake is writing the calls in a loop that awaits each one before starting the next, silently defeating the pattern's whole purpose. ## Where it shows up A well-known real-world case is how Netflix's API/edge layer historically composed device-specific home-screen responses by calling many backend services -- catalog, personalization, viewing history, artwork selection -- and merging them into one payload tailored to each device, because a TV app, a mobile app, and a web client each needed a different shape and subset of that data and could not afford dozens of round trips per screen render. Similarly, an e-commerce product page gateway might fan out to pricing, inventory, and reviews services in parallel and stitch the three responses into the single JSON object the page's rendering code expects, turning three mobile round trips into one.

  • What happens if the reviews service call times out but the other four downstream calls succeed?
    A well-designed aggregator returns a 200 with the four available sections and either omits the reviews section, marks it as temporarily unavailable, or serves a cached value, rather than failing the whole request. Failing the entire response because one non-critical dependency timed out turns a minor backend blip into a full client-facing outage, which defeats much of the pattern's resilience benefit.
  • Why does the fan-out need to happen in parallel rather than sequentially?
    If the gateway calls each downstream service one after another, the aggregate response time becomes the sum of the individual call latencies, which can be worse than several of the original client-to-service round trips combined. Calling them concurrently, via async/await, reactive streams, or a bounded thread pool, makes the total response time roughly equal to the single slowest call, which is the whole point of relocating the fan-out inside a low-latency network.
  • Does gateway aggregation reduce the total number of network calls made?
    No -- the same number of calls to the backend services still happen, they are just relocated. What decreases is the number of high-latency round trips the client itself has to make; the fan-out calls happen over a fast internal network instead.

Like a restaurant server who takes one order from the table, relays separate parts to the grill, the bar, and the salad station, then brings back a single tray with everything plated together -- the diner never has to walk to three counters.

saying these in an interview costs you the question

  • says aggregation reduces total backend calls rather than client round trips
  • designs the fan-out as a sequential loop and calls it aggregation
  • fails the whole response whenever any one downstream call fails
  • confuses aggregation with routing a request to the correct single backend
  • no mention of timeouts on the downstream calls

context