skip to content

Describe the request-aggregation pattern at an API Gateway — where the gateway fans a single incoming client request out to multiple backend services and combines their responses into one — and the main risk this introduces compared to simple 1:1 routing.

level: seniorimportance: must knowfreq 55%

answer

  1. fan-out to N services, merge one response
  2. parallel calls bounded by slowest, not summed
  3. partial-failure policy: graceful degradation vs fail-all
  4. per-call timeout required
  5. combined availability multiplies (worse than any single service)

basics

~20 s

Instead of the client making five separate calls to five services, it makes one call to the gateway, and the gateway calls all five services itself, combines the answers, and sends back one response. The risk is that the slowest of those five calls determines how long the client waits, and if any one fails, the whole combined answer can fail too.

solid answer

~50 s

In aggregation, the gateway (or a dedicated aggregation layer behind it) receives one client request, issues multiple downstream calls — often in parallel — to different services, and merges the results into a single composed response, saving the client from making several round trips over a potentially slow network. This is distinct from plain path-based routing, where each request maps 1:1 to one service. The main risk is that the aggregated response is only as fast and as available as its slowest/least-reliable dependency, and a naive implementation that waits on all calls sequentially or without timeouts turns one flaky downstream service into a failure of the whole composed response. Production implementations parallelize the fan-out, set per-call timeouts, and decide a partial-failure policy (return partial data with a flag, or fail the whole request) rather than letting one dependency block or break everything.

go deeper

for a junior

Understands that the gateway can call multiple services for one client request and combine the answers into one response.

for a middle

Knows aggregation should be parallelized rather than sequential to control latency, and that timeouts are needed per downstream call.

for a senior

Reasons quantitatively about combined availability degradation and designs an explicit partial-failure policy per field/use case rather than defaulting to all-or-nothing.

for a principal

Decides organizationally where aggregation logic should live (shared gateway vs. dedicated composition service vs. BFF) based on how volatile and client-specific the aggregation logic is expected to become.

## What the gateway does with one incoming request Request aggregation is a pattern where the API Gateway (or a purpose-built aggregation/composition layer sitting just behind it) accepts a single incoming client request and, rather than forwarding it 1:1 to one backend service, issues multiple calls to several different backend services on the client's behalf, then merges those results into a single combined response before returning it. Mechanically, this typically works by the gateway: 1. **parsing the request**, determining which set of backend calls are needed to satisfy it, 2. **dispatching those calls** — ideally concurrently rather than one after another — 3. **waiting for them to return** (each with its own timeout), 4. and then **assembling a merged payload** that stitches together fields from each service's response before sending one reply back over the client's single connection. For example, a 'product detail page' request might need catalog data from a product service, current price from a pricing service, and stock level from an inventory service; aggregation lets the gateway make those three calls itself and return one combined object, rather than making the client open three separate connections and assemble the view itself. ## Why it exists: round trips are the expensive part This exists to solve a specific class of network-efficiency problem: over a slow, high-latency, or metered network — mobile clients are the canonical case, but any client far from the datacenter benefits — round trips are expensive, and every additional serial request adds a full network round-trip's worth of latency. If a client needs data from five services to render one screen, five sequential round trips over a mobile network (each potentially 100-300ms) can dominate perceived load time. Doing the fan-out server-side, inside the datacenter where inter-service latency is typically single-digit milliseconds, and returning one payload to the client collapses that cost dramatically. It also simplifies client code: the client doesn't need to know about five services' existence, their individual failure modes, or how to reconcile partial results — it gets one contract to code against. ## The trade-off: combined availability The core trade-off is a reliability and complexity shift. In plain 1:1 routing, the availability of a single client request depends on exactly one backend service being up; if it's not, that one feature fails and everything else keeps working. Once you aggregate, a single client request's success now depends on the combined availability of every service it fans out to — and by basic probability, calling five services each with 99.9% independent availability naively (as an unconditional AND) yields roughly 99.5% combined availability for the composed response, a meaningfully worse number, unless the aggregation layer is deliberately built to tolerate partial failure. This is why real aggregation implementations almost always: - **parallelize the fan-out** (so total latency is bounded by the slowest single call, not the sum of all calls), - **set an explicit timeout** per downstream call, and - **define an explicit partial-failure policy**: return the fields that succeeded with a null/omitted section for the ones that didn't (graceful degradation), versus failing the entire composed request if any dependency is missing (strict consistency) — the right choice is a business decision made per use case, not a default the gateway should silently pick. ## Failure modes in production Failure modes in production follow directly from ignoring these design points. - **A naive sequential fan-out** turns what should be a bounded-by-the-slowest-call latency into a sum-of-all-calls latency, and this is a very common first-implementation bug. - **A missing per-call timeout** means one hung downstream dependency can hold the whole aggregated request open indefinitely, exhausting the gateway's own connection pool under load — the same cascading-overload pattern seen with synchronous auth checks, just triggered by a different concern. - **An aggregation layer with no partial-failure policy** tends to fail the entire user-facing feature (blank page) because of a single non-critical dependency being down, when a graceful-degradation policy would have been the better default for many use cases. ## Aggregation versus Backend-for-Frontend It's worth being precise about scope here: generic aggregation — combining several services' data into one response for efficiency — is what's described above, and is a gateway-level concern regardless of client type. Tailoring a distinct aggregated response shape per client type (a lean payload for mobile vs. a richer one for a desktop web client) is a related but separate pattern usually implemented as a **Backend-for-Frontend** rather than in the shared, general-purpose gateway. ## Where you have seen it A concrete real-world example of plain aggregation: **Netflix's** edge API layer historically composed data from dozens of backend services into single responses for device clients, driven precisely by the mobile/TV network-efficiency problem described above.

  • If aggregation calls five downstream services each with 99.9% independent availability, why is the naive combined availability of the aggregated response worse than any single one?
    If failures are independent and the aggregation requires all five to succeed, the probability all five are up simultaneously is roughly 0.999^5, about 99.5% — worse than any individual 99.9% service. This is a basic AND-of-independent-probabilities effect, and it's the mathematical reason naive full-dependency aggregation degrades overall reliability.
  • How would you decide whether a composed response should fail entirely or return partial data when one downstream call fails?
    It's a product/business decision tied to how essential that piece of data is to the response's usefulness — e.g., missing a 'related products' section is fine to omit, but missing the price on a checkout page is not. Critical fields get strict fail-fast handling, non-critical/enrichment fields get graceful degradation with the field omitted or nulled and ideally a flag signaling partial data.
  • Why is it important to parallelize the fan-out calls rather than making them sequentially?
    Sequential calls sum their latencies, so five 50ms calls become a 250ms response; parallel calls are bounded by the slowest single call, so the same five calls return in roughly 50-60ms. This difference compounds directly into user-perceived latency on every aggregated request.

Like ordering a combo meal from one cashier who relays your order to the grill, the fryer, and the drink station behind the counter and hands you one tray — versus you personally walking to three separate counters and assembling your own tray.

saying these in an interview costs you the question

  • Doesn't recognize that combined availability across dependencies is worse than any single dependency's availability
  • Assumes sequential fan-out and parallel fan-out have the same latency characteristics
  • No concept of a partial-failure/graceful-degradation policy — assumes all-or-nothing is the only option
  • Confuses generic response aggregation with client-specific tailoring (BFF)
  • Forgets to set per-call timeouts, assuming downstream calls always return quickly

context