An aggregation gateway composes a checkout confirmation page by calling an order-status service, which is required to render the page at all, and a recommendations service for a 'you might also like' section, which is a nice-to-have. If the recommendations call times out, should the gateway fail the whole request or return something else? Walk through how you'd design that behavior.
answer
- required vs optional dependency
- timeout shorter than overall SLA budget
- degrade, don't fail, on optional dependencies
- composite availability multiplies down
- flag a section as unavailable, don't silently blank it
basics
~20 sOnly fail the whole response if a truly required piece of data is missing. For optional extras like recommendations, return the page without that section, or with a placeholder, rather than blocking the whole page on a non-critical service.
solid answer
~40 sClassify each downstream dependency as required or optional for the aggregate response. A required dependency, order-status here, failing should produce an error response, since the page is meaningless without it. An optional dependency, recommendations here, should have a strict timeout, and on failure or timeout the gateway omits that section, substitutes a cached or stale value, or marks it as temporarily unavailable, then still returns a success response with the rest of the composed payload. This needs per-call timeouts shorter than the overall response budget, and the merge logic must treat a missing optional field as a normal, expected case rather than an exception.
go deeper
Should intuitively say that a 'nice to have' piece of data shouldn't block the whole page from loading.
Should describe required versus optional classification and that optional failures should degrade rather than fail the whole request.
Should reason about composite availability math, per-call timeout budgets, and how to signal degraded state to the client explicitly.
Should weigh the organizational risk of dependency classification drifting over time as new dependencies default to required, and design conventions or tooling, schema flags, per-dependency SLOs, that keep the classification correct as the system grows.
## Classify every dependency first Designing partial-failure behavior starts with classifying every downstream dependency the aggregator calls as either required or optional for that particular composed response. | Class | What it covers | |---|---| | **Required** | Required dependencies are ones without which the response has no meaning at all, order status on a checkout confirmation page being the clear example: if the gateway can't confirm the order was placed, there's nothing useful to show. | | **Optional** | Optional dependencies add value but aren't load-bearing, like a recommendations widget. | Each optional call gets its own timeout, strictly shorter than the gateway's overall response budget, and its own explicit fallback behavior defined ahead of time: - omit the field - serve a cached value - return a status flag saying the section is temporarily unavailable The merge step that assembles the final response must treat a missing optional field as a normal branch of logic, not as an exceptional case that needs a try/catch bolted on as an afterthought. ## Why the classification exists This design exists because treating every dependency as required multiplies failure probability across all of them. If five downstream services each have 99.9% independent availability and the gateway requires all five to succeed before returning anything, the composite availability is roughly 0.999 to the fifth power, about 99.5%, worse than any single dependency on its own. As more dependencies get added to a page over time, an all-required policy makes the aggregate endpoint strictly less reliable than its shakiest component. Classifying non-essential calls as optional and degrading gracefully on their failure protects the composite availability from this multiplication effect, so the checkout page stays up even when the recommendations service is having a bad day. ## What it costs, and what it buys The cost of building this is real complexity -- the gateway needs: - an explicit criticality classification for every dependency - per-branch error handling instead of one blanket try/catch - a convention for signaling to the client that a section is degraded rather than legitimately empty - tests that cover the combinations of which optional dependencies are up or down The payoff is a far higher composite availability and a much smaller blast radius for any single backend incident, which is usually worth the added design and testing surface for anything client-facing and revenue-relevant like a checkout flow. ## Failure modes 1. **A new dependency defaulting to required.** In production, a common failure mode is a new dependency getting added to a page and defaulting, by omission, to being treated as required, so a low-stakes addition like a marketing banner service accidentally makes the whole checkout page fail whenever the banner backend has any hiccup. 2. **Stale-cache fallback.** Another is stale-cache fallback with no expiry: during a prolonged outage, the gateway can keep serving increasingly outdated cached data indefinitely with nobody noticing, masking a real incident behind an apparently-working page. 3. **A timeout close to the overall budget.** A third is setting the optional dependency's timeout too generously, close to the overall budget, so every degraded response still pays nearly the full timeout cost before falling back, adding needless latency exactly when the system is already struggling. ## Where it shows up A well-known real-world pattern of this strategy is how large e-commerce product pages treat auxiliary widgets, recommendations, 'customers also bought,' review counts, as optional and non-blocking, so the core buy-and-checkout flow keeps rendering even if some auxiliary backend is degraded or entirely down. This is a deliberate, load-bearing architectural decision: the page's core commercial function is insulated from the availability of every ancillary feature bolted onto it, which is exactly the required-versus-optional classification applied at scale across dozens of composed sections rather than just two.
- If five downstream services each have 99.9% independent availability and the gateway requires all five to succeed to return any response, what is the gateway's resulting availability?Roughly 0.999 to the fifth power, about 99.5%, which is worse than any single dependency's own availability. Treating every dependency as required multiplies the failure probability across all of them, which is exactly why classifying non-essential calls as optional and degrading gracefully protects the composite availability.
- How should the gateway communicate to the client that a section of the response is degraded rather than just silently omitting it?A common approach is an explicit status or availability flag per section in the response schema, for example a recommendations field with a null value and a status of unavailable, so the client can render an appropriate placeholder or intentionally hide the widget instead of not knowing whether the field is legitimately empty or the underlying service failed.
- What risk does serving a stale cached value on optional-dependency failure introduce?If the cache has no expiry or staleness signal, the gateway can keep serving old data indefinitely during a prolonged outage without anyone noticing, masking a real incident. A better design caps how stale a fallback value is allowed to be and surfaces staleness in monitoring so a genuine ongoing outage doesn't hide behind an apparently-functioning page.
Like a tour guide who won't start the bus tour without the driver, a required dependency, but will happily leave without the gift-shop coupon flyer, an optional one, if it's not ready in time, rather than cancelling the whole tour.
saying these in an interview costs you the question
- treats every downstream dependency identically with no required/optional classification
- fails the entire response whenever any single call errors, regardless of that call's importance
- no timeout distinct from the overall response SLA for optional calls
- doesn't mention how the client is told a section is missing versus legitimately empty