skip to content

A checkout page normally shows personalized product recommendations pulled from a separate recommendation service. If that service is slow or completely down, what should the checkout page do instead of failing the whole page, and what is the general name for this strategy?

level: juniorimportance: must knowfreq 65%

answer

  1. hide non-critical widget, keep core flow
  2. fallback = substitute response
  3. stale cache vs default value vs skip
  4. must-have vs nice-to-have features
  5. mask failure from user, not from monitoring

basics

~20 s

Skip or replace the broken part instead of crashing everything — show checkout without recommendations, or with generic ones, so the customer can still buy. This is called graceful degradation, and the substitute response is a fallback.

solid answer

~40 s

The page should catch the recommendation-service failure at the call site, skip or hide that widget, and render the rest of the page normally — checkout is the core flow, recommendations are a nice-to-have. This is graceful degradation: the system keeps its critical path working while shedding non-critical features when a dependency is unavailable. The fallback itself can be 'nothing' (hide the widget), a static default (e.g. 'bestsellers this week'), or a cached last-good response. The key design point is isolating the failure so it can't propagate — a slow recommendation call shouldn't hold open a thread/connection or block the checkout request, so it needs its own timeout and shouldn't be awaited synchronously if avoidable.

go deeper

for a junior

Should recognize that failing a non-critical dependency shouldn't fail the whole request, and can name at least one concrete fallback (default value, cached data, or hiding the feature).

for a middle

Should be able to implement the try/timeout/fallback pattern in code, know to add a timeout so the fallback doesn't get delayed too, and know to log/monitor when the fallback fires.

for a senior

Should proactively design which features are fallback-eligible during system design, weigh staleness vs correctness for each, and call out the risk of fallback masking a real outage.

for a principal

Should set org-wide policy for what 'critical path' means, ensure fallback usage is tracked as an SLO/error-budget signal, and prevent scope creep of fallback logic into correctness-sensitive paths.

## What graceful degradation is Graceful degradation is the practice of designing a system so that, when a dependency it needs becomes slow or unavailable, it deliberately turns off or downgrades the features that depend on it while keeping the rest of the system working. A **fallback** is the substitute value or behavior used in place of the real response: - **nothing at all** — simply hide a UI element; - **a static default** — a generic message or precomputed list; - **a previously cached response** that is now stale. Mechanically, the calling code wraps the risky call in error handling — a `try/catch`, a timeout, or a resilience-library decorator — and on failure (or timeout) branches to the fallback path instead of propagating the exception up to the user. ## Why it exists Without a deliberate fallback, a failure in a small, non-essential component often takes down the entire request because most code paths assume every call succeeds and simply throw when one doesn't. In a system built from many independent services, any one of dozens of dependencies can be degraded at any moment; if every dependency failure is allowed to fail the whole request, overall availability multiplies down fast (a page depending on 10 services each 99.9% available is only around 99% available end-to-end unless failures are isolated). Fallback and graceful degradation exist to decouple the availability of the whole request from the availability of every individual piece, so the business gets to choose which features are 'must work' versus which are 'nice to have': | Tier | Belongs there | |---|---| | 'must work' — the core flow | e.g. checkout, login, search | | 'nice to have' | recommendations, reviews count, avatar images | ## The trade-off The core trade-off is **correctness/freshness versus availability**. - A **cached fallback** keeps the request succeeding but may show data that's minutes or hours old — fine for 'bestseller list,' dangerous for 'current inventory count' or 'account balance.' - A **static default fallback** (e.g. a generic shipping estimate) is always available and cheap but can mislead users if not clearly labeled as an estimate. - There's also an **engineering cost**: every fallback path is more code, another thing to test, and another behavior that can silently drift out of sync with the primary path. Finally, using a fallback masks the underlying failure from the end user, which is good for UX but risks masking it from the engineering team too if it isn't paired with monitoring. ## Failure modes in production 1. **Silent staleness** — the most common production failure: the fallback keeps firing far longer than intended because nobody notices the primary path is broken, since users never see an error. 2. **Scope creep** of what's considered 'safe to degrade' — teams start with recommendations as fallback-able, then under pressure quietly add fallback logic to something that actually needs correctness (e.g. defaulting a payment amount), which can cause real financial or safety incidents. 3. **Treating the fallback path as an untested afterthought** — a third failure mode: because it only executes during dependency failure, it can accumulate bugs (null pointer errors, wrong cache keys) that surface for the first time during an actual incident, doubling the outage. ## Where it shows up - **Netflix** is a well-known example. Its Hystrix library (and successors like resilience4j) popularized wrapping every remote call with a fallback method — if a 'similar titles' service fails, the UI falls back to a generic, precomputed list rather than showing a blank row or an error page; the home page keeps rendering. - **Amazon's** retail site is often described in engineering talks as prioritizing 'can add to cart and check out' above almost everything else — reviews, recommendations, and 'customers also bought' widgets are all designed to degrade or disappear independently without touching the buy button. This tiered design — core transactional flow protected, decorative/personalization features degrade first — is the general pattern behind graceful degradation in large-scale systems.

  • How would you decide, in a new checkout flow, which sub-features are allowed to have a fallback and which are not?
    Draw the line at anything the core business transaction depends on for correctness or legality — payment amount, item availability, shipping address — those must fail loudly rather than default silently. Anything purely presentational or personalization-related (recommendations, reviews, ratings) is a safe fallback candidate. In practice teams tag each dependency call as 'critical' or 'best-effort' in code/config so the policy is explicit rather than ad hoc.
  • If the checkout page hides the recommendation widget on failure, how do you make sure engineers still find out the recommendation service is down?
    Emit a metric or log every time the fallback path executes, and alert on a spike or sustained rate of fallback usage rather than only on hard errors — the failure is invisible to users but should not be invisible to observability. Many teams also tag the response with an internal 'degraded' flag so it can be traced even though the user-facing page looks normal.
  • What's the risk of always returning a hardcoded default value (like 'estimated delivery: 5 days') when the shipping-estimate service is down?
    It can quietly become wrong for a large share of real orders (e.g. during a regional outage that also affects real delivery times), misleading customers and creating support/refund costs. Defaults for anything financial or time-committed need to be conservative and periodically reviewed, not just picked once and forgotten.

Like a restaurant that's out of a menu item: the kitchen doesn't shut down for the whole table, it just tells you that dish is unavailable and offers a substitute so you can still eat.

saying these in an interview costs you the question

  • Says the whole page/request should just error out instead of degrading
  • Can't name a concrete fallback response type (cache/default/skip)
  • Assumes fallback logic doesn't need its own testing
  • Doesn't distinguish which features are safe to fall back on vs not
  • Thinks masking the failure from the user means no monitoring is needed

context