skip to content

A team adds a fallback so that whenever the primary recommendation service is slow, requests fall through to a secondary 'simple trending list' service instead. During a major primary outage, the secondary service — which was sized only for occasional traffic — collapses under the full redirected load, and now both are down. What went wrong with this fallback design, and how would you fix it?

level: seniorimportance: should knowfreq 45%

answer

  1. fallback traffic = sudden full-load spike on secondary
  2. capacity-plan the fallback like a real dependency
  3. rate-limit/shed in front of the fallback
  4. cascading/fallback-stampede anti-pattern
  5. final cheap fallback-of-last-resort to stop the chain

basics

~20 s

The backup service wasn't built to handle all the traffic at once, so when everyone switched to it, it broke too. Fix it by making sure the fallback can actually handle full load, or by rate-limiting/shedding instead of just dumping everything on it.

solid answer

~40 s

The team treated the fallback as a magic escape hatch without capacity-planning it as a real dependency in its own right — it was sized for the occasional trickle of requests during brief blips, not for 100% of primary traffic during a sustained outage. This is the same problem as the primary having a single point of failure, just moved one hop over. The fix is to either provision the fallback for worst-case full-traffic load, add rate-limiting/load-shedding in front of the fallback so it degrades gracefully instead of collapsing, or make the fallback itself cheaper than the primary (e.g. a static cached list requiring near-zero compute) so it scales to full traffic naturally. I'd also want a fallback-of-the-fallback (e.g. an empty/static default) so a secondary failure doesn't cascade further.

go deeper

for a junior

Should recognize in plain terms that dumping all failed traffic onto a backup that wasn't sized for it can break the backup too.

for a middle

Should propose rate-limiting or capacity-planning the fallback and giving it its own timeout, when prompted.

for a senior

Should proactively identify this failure mode when designing a fallback chain, propose layered mitigations (capacity plus rate limiting plus independent timeout plus last-resort default), and know it as a named anti-pattern.

for a principal

Should mandate failover load-testing/chaos drills as part of the resilience review process org-wide, and weigh the cost trade-offs (always-on redundant capacity vs. graceful shedding) as an architectural policy decision.

## Why the fallback becomes the next outage A fallback that redirects failed traffic to a secondary service is architecturally identical, from the secondary's point of view, to a sudden, sustained traffic spike: every request that would have gone to the primary now goes to it instead, for as long as the primary stays unhealthy. If the secondary wasn't explicitly capacity-planned and load-tested for 100% of primary's traffic indefinitely, it experiences the same kind of overload the primary just suffered — CPU/memory exhaustion, connection pool saturation, queueing delay that itself starts timing out — and can fail in exactly the same way, sometimes faster, because it typically has fewer resources than the primary it's backing up. ## Why teams miss it This happens because fallback design is often done at the level of 'what value do we return on failure' rather than 'what does this fallback path cost, and can the target sustain that cost.' Teams reasonably capacity-plan the primary path carefully, since that's the one carrying traffic every day, but treat the fallback as a rarely exercised code path and under-invest in its operational readiness — it's tested for correctness (does it return a reasonable list) far more often than it's load-tested for capacity (can it survive an hour of 100% traffic). ## The trade-off There's a real cost trade-off in fixing this properly: provisioning a secondary service to handle 100% of primary traffic means paying for capacity that, most of the time, sits idle — the same always-on redundant capacity cost that motivates active-active architectures generally. The cheaper alternatives all carry their own downside: | Cheaper option | What it costs you | |---|---| | rate-limiting/shedding requests to the fallback | protects it from collapse but means some fraction of users get a harder failure (empty state or error) instead of the fallback content during a bad outage | | making the fallback deliberately cheaper (e.g. a static file instead of a live service) | reduces personalization/quality of the fallback response | There's no free option — every mitigation trades either money, user experience, or engineering complexity for safety. ## Failure modes in production 1. **In production, this shows up as a cascading outage**: primary fails, traffic shifts to secondary, secondary fails from the redirected load, and now there's no fallback left, often producing a worse outage than if the fallback hadn't existed at all, because the secondary might also have been serving some of its own, unrelated legitimate traffic that's now also impacted. 2. **A related, subtler failure mode is a fallback that 'works' during testing** (small-scale failure injection) but has never been tested under full-traffic failover, so its true capacity ceiling is unknown until the real incident. 3. **Another is a fallback with no independent circuit breaker/timeout of its own** — if it starts to degrade under the redirected load, requests hang waiting on the now-also-struggling secondary instead of failing fast to a final, cheap default. ## The layered fix This is a well-known anti-pattern sometimes discussed under retry storms/cascading failure, and it's part of why resilience libraries such as Hystrix historically, and resilience4j today, pair fallbacks with bulkheads and rate limiters rather than leaving the fallback's inbound traffic unbounded. Well-architected resilience guidance from major cloud providers specifically calls out designing fallback/backup paths with the same rigor (capacity planning, load testing, independent scaling) as the primary path, precisely because of incidents like this one. The concrete fix is layered: 1. capacity-plan and load-test the secondary for full-failover traffic, or explicitly cap it with a rate limiter/load shedder so it degrades — serving a static default to excess requests — rather than collapsing; 2. give the fallback call itself a timeout and circuit breaker so the primary system doesn't queue up waiting on a struggling secondary; 3. add a final, near-zero-cost fallback-of-last-resort (a static/empty response) so the failure chain terminates instead of cascading indefinitely.

  • How would you load-test a fallback path to catch this problem before it happens in production?
    Run a controlled failover drill (chaos engineering) where you deliberately fail the primary in a staging or canary environment and drive full realistic production-level traffic at the fallback for a sustained period, not just a quick smoke test, measuring whether it holds up over minutes/hours, not seconds.
  • If provisioning the secondary for 100% of primary's traffic is too expensive, what's a cheaper mitigation?
    Put a rate limiter or load shedder in front of the fallback so it accepts only as much traffic as it can safely handle and serves a cheap static default (or a clear degraded response) to the overflow, rather than accepting unlimited redirected traffic and collapsing under it.
  • Why does the fallback call itself need its own timeout and circuit breaker, separate from the primary's?
    If the secondary starts to struggle under the redirected load, calls to it can hang just like the primary did; without its own timeout/breaker, requests queue up waiting on a now-also-failing dependency, delaying the fall-through to a final cheap default and prolonging the outage instead of degrading quickly.

Like everyone in a stalled elevator bank taking the one working staircase at once — the staircase wasn't built for the whole building's foot traffic, and now it's jammed too.

saying these in an interview costs you the question

  • Assumes a fallback service doesn't need its own capacity planning/load testing
  • No mention of rate-limiting or shedding traffic into the fallback
  • Doesn't recognize the fallback needs its own timeout/circuit breaker independent of the primary's
  • Treats this as unfixable/inherent rather than a capacity-planning gap
  • No fallback-of-last-resort to stop a cascading chain

context