skip to content

You're designing the resilience strategy for a video-streaming home page that renders: (1) 'continue watching' with playback position, (2) personalized recommendation rows, (3) trending/most-popular rows, (4) user profile avatar and account menu. Under heavy backend stress, which of these would you shed first and which would you protect longest, and how would you implement the prioritization?

level: seniorimportance: must knowfreq 70%

answer

  1. rank by core-to-task + cost-to-compute
  2. bulkhead isolation prevents cross-section starvation
  3. shed before failure via load signals, not just react
  4. product must agree on shedding order
  5. Netflix/Amazon independent widget fallback

basics

~10 s

Keep the things people actually came for (continue watching, account access) working the longest, and turn off the fancier extras (personalized recommendations, trending rows) first since the page still works fine without them.

solid answer

~50 s

I'd rank by how core each feature is to the primary job of the page — getting the user back into a show — and how expensive each is to compute. Continue-watching and account/profile access are core and cheap (mostly lookups), so they're protected longest. Personalized recommendations are the most expensive (real-time ML scoring) and least essential, so they're the first to shed — falling back to a cheap, precomputed trending/popular list. Trending is itself a good fallback target for recommendations but would be the next thing shed if things got worse, replaced by a fully static row. I'd implement this with independent timeouts and circuit breakers per section so one slow section can't block the others, feature flags to disable a section outright under load, and a load-shedding policy (shed in priority order as error rate or latency crosses thresholds) rather than an all-or-nothing page failure.

go deeper

for a junior

Should be able to say which 1-2 features matter most and least on an example page and that the least important ones should be turned off first under stress.

for a middle

Should know to give each section its own timeout/fallback and roughly explain why a shared resource pool is a problem.

for a senior

Should design the full priority ordering, propose bulkheads per section, and propose a proactive load-shedding trigger based on system signals, not just per-call timeouts.

for a principal

Should establish a reusable criticality/priority framework across the whole product surface, get product buy-in on the ordering, and ensure shedding policy is testable and exercised outside real incidents (e.g. via chaos/load testing).

## How feature shedding works Feature shedding under stress works by treating a composite page (or API response) as a set of independently callable sections rather than one monolithic must-all-succeed unit. - Each section — continue-watching, recommendations, trending, account menu — is fetched with its own timeout and its own fallback, typically behind a resilience wrapper (bulkhead plus timeout plus fallback, sometimes paired with a circuit breaker for the 'is this dependency currently healthy' decision). - The page-rendering layer then assembles whatever sections succeeded and quietly omits or substitutes the ones that didn't, rather than failing the whole page if any one section is slow. A **load-shedding policy** goes a step further: rather than waiting for each section to individually fail, the system proactively decides — based on current load, latency, or error-rate signals — to skip expensive or optional sections before they're even attempted, freeing capacity for the sections that matter most. ## Why prioritize at all This exists because a composite page has a much lower effective availability than any single dependency if every section is treated as required — the **weakest-link problem**. It also exists because, under real load spikes, not all features cost the same: an ML-driven recommendation call might be tens to hundreds of milliseconds of expensive scoring work per user, while an account-menu lookup is a cheap keyed read. When the system is under stress, continuing to spend capacity on the expensive, least-essential feature actively worsens the outage for the cheap, most-essential ones, so prioritized shedding directly protects the features users care about most by starving the ones they care about least. ## The trade-off The trade-off is **user-experience richness versus resilience and engineering complexity**. A page that always renders everything is a richer product on a good day, but every optional section is a dependency that can degrade the whole experience on a bad day unless it's explicitly isolated — and isolating N sections means N sets of timeouts, fallbacks, and monitoring to build and maintain, not one. There's also a product trade-off in the shedding order itself: recommendations drive engagement and revenue, so a business may resist ranking them as 'shed first' even though technically they're the most expensive and least essential to the immediate task — this decision has to be made jointly with product owners, not unilaterally by engineering. ## Failure modes in production 1. **The most common production failure is not defining a shedding order at all** — every section is fetched with the same priority and the same generous timeout, so under load, all sections degrade together and the whole page becomes slow rather than gracefully losing its least essential parts first. 2. **A second failure mode is a missing bulkhead** — if all sections share the same thread pool or connection pool, a slow recommendations call can exhaust the pool and starve the cheap continue-watching call even though they're logically independent: the fallback logic exists but never gets a chance to run because the request never even gets a thread. 3. **A third is shedding logic that's untested** and only exercised during real incidents, so the 'shed recommendations' code path itself has a bug that surfaces for the first time under the worst possible conditions, compounding the outage instead of mitigating it. ## Where it shows up - This is essentially how **Netflix's** home page and **Amazon's** product page are known to be built — independently fetched, independently degradable rows and widgets, each with its own fallback (e.g. Netflix's 'similar titles' falling back to a generic precomputed list), so a slowdown in one personalization service degrades one row instead of the entire page. - **Amazon engineers** have described designing the retail site so the buy box (price, availability, add-to-cart) is isolated from decorative elements like reviews and recommendations specifically so a reviews-service outage never threatens the ability to purchase. - **Load-shedding by priority** under stress is also formalized in patterns like Google's request-criticality levels in its internal RPC framework, where requests and features are tagged (e.g. critical, sheddable) so that during overload, infrastructure automatically rejects the lowest-priority traffic first to protect capacity for the highest-priority traffic — the same idea applied at the infrastructure layer rather than per page section.

  • Why isn't it enough to just give each section its own timeout and fallback — why do you also need a bulkhead (resource isolation)?
    Without resource isolation, a slow section can exhaust a shared thread pool, connection pool, or queue, so the request for a 'protected' section never even gets a chance to execute — the timeout and fallback code exists but is starved of resources before it can run. A bulkhead gives each section its own resource budget so one section's slowness can't starve another's capacity.
  • How would you decide, technically, when to proactively shed a feature versus waiting for it to individually time out?
    Watch a system-level signal — CPU or queue depth, overall error rate, or p99 latency — and define thresholds at which lower-priority sections are skipped outright before being attempted, rather than fetched and allowed to time out. This is more efficient than reactive per-call timeouts because it avoids spending any capacity on a call you already know is unlikely to help.
  • What would you do if product stakeholders push back on recommendations being 'shed first' because it drives revenue?
    Bring data: quantify the cost of a full page outage (which also loses the recommendation revenue, plus damages trust) versus the cost of briefly showing a generic trending list during a stress event, and propose a middle ground like a cheap non-personalized fallback for recommendations rather than removing the row outright, so the section degrades instead of disappearing.

Like a ship's crew during flooding: they seal off and abandon the less essential compartments first (a lounge) to keep power to the engine room and bridge, rather than letting water spread evenly and sink everything at once.

saying these in an interview costs you the question

  • Treats all sections of the page as equally critical with no prioritization
  • Doesn't mention resource isolation (bulkhead) alongside timeouts/fallbacks
  • Assumes shedding decisions are purely a technical call with no product input
  • No distinction between reactive (wait for failure) and proactive (shed before attempting) strategies
  • Can't name a concrete fallback for the shed feature (e.g. recommendations to trending)

context