One of your three availability zones fails and the surviving capacity can serve only about 70% of peak traffic. Would you let all users experience a slow, partly broken service, or deliberately serve 70% of them fully and reject the rest? How would you decide, and who has to agree to that beforehand?
answer
- everyone at 70% or 70% at full
- half-served users come back
- journeys complete, requests do not
- hash the session, not the request
- a lever nobody may pull is not a lever
basics
~20 sDeliberate shedding usually wins: partially served users retry and consume more capacity than rejected ones, so uniform degradation costs more than it saves. Shed by session rather than per request, rank traffic by pre-agreed business criticality, and get that ladder signed off before the incident.
solid answer
~50 sSpreading a 30% shortfall evenly means every user gets a page where some parts fail, and users respond to broken pages by reloading them — so uniform brownout manufactures extra load precisely when you have none to spare, and no user journey completes. Deterministic shedding gives a defined majority a fully working product while the remainder gets a clear, honest failure. The critical implementation detail is the unit: shed by user or session hash, not per request, or you break a checkout at step three after the card was charged and that user retries the whole flow. Ranking comes from a degradation ladder agreed with product, support and legal in advance — which flows are protected to the end, which features go dark first, whether paying and free tiers are treated differently. And the ladder is only a control if on-call is pre-authorised to pull it at 3am without waiting for an executive, and has pulled it before in a drill.
go deeper
Know that when capacity is short, a service can either give everyone a worse experience or give some people the full experience, and that this is a deliberate choice rather than an accident.
Be ready to explain why partially served users generate extra load by retrying, and why shedding on a stable session key keeps multi-step flows completing while per-request shedding does not.
Show the operational reasoning: journey completion rate as the metric that matters, protecting in-flight transactions, keeping telemetry legible by making the split clean, and knowing the case where degrading everyone genuinely wins.
Own the policy and its authorisation. Decide the degradation ladder with product, support and legal ahead of time, resolve the fairness question about who stays excluded, pre-authorise on-call to pull it, and require it to be rehearsed in production.
## The choice, stated honestly A capacity shortfall forces a distribution decision that most teams have never made explicitly: with 70% of the capacity and 100% of the demand, do you give everybody 70% of a service, or 70% of people 100% of a service? The instinct — and the default behaviour of a system with no shedding at all — is the first. It feels fairer, and it avoids the uncomfortable act of turning users away. It is usually the worse choice, and being able to say why is the point of the question. ## Why uniform degradation costs more than it saves **Incomplete work is repeated work.** A user whose page half-loaded reloads it. A user whose checkout failed at the payment step tries again, often from the start. A partially served request has consumed capacity and produced nothing durable, and it usually comes back. A cleanly rejected user with an honest error message is far more likely to wait or leave. So uniform degradation generates additional offered load at the exact moment there is none available — the same self-amplifying dynamic that turns overload into collapse. **Nothing completes end to end.** Product value lives in journeys, not requests. If every step of a five-step flow has a 30% chance of failing, roughly 17% of journeys complete. Rejecting 30% of *users* and serving the rest cleanly completes about 70% of journeys. The same capacity, four times the delivered outcome. **Diagnosis becomes impossible.** Universal partial failure means every dashboard is amber and every service looks slightly ill, so the on-call responder cannot tell which component is the constraint. A clean split — a defined share served correctly, the rest rejected at the edge with a known reason — leaves the telemetry legible. The honest counter-case exists and you should name it: where partial results are genuinely useful and stateless — a search returning fewer results, a feed with fewer items, video at a lower bitrate — degrading everyone can beat excluding anyone. The rule of thumb is that degradation wins when a partial answer is still a usable answer, and shedding wins when the product is a multi-step transaction. ## Choosing the unit of shedding This is where designs quietly fail. Shedding per request produces the worst of both worlds: each user experiences a random 30% of their requests failing, which is uniform degradation wearing a load-shedding costume. Shed on a stable key — a hash of the user or session identifier — so that a given user is either in or out for the duration, and the ones who are in complete their journeys. That immediately raises fairness, which is a policy question rather than a technical one. Consistent hashing means the same unlucky users are excluded for the entire incident; rotating the excluded set spreads the pain but breaks journeys mid-flight. Common resolutions: rotate slowly, at a period longer than a typical session; never shed a session that is already inside a transactional flow; and always let already-authenticated, in-progress journeys drain. ## Where the ranking comes from The ordering is a business decision, not an engineering one, and it must be written down before it is needed. A workable degradation ladder names, in order: the flows protected to the last drop of capacity (usually payment, safety, regulatory, and anything that would corrupt state if interrupted); the features turned off first (recommendations, personalisation, analytics beacons, non-essential enrichment); how customer tiers rank against each other, informed by contractual commitments; and how internal traffic — batch jobs, backfills, cache warming — is stopped outright before any user traffic is touched. The stakeholders who must sign it are product (which features may go dark), support (what users are told, and what the status page says), and legal or commercial (where contractual commitments constrain who may be shed). Engineering owns the mechanism; it does not own the ranking, and a plan engineering wrote alone will be renegotiated live during the incident, which is the worst possible time. ## What makes it a real control Three things, all of which are usually missing: **Pre-authorisation.** On-call must be able to invoke the ladder without an approval chain. A lever that requires a vice-president at 3am is not a lever. **Rehearsal.** A shed path exercised for the first time during an incident will not work — the flag is stale, the config path is broken, the fallback page 500s. It needs to be pulled deliberately, in production, on a schedule. **Observability of the shed.** You must be able to see how much is being shed and who is being shed, or you will not know when to restore, and you will not notice when the mechanism silently stops shedding. ## The interview point The candidate to hire is the one who reaches for the deterministic split, justifies it with the retry dynamics rather than a slogan about fairness, picks the session as the unit, and then spends most of the answer on the part that is not technical at all: who agreed to this ranking, when, and whether anyone has ever pulled the lever in anger. Sizing headroom so the zone loss is absorbed in the first place is the preventive discipline, but this question assumes it was insufficient — which, on the day of a real regional event, it frequently is.
- When would you actually prefer to degrade every user rather than shed a fraction of them?When a partial answer is still genuinely useful and the interaction is not transactional: search returning fewer results, a feed with fewer items, video stepped down to a lower bitrate, a map with less detail. Everyone gets a usable if lesser product and nobody is locked out. The moment the flow is a multi-step transaction that can fail halfway and leave state behind, that logic reverses.
- If you shed by hashing the user identifier, the same users are excluded for the whole incident. Is that acceptable?It is a policy call, and it needs an explicit answer before the incident. Stability protects journeys but concentrates the harm; rotation spreads it but breaks flows in flight. The usual compromise is to rotate on a period longer than a typical session, never shed a session already inside a transaction, and let in-progress journeys drain rather than cutting them mid-way.
- What single piece of evidence would convince you a documented degradation ladder is real rather than aspirational?A timestamped record of the last time it was deliberately exercised in production — a drill where on-call pulled it without an approval chain and the shed traffic showed up on a dashboard. Absent that, assume the flag is stale, the fallback path is broken, and the first attempt will happen under maximum pressure with an audience.
saying these in an interview costs you the question
- Assumes degrading everyone is automatically the fairer choice
- Sheds per request, breaking journeys midway
- Leaves the ranking of flows to be decided during the incident
- Requires executive approval before on-call may shed
- Counts a never-rehearsed kill switch as a working control