skip to content

Overload Protection & Load Shedding

What a service does when demand exceeds capacity anyway: degrade gracefully instead of collapsing. Interviewers love cascading-failure scenarios — retry storms and priority-based shedding are the concepts they are fishing for.

on this pageshow

questions

5

Your service is over capacity and must reject some fraction of its traffic. Why is shedding a random 10% of requests worse than shedding a deliberately chosen 10%, and what does a system need in place to be able to choose?

level: seniorimportance: must knowfreq 60%

answer

  1. the amount is fixed; the selection is not
  2. label where business meaning still exists
  3. the tag must travel downstream
  4. a retry outranking a first attempt
  5. everything labelled critical ranks nothing

basics

~20 s

Random shedding drops checkout traffic and background prefetches at the same rate, so it damages revenue-bearing work to save cheap work. Choosing requires a criticality label attached at the entry point and propagated to every downstream call, with each server shedding lowest tier first.

solid answer

~50 s

Random shedding treats a user's payment and a speculative prefetch as equally valuable, so at 10% shed you break 10% of checkouts to protect work nobody would miss. Priority-based shedding fixes the *selection*, not the amount: you tag each request with a criticality at the point where its business meaning is still known — the edge, or the job that scheduled it — and then every server drops from the lowest tier upward as its own saturation signal worsens. The label has to travel with the request into downstream RPCs, or your database and your caches keep burning scarce capacity on traffic the front door already considers droppable. Google's SRE book describes four such values in its RPC system, roughly `CRITICAL_PLUS`, `CRITICAL`, `SHEDDABLE_PLUS` and `SHEDDABLE`. Two rules make it real in practice: a retry must never outrank a first attempt, and tiers need enforced quotas — the moment every team labels its own traffic critical, you are back to shedding at random.

code

python · 12 lines
python
TIERS = {"critical_plus": 0, "critical": 1, "sheddable_plus": 2, "sheddable": 3}
# as local pressure rises, keep a numerically lower (more important) tier only
SHED_ABOVE = {0.70: 3, 0.85: 2, 0.95: 1}

def admit(criticality, pressure):
    keep = 3
    for threshold, tier in sorted(SHED_ABOVE.items()):
        if pressure >= threshold:
            keep = tier - 1
    return TIERS[criticality] <= keep

print(admit("sheddable", 0.90), admit("critical_plus", 0.90))

go deeper

for a junior

Know that not all traffic is equally valuable, and that a service can be told which requests matter most so it drops the least important ones first when it is out of capacity.

for a middle

Be ready to explain where the label is assigned, how it rides along with the request into downstream calls, and why a server drops the lowest tier first against its own saturation signal rather than a fixed request rate.

for a senior

Show the operational depth: retries demoted below first attempts, criticality carried in RPC metadata next to the deadline, and shedding decided per server from a measured signal because request cost varies across endpoints and deploys.

for a principal

Own the governance. Decide how many tiers exist, what the default is, how per-tier quotas are enforced and audited across teams, and how you prevent the inflation that quietly turns priority shedding back into random shedding.

## The selection problem, not the volume problem When a service is over capacity, the amount it must reject is fixed by arithmetic — offered load minus capacity. The only free variable is *which* requests get rejected, and that variable is worth a great deal. A random 10% shed is a 10% failure rate for payments, a 10% failure rate for login, and a 10% failure rate for a background image prefetch. Nobody would choose that distribution if asked. Priority-based load shedding is simply the mechanism that lets the system make the choice a human would make, at machine speed, without anyone being paged first. ## What criticality is, and where it comes from Criticality is a property of the *request*, assigned where its business meaning is still visible. That is almost never inside the service being protected — a storage node cannot tell from an RPC whether the row it is fetching backs a checkout page or a nightly report. So the label is stamped at the origin: the API gateway or front-end handler that knows this is an interactive checkout; the scheduler that knows this is a batch backfill; the cache-warming job that knows its work is speculative. Google's SRE book (2016) documents four values used in its internal RPC system as an example: `CRITICAL_PLUS` for the most important user-facing traffic where failure is directly visible, `CRITICAL` as the default for user-facing requests, `SHEDDABLE_PLUS` for work with partial-unavailability tolerance such as batch jobs that can be retried later, and `SHEDDABLE` for work that is expected to fail routinely. The specific names are not an industry standard; what matters in an interview is the shape — a small number of ordered tiers, few enough that people can reason about them, and coarse enough that they map onto real product behaviour. ## Propagation is the part people forget A criticality label that stops at the edge protects only the edge. A single user-facing request typically fans out to a dozen internal services, and the ones nearest to a shared resource — the database, the cache tier, an authentication service — are the ones most likely to be the actual bottleneck. If the label does not travel in RPC metadata alongside the trace context and the deadline, those services shed uniformly and the shared bottleneck keeps spending its scarce capacity on sheddable work. Propagation also means the label must be *inherited*: any call a service makes on behalf of a request carries that request's criticality, and it may lower it but should never raise it. ## Retries and criticality A retry is the single most dangerous thing to mis-tier. If retried requests keep or gain priority, then a service under stress fills up with retries that displace fresh first attempts, and the failure sustains itself. The rule is that a retry is never more important than the original — many designs deliberately mark retried requests one tier lower, so that under pressure the system serves new work and sheds the amplification. ## Shedding against a local signal Each server decides independently, using its own saturation measure — queue wait, in-flight count, CPU pressure — and progressively raises the tier below which it refuses. At mild pressure it drops `SHEDDABLE`; as pressure worsens it drops `SHEDDABLE_PLUS`, and only under extreme stress does it touch user-facing traffic. This has to be against a measured local signal rather than a static request-per-second cap, because the cost of a request varies by an order of magnitude across endpoints and changes with every deploy. ```python # shed the numerically lowest-importance tier first as pressure rises if pressure >= 0.95: keep_tier = 0 # only the most critical elif pressure >= 0.85: keep_tier = 1 elif pressure >= 0.70: keep_tier = 2 else: keep_tier = 3 ``` ## The failure mode: tier inflation The predictable organisational failure is that every team labels its own traffic with the highest tier, because from inside any one team the traffic genuinely is important. Once that happens the tiers rank nothing and shedding is random again with extra machinery. The defences are enforced quotas — a service is allowed only so much of its traffic at each tier, verified from telemetry rather than from intent — plus a periodic audit of tier distribution, and a default that is *not* the top tier. Treat the criticality distribution as a metric you graph, not as a configuration you set once. ## The interview point The decision with a cost is what you are willing to sacrifice first, and having decided it *before* the incident. The strong answer names the label, insists it is stamped at the origin and propagated end to end, handles retries explicitly, and then admits the governance problem — because the mechanism is easy and keeping the tiers honest is the part that actually fails.

  • How do you stop every team from labelling its own traffic as the highest criticality?
    Do not rely on intent. Give each service a quota of traffic per tier, measure the actual distribution as a graphed metric, and make the default tier something below the top so claiming priority is an explicit act. Review the distribution periodically the way you would review any budget, and treat a service whose traffic is 100% top-tier as a finding, not as a fact.
  • Where should the criticality of an automatically retried request sit relative to the original?
    No higher, and often one tier lower. Retries are exactly the traffic that appears when a service is weakest, so letting them keep full priority means a struggling service fills with amplification and starves fresh first attempts. Demoting retries makes the system shed the duplicated work first and keeps new user requests moving.
  • Why not let each service infer criticality from the endpoint being called instead of propagating a label?
    Because the same endpoint serves different callers. A single row fetch may back an interactive checkout or a nightly report, and the storage node cannot tell them apart. Endpoint-based inference also breaks the moment a new caller appears, whereas a propagated label stays correct because it was assigned where the business meaning was known.

saying these in an interview costs you the question

  • Assigns criticality inside the service being protected
  • Lets the label stop at the edge instead of propagating it
  • Gives retried requests the same or higher priority
  • Defines a dozen tiers nobody can reason about
  • Treats tier labels as trustworthy without quotas or audits

context

open as a page

Requests hitting a backend service triple within a minute while its success rate falls, and after you restart it, it saturates again within seconds. How do you confirm this is a retry storm rather than a genuine traffic surge, and what do you do to break it?

level: seniorimportance: must knowfreq 68%

basics

~20 s

Retry storms show load rising as success falls, with internal call volume spiking while edge traffic stays flat — organic surges raise both together. Breaking one requires cutting offered load below the degraded service's reduced capacity, then restoring in steps.

open as a page

To make sure no request is ever rejected, a team makes their service's own inbound request queue unbounded. During the next traffic spike, what actually happens to latency, memory and success rate, and what queue design would you use instead?

level: middleimportance: should knowfreq 45%

basics

~20 s

An unbounded inbound queue converts an overload into unbounded latency plus a memory incident: the queue grows without limit, every request waits longer than its client's timeout, and the process eventually dies. A bounded, deadline-aware queue rejects immediately instead.

open as a page

A service sized for about 10,000 requests per second is suddenly offered 30,000. Instead of serving roughly 10,000 of them successfully and failing the rest, nearly every request now times out. Explain what is happening inside the server, and what changes if it sheds the excess load instead.

level: middleimportance: should knowfreq 58%

basics

~20 s

Accepting more work than it can finish makes a server spend its capacity on requests whose callers have already timed out, so goodput — useful completed work — collapses toward zero. Shedding excess load early and cheaply keeps the fraction it does serve healthy.

open as a page

One of your three availability zones fails and the surviving capacity can serve only about 70% of peak traffic. Would you let all users experience a slow, partly broken service, or deliberately serve 70% of them fully and reject the rest? How would you decide, and who has to agree to that beforehand?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Deliberate shedding usually wins: partially served users retry and consume more capacity than rejected ones, so uniform degradation costs more than it saves. Shed by session rather than per request, rank traffic by pre-agreed business criticality, and get that ladder signed off before the incident.

open as a page