skip to content

When a service decides it can't safely accept a request within its throttling policy, its three main options are to reject it outright, queue it briefly, or serve a degraded response. Walk through the trade-offs of each and how you'd decide which to apply to a given request type.

level: seniorimportance: must knowfreq 65%

answer

  1. reject = cheap, pushes cost to caller
  2. queue = latency not error, must be bounded
  3. degrade = cheaper answer, costs correctness
  4. priority-based shedding by request criticality
  5. queueing here ≠ durable message-queue buffering

basics

~20 s

You can say no right away, make the request wait a bit, or give a simpler/cheaper answer instead of the full one. Saying no is cheap but annoys the caller; waiting is nicer but risky if too many wait at once; a simpler answer keeps things running but might be less accurate.

solid answer

~50 s

Rejecting fails fast and costs the service almost nothing, pushing the burden of retrying onto the caller — best for requests where a quick, clear failure lets the caller recover cleanly (idempotent calls, non-critical enrichments). Queueing hides the overload as added latency instead of an error, which is a better experience for short delays, but only works with a small, bounded queue and an arrival rate expected to drop soon — an unbounded queue just delays the same overload and adds unpredictable tail latency. Degrading serves a cheaper, partial, cached, or lower-fidelity response instead of doing full work, preserving availability and flat latency at the cost of correctness or completeness, which only fits requests where an approximate answer is actually acceptable. The choice depends on the request's criticality, whether a slightly-wrong answer is tolerable, and whether the caller can handle an error gracefully.

go deeper

for a junior

Should be able to name at least two of the three options (reject, queue, degrade) and give one reason to use each.

for a middle

Should describe all three with a basic trade-off for each and give at least one example request type suited to each.

for a senior

Should reason explicitly about bounding the queue, the risk of an unbounded queue turning into slow failures, and pick strategies based on request criticality/idempotency rather than one-size-fits-all.

for a principal

Should discuss priority-based shedding across an entire request mix during a single overload event, and be able to draw the boundary between admission-time queueing and durable buffering patterns without conflating them.

## Three answers to one decision Once a throttling decision determines that a request exceeds the currently sustainable capacity, the service needs a concrete answer for what to actually do with that specific request, and there are three broad options, each appropriate for a different kind of request and a different kind of caller. | Option | What the service does with the request | What it costs the caller | |---|---|---| | **Rejecting** | returning an explicit failure immediately | the cost of rejection is entirely pushed onto the caller | | **Queueing** | holding the request briefly at the admission boundary | added latency instead of an outright error | | **Degrading** | serving a deliberately cheaper version of the response | costs correctness, freshness, or completeness | ## Rejecting **Rejecting** means returning an explicit failure immediately, without doing further work. Mechanically this is the cheapest option for the service — it spends almost no CPU, memory, or downstream capacity handling a rejected request, which is exactly the point: the whole reason to reject is to avoid spending resources on work the system can't currently sustain. The cost of rejection is entirely pushed onto the caller, who now has to detect the failure and decide what to do — retry later, show the end user an error, or fall back to some other path. Rejection is the right default: - for requests that are safe to retry (idempotent operations) - for callers sophisticated enough to back off sensibly instead of hammering the service again immediately - and for any request where a quick, unambiguous failure is preferable to an uncertain wait — a client waiting on a slow response often has a worse experience than one that gets an immediate, clear error it can act on. ## Queueing **Queueing** means holding the request briefly at the admission boundary rather than immediately answering it, and serving it once capacity frees up, so the caller experiences the overload as added latency instead of an outright error. This is a genuinely different user experience — many callers would rather wait two extra seconds than see an error — but it only works under two conditions: - the queue has to be small and bounded - the arrival rate has to be expected to drop back below the drain rate soon This is deliberately a short-lived, admission-time holding mechanism, distinct from the durable, application-level buffering used to smooth bursts over a longer window (a separate pattern built around a persistent message queue between producer and consumer) — the throttling-level queue exists only to smooth momentary admission spikes, not to durably absorb sustained overload. If the queue isn't bounded, or the overload doesn't actually subside, queueing doesn't prevent overload, it just delays and disguises it — held requests keep accumulating, memory pressure grows, and by the time requests are finally served, their latency may already have breached the caller's own timeout, so the caller experiences a slow failure instead of a fast one, which is often worse than an immediate rejection would have been. ## Degrading **Degrading** means serving a deliberately cheaper version of the response instead of doing the full work — returning cached or slightly stale data instead of a freshly computed answer, skipping a non-critical enrichment step, or serving a simplified result. This preserves both availability and flat latency, since the degraded path is by design cheap to serve, but it costs correctness, freshness, or completeness, so it's only appropriate where an approximate or slightly-out-of-date answer doesn't cause real harm to the caller. A product recommendation list can safely fall back to a generic, cached list under load; a payment authorization amount cannot be 'approximately right.' Deciding which requests are eligible for degradation is a product-level judgment as much as a technical one, and it typically requires having actually built the cheaper fallback path in advance, which is engineering effort spent before the incident, not during it. ## Choosing for a given request type In practice, the choice among the three tracks two questions about the specific request type: 1. how bad is a wrong-but-plausible answer compared to no answer at all 2. and how well can the caller handle each kind of response? - **A background batch job** calling an internal API can usually tolerate rejection with backoff just fine. - **A user-facing checkout total** must be either fully correct or explicitly failed — never approximated — making it a strong candidate for rejection (or, if slightly delayed completion is safe, a short bounded queue) but a poor candidate for degradation. - **A homepage's 'related products' widget** is the opposite: correctness barely matters, so serving a stale cached list under load (degrade) is clearly better than making the user wait or showing an error for a non-critical element. ## Priority-based shedding Systems in production often combine strategies by priority: under load, low-priority, degradable requests (recommendations, non-critical background sync) get degraded or rejected first, freeing capacity to keep serving the high-priority path (checkout, authentication) at full fidelity — a form of load shedding that treats not all requests as equally worth protecting.

  • Why is an unbounded queue at the throttling boundary considered a failure mode rather than a solution?
    An unbounded queue doesn't reduce demand, it just delays it, so if the arrival rate stays above the drain rate the backlog keeps growing indefinitely, consuming memory and eventually breaching callers' own timeouts anyway. At that point the caller experiences a slow, wasted failure instead of a fast one, which is usually a worse outcome than an immediate rejection would have been.
  • How would you decide whether a given request type is a good candidate for degradation versus rejection?
    Ask whether an approximate, stale, or partial answer is actually acceptable to whatever consumes the response — if yes, and a cheap fallback can be pre-built, degradation preserves a better experience than an error. If the request needs to be exactly correct or not answered at all, like a financial calculation, degradation isn't a safe option and rejection (possibly with a short bounded queue) is the right choice instead.
  • What role does priority play in choosing among reject, queue, and degrade under a single overload event?
    Not every request competing for the same capacity is equally important, so systems often shed or degrade lower-priority requests first — background jobs, non-critical enrichments — specifically to preserve full-fidelity capacity for higher-priority ones like checkout or auth. This turns a blunt, uniform throttle into a form of triage that protects what matters most during a spike.

Like a restaurant at capacity: turning walk-ins away at the door (reject) costs the restaurant nothing extra; a short waitlist with a real wait time (queue) works only if tables free up soon; offering a simplified express menu to keep the kitchen moving (degrade) keeps everyone fed, just not with their first choice.

saying these in an interview costs you the question

  • Treats 'reject' as the only real option, missing queue and degrade entirely
  • Recommends unbounded queueing as a safe way to never lose a request
  • Suggests degrading a financially or safety-critical response (e.g., a payment amount) under load
  • Doesn't connect the choice of strategy to whether the request is idempotent/retry-safe
  • Confuses the short admission-time queue with a durable, persistent message-queue buffering system

context