skip to content

questions

6

In a cloud service, what is throttling and why would a service deliberately reject or delay some incoming requests instead of trying to handle all of them?

level: juniorimportance: must knowfreq 70%

answer

  1. measure vs budget vs policy
  2. knee of the latency curve
  3. retry storm / metastable failure
  4. reject vs queue vs degrade
  5. per-tenant fair share

basics

~20 s

Throttling means a service limits how much work it accepts. If too many requests come in at once, it slows down, delays, or rejects some so it doesn't crash and can keep serving everyone reasonably well.

solid answer

~40 s

Throttling caps the rate of consumption a service accepts from its callers so that resource usage stays within safe operating limits. Instead of accepting unbounded load and degrading unpredictably (or crashing) under a spike, the service measures incoming demand against known capacity and, once a threshold is crossed, applies a policy: reject excess requests, queue them for later, or serve a cheaper/degraded response. The goal isn't to punish callers — it's to protect the service's SLOs (latency, availability) for everyone, including callers that are behaving reasonably, by shedding or postponing the load that would otherwise push the system past its stable operating point.

go deeper

for a junior

Should be able to explain in plain terms that throttling caps how much work a service accepts and name at least one thing it does with excess (reject/delay). Doesn't need to discuss SLOs or fairness by name.

for a middle

Should name all three excess-load strategies (reject/queue/degrade) and explain why unmanaged overload is worse than a controlled rejection, referencing latency or availability impact.

for a senior

Should discuss where in the request path throttling belongs, connect it explicitly to protecting SLOs, and raise per-tenant fairness unprompted as a concern in shared systems.

for a principal

Should reason about system-wide failure dynamics (retry storms, metastable failure), how throttling interacts with client retry behavior and circuit breakers, and how to set policy across an organization with many services and shared dependencies.

## What throttling actually is Throttling is an **admission-control pattern**: before a service commits real resources to handling a request, it checks whether accepting that request would push its current load past a **capacity budget** it knows it can sustain, and if so, it applies a policy other than 'just do the work.' ## The mechanism The mechanism has three parts. 1. **First**, the service (or a component in front of it, like a gateway or sidecar) continuously **measures consumption** — requests per second, concurrent connections, CPU or memory pressure, database connections in use, or an abstract 'cost' unit that weights expensive operations more heavily than cheap ones. 2. **Second**, it compares that measurement against a **budget derived from known capacity**: how many requests the downstream database, thread pool, or dependency can absorb while still meeting its latency targets. 3. **Third**, once the measured consumption crosses the budget (or a safety margin below it), the service applies one of a small set of responses to the excess: **reject it outright** and fail fast, **queue it briefly** so it can be served once capacity frees up, or **degrade it** — serve a cheaper, partial, cached, or lower-fidelity version of the response instead of doing the full expensive work. ## Why the pattern exists The reason this pattern exists is that unmanaged demand growth doesn't degrade a service gracefully — it degrades it catastrophically. Most services have a stable operating region where latency stays roughly flat as load increases, followed by a **knee** where finite resources (thread pools, connection pools, CPU, downstream databases) start to saturate and latency climbs steeply, often non-linearly. Past that knee, clients start timing out and retrying, which adds more load on top of an already-struggling system — a feedback loop sometimes called a **retry storm** or, in the worst case, a **metastable failure**, where the system never recovers on its own even after the original traffic spike subsides, because retry traffic alone is enough to keep it pinned past capacity. Throttling exists to keep the system on the flat, stable side of that knee by refusing to let accepted work exceed what the system can sustainably process, protecting the SLOs (latency and availability targets) that matter to the business. ## What to do with the excess The trade-offs differ by which response the throttle chooses for excess load. - **Rejecting** is the cheapest and most predictable option operationally — the service does almost no work for a rejected request, so it costs it almost nothing — but it pushes the burden onto the caller, who now has to handle a failure (retry with backoff, show an error, fall back to a cached result). - **Queueing** hides the failure from the caller by adding latency instead of an error, which is a better user experience for short delays, but it only works if the queue is bounded and the arrival rate is expected to drop back below the drain rate soon; an unbounded or long-lived queue just delays the same overload problem and adds memory pressure and unpredictable tail latency. - **Degrading** — serving stale cached data, a simplified response, or skipping a non-critical enrichment step — preserves availability and keeps latency flat, but it costs correctness or completeness, and only works for operations where a slightly worse answer is acceptable (a checkout total can't be 'approximately' right, but a product recommendation list can be). ## Fairness across callers A dimension often missed is fairness across callers. A single global budget protects the service as a whole but does nothing to stop one very active caller — a tenant with an unusually large customer base, or a client with a bug causing a retry loop — from consuming the entire budget and starving every other caller, a classic **'noisy neighbor'** problem in multi-tenant systems. Robust throttling therefore usually layers a **per-tenant or per-API-key budget** on top of (or instead of) a single global one, so that no one caller's excess demand degrades service for everyone else, at the cost of extra bookkeeping (tracking consumption per key rather than in aggregate) and of having to decide what a 'fair share' is: - equal allocation - usage-based allocation - or a paid tier's larger allocation. ## Failure modes in production In production, failure modes show up in predictable ways: - a threshold set **too conservatively** rejects perfectly legitimate traffic during an ordinary daily peak, which looks to the business like an outage even though the system never actually overloaded - a threshold set **too loosely** fails to trigger before the system tips past its knee, so the throttle never actually protects anything - and a throttle applied **without per-tenant fairness** lets one customer's traffic spike take down service for every other customer sharing the pool. ## Where it shows up A well-known real-world instance is a payments platform like Stripe, which applies per-account request quotas so that a burst of traffic from one large merchant's flash sale can't degrade API latency for every other merchant on the shared platform — the throttle protects the platform's SLOs for the tenants who are behaving normally, not just the platform's aggregate uptime number.

  • Where in the request path should the throttling decision be made, and why does that placement matter?
    As early as possible — ideally at the edge or gateway, before the request has consumed expensive resources like a database connection or a downstream call. Throttling after the request has already done most of its work saves nothing, because the resource has already been spent by the time the request is rejected. Early placement also means a single overloaded caller can be shed before it ever reaches the services it would otherwise harm.
  • How does throttling differ from a circuit breaker?
    A circuit breaker reacts to a dependency's own health — it trips when calls to that dependency are failing or timing out, protecting the caller from a dependency that is already unhealthy. Throttling instead controls how much load the throttling service itself accepts from its callers, independent of whether any dependency has failed yet. They're complementary: throttling caps demand before it becomes a problem, while a circuit breaker responds after a dependency is already struggling.
  • Can throttling alone prevent a metastable failure?
    Not by itself — it needs to be paired with well-behaved clients that back off on rejection rather than retrying immediately, otherwise a rejected burst just turns into an even bigger retry burst. Throttling controls what the service accepts, but the client's retry policy controls what gets offered back, and both sides have to cooperate for the system to actually recover.

Like a nightclub bouncer capping how many people are let in at once: past capacity, the bouncer doesn't let the room turn into a crush — new arrivals wait in line (queue), get turned away (reject), or are let into a side room with limited music (degrade), so everyone already inside still has a good time.

saying these in an interview costs you the question

  • Says throttling and rate limiting are only about stopping abuse/DoS, missing SLO protection under normal legitimate load
  • Can't name any option besides 'return an error' for handling excess load
  • Assumes one global threshold is enough and doesn't mention the noisy-neighbor / per-tenant problem
  • Thinks throttling happens after the expensive work is done rather than at admission time
  • Doesn't recognize that rejecting traffic without caller backoff can make things worse

context

open as a page

In a throttling policy, what's the practical difference between a hard limit and a soft limit, and when would a team choose each?

level: middleimportance: must knowfreq 60%

basics

~20 s

A hard limit is a strict cap that's never crossed, even if there's spare capacity. A soft limit is a flexible cap that can be exceeded temporarily if the system has room, but tightens up when things get busy.

open as a page

When a service decides it can't safely accept a request within its throttling policy, its three main options are to reject it outright, queue it briefly, or serve a degraded response. Walk through the trade-offs of each and how you'd decide which to apply to a given request type.

level: seniorimportance: must knowfreq 65%

basics

~20 s

You can say no right away, make the request wait a bit, or give a simpler/cheaper answer instead of the full one. Saying no is cheap but annoys the caller; waiting is nicer but risky if too many wait at once; a simpler answer keeps things running but might be less accurate.

open as a page

In a multi-tenant platform where all tenants share one backend capacity pool, what techniques let you throttle for per-tenant fairness so a single heavy tenant can't degrade service for everyone else?

level: middleimportance: should knowfreq 55%

basics

~10 s

Instead of one big shared limit for everybody, give each customer their own smaller limit. That way one customer sending a lot of traffic can only use up their own share, not everyone else's.

open as a page

How should a team decide where to set a throttling threshold relative to a service's SLOs, and what goes wrong in production if that threshold is set too aggressively versus too loosely?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Set the limit based on real measured capacity, with some safety margin, not a guess. Too strict and you block normal traffic for no reason; too loose and the throttle never actually kicks in before things break.

open as a page

In a large distributed system with an edge/gateway layer, individual backend services, and per-tenant concerns all in play, how would you design a throttling policy across these layers, and what goes wrong if throttling is only applied at one of them?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Put limits at more than one point — at the front door, inside each service, and per customer — instead of just one place. If you only limit at one spot, problems from a different spot slip through and can still overload something downstream.

open as a page