In a cloud service, what is throttling and why would a service deliberately reject or delay some incoming requests instead of trying to handle all of them?
answer
- measure vs budget vs policy
- knee of the latency curve
- retry storm / metastable failure
- reject vs queue vs degrade
- per-tenant fair share
basics
~20 sThrottling means a service limits how much work it accepts. If too many requests come in at once, it slows down, delays, or rejects some so it doesn't crash and can keep serving everyone reasonably well.
solid answer
~40 sThrottling caps the rate of consumption a service accepts from its callers so that resource usage stays within safe operating limits. Instead of accepting unbounded load and degrading unpredictably (or crashing) under a spike, the service measures incoming demand against known capacity and, once a threshold is crossed, applies a policy: reject excess requests, queue them for later, or serve a cheaper/degraded response. The goal isn't to punish callers — it's to protect the service's SLOs (latency, availability) for everyone, including callers that are behaving reasonably, by shedding or postponing the load that would otherwise push the system past its stable operating point.
go deeper
Should be able to explain in plain terms that throttling caps how much work a service accepts and name at least one thing it does with excess (reject/delay). Doesn't need to discuss SLOs or fairness by name.
Should name all three excess-load strategies (reject/queue/degrade) and explain why unmanaged overload is worse than a controlled rejection, referencing latency or availability impact.
Should discuss where in the request path throttling belongs, connect it explicitly to protecting SLOs, and raise per-tenant fairness unprompted as a concern in shared systems.
Should reason about system-wide failure dynamics (retry storms, metastable failure), how throttling interacts with client retry behavior and circuit breakers, and how to set policy across an organization with many services and shared dependencies.
## What throttling actually is Throttling is an **admission-control pattern**: before a service commits real resources to handling a request, it checks whether accepting that request would push its current load past a **capacity budget** it knows it can sustain, and if so, it applies a policy other than 'just do the work.' ## The mechanism The mechanism has three parts. 1. **First**, the service (or a component in front of it, like a gateway or sidecar) continuously **measures consumption** — requests per second, concurrent connections, CPU or memory pressure, database connections in use, or an abstract 'cost' unit that weights expensive operations more heavily than cheap ones. 2. **Second**, it compares that measurement against a **budget derived from known capacity**: how many requests the downstream database, thread pool, or dependency can absorb while still meeting its latency targets. 3. **Third**, once the measured consumption crosses the budget (or a safety margin below it), the service applies one of a small set of responses to the excess: **reject it outright** and fail fast, **queue it briefly** so it can be served once capacity frees up, or **degrade it** — serve a cheaper, partial, cached, or lower-fidelity version of the response instead of doing the full expensive work. ## Why the pattern exists The reason this pattern exists is that unmanaged demand growth doesn't degrade a service gracefully — it degrades it catastrophically. Most services have a stable operating region where latency stays roughly flat as load increases, followed by a **knee** where finite resources (thread pools, connection pools, CPU, downstream databases) start to saturate and latency climbs steeply, often non-linearly. Past that knee, clients start timing out and retrying, which adds more load on top of an already-struggling system — a feedback loop sometimes called a **retry storm** or, in the worst case, a **metastable failure**, where the system never recovers on its own even after the original traffic spike subsides, because retry traffic alone is enough to keep it pinned past capacity. Throttling exists to keep the system on the flat, stable side of that knee by refusing to let accepted work exceed what the system can sustainably process, protecting the SLOs (latency and availability targets) that matter to the business. ## What to do with the excess The trade-offs differ by which response the throttle chooses for excess load. - **Rejecting** is the cheapest and most predictable option operationally — the service does almost no work for a rejected request, so it costs it almost nothing — but it pushes the burden onto the caller, who now has to handle a failure (retry with backoff, show an error, fall back to a cached result). - **Queueing** hides the failure from the caller by adding latency instead of an error, which is a better user experience for short delays, but it only works if the queue is bounded and the arrival rate is expected to drop back below the drain rate soon; an unbounded or long-lived queue just delays the same overload problem and adds memory pressure and unpredictable tail latency. - **Degrading** — serving stale cached data, a simplified response, or skipping a non-critical enrichment step — preserves availability and keeps latency flat, but it costs correctness or completeness, and only works for operations where a slightly worse answer is acceptable (a checkout total can't be 'approximately' right, but a product recommendation list can be). ## Fairness across callers A dimension often missed is fairness across callers. A single global budget protects the service as a whole but does nothing to stop one very active caller — a tenant with an unusually large customer base, or a client with a bug causing a retry loop — from consuming the entire budget and starving every other caller, a classic **'noisy neighbor'** problem in multi-tenant systems. Robust throttling therefore usually layers a **per-tenant or per-API-key budget** on top of (or instead of) a single global one, so that no one caller's excess demand degrades service for everyone else, at the cost of extra bookkeeping (tracking consumption per key rather than in aggregate) and of having to decide what a 'fair share' is: - equal allocation - usage-based allocation - or a paid tier's larger allocation. ## Failure modes in production In production, failure modes show up in predictable ways: - a threshold set **too conservatively** rejects perfectly legitimate traffic during an ordinary daily peak, which looks to the business like an outage even though the system never actually overloaded - a threshold set **too loosely** fails to trigger before the system tips past its knee, so the throttle never actually protects anything - and a throttle applied **without per-tenant fairness** lets one customer's traffic spike take down service for every other customer sharing the pool. ## Where it shows up A well-known real-world instance is a payments platform like Stripe, which applies per-account request quotas so that a burst of traffic from one large merchant's flash sale can't degrade API latency for every other merchant on the shared platform — the throttle protects the platform's SLOs for the tenants who are behaving normally, not just the platform's aggregate uptime number.
- Where in the request path should the throttling decision be made, and why does that placement matter?As early as possible — ideally at the edge or gateway, before the request has consumed expensive resources like a database connection or a downstream call. Throttling after the request has already done most of its work saves nothing, because the resource has already been spent by the time the request is rejected. Early placement also means a single overloaded caller can be shed before it ever reaches the services it would otherwise harm.
- How does throttling differ from a circuit breaker?A circuit breaker reacts to a dependency's own health — it trips when calls to that dependency are failing or timing out, protecting the caller from a dependency that is already unhealthy. Throttling instead controls how much load the throttling service itself accepts from its callers, independent of whether any dependency has failed yet. They're complementary: throttling caps demand before it becomes a problem, while a circuit breaker responds after a dependency is already struggling.
- Can throttling alone prevent a metastable failure?Not by itself — it needs to be paired with well-behaved clients that back off on rejection rather than retrying immediately, otherwise a rejected burst just turns into an even bigger retry burst. Throttling controls what the service accepts, but the client's retry policy controls what gets offered back, and both sides have to cooperate for the system to actually recover.
Like a nightclub bouncer capping how many people are let in at once: past capacity, the bouncer doesn't let the room turn into a crush — new arrivals wait in line (queue), get turned away (reject), or are let into a side room with limited music (degrade), so everyone already inside still has a good time.
saying these in an interview costs you the question
- Says throttling and rate limiting are only about stopping abuse/DoS, missing SLO protection under normal legitimate load
- Can't name any option besides 'return an error' for handling excess load
- Assumes one global threshold is enough and doesn't mention the noisy-neighbor / per-tenant problem
- Thinks throttling happens after the expensive work is done rather than at admission time
- Doesn't recognize that rejecting traffic without caller backoff can make things worse