skip to content

A service sized for about 10,000 requests per second is suddenly offered 30,000. Instead of serving roughly 10,000 of them successfully and failing the rest, nearly every request now times out. Explain what is happening inside the server, and what changes if it sheds the excess load instead.

level: middleimportance: should knowfreq 58%

answer

  1. throughput is not the same as usefulness
  2. work finished after the caller left
  3. goodput, not requests per second
  4. rejection must cost far less than serving
  5. shed before you allocate anything expensive

basics

~20 s

Accepting more work than it can finish makes a server spend its capacity on requests whose callers have already timed out, so goodput — useful completed work — collapses toward zero. Shedding excess load early and cheaply keeps the fraction it does serve healthy.

solid answer

~60 s

A server that accepts everything does not degrade gracefully; it collapses. Work in progress is bounded by Little's Law — concurrency equals arrival rate times latency — so when latency climbs, in-flight requests climb with it, and each extra one adds memory, context switches, lock contention and connection-pool pressure that push latency higher again. That is a positive feedback loop. Worse, by the time a request finally completes, its caller has usually timed out and thrown the response away, so the machine is busy at 100% CPU producing almost no `goodput` — successful work someone is still waiting for. The fix is admission control: measure a local saturation signal such as in-flight count or queue wait, and reject beyond it. The rejection has to be cheap and early — before you allocate a worker, parse the body or touch the database — because if rejecting costs a tenth of serving, a 10x overload still saturates you. Return a fast 503 or 429 and keep serving the fraction you can actually finish.

code

python · 18 lines
python
import threading, time

MAX_IN_FLIGHT = 200          # measured concurrency at the knee, not a guess
gate = threading.BoundedSemaphore(MAX_IN_FLIGHT)

def do_work(request):
    time.sleep(0.01)
    return "ok"

def handle(request):
    if not gate.acquire(blocking=False):   # cheap: no worker, no parse, no DB
        return 503, "overloaded"
    try:
        return 200, do_work(request)
    finally:
        gate.release()

print(handle({"path": "/checkout"}))

go deeper

for a junior

Know that a server which accepts more work than it can finish gets slower for everyone rather than serving a steady subset, and that returning an error quickly is a legitimate, deliberate behaviour.

for a middle

Be ready to explain the feedback loop: rising latency raises in-flight concurrency, which raises latency again, and completed work that the caller already abandoned is wasted capacity. Name goodput and relate it to Little's Law.

for a senior

Show that you would place the admission check where a rejection is genuinely cheap, drive it from a measured saturation signal rather than a hardcoded RPS number, and reason about the mix of request costs behind a single limit.

for a principal

Own the tradeoff between an early, cheap, ignorant shed at the edge and a late, informed, expensive one at the service, and decide what the platform provides by default so every team is not re-deriving admission control on its own.

## Throughput is not the same as usefulness Every server has a load level past which it can no longer finish work as fast as work arrives. The intuition most people carry is that throughput plateaus there: offered 30,000 requests per second, a 10,000-per-second service serves 10,000 and fails 20,000. Real systems that accept everything do something much worse — total successful work falls, often to a small fraction of nominal capacity, while the machine reports full CPU utilisation. The name for the quantity that collapses is **goodput**: work that completes *and* is still wanted by the caller who asked for it. Throughput counts responses produced; goodput counts responses somebody received in time to use. ## Why accepting everything is self-amplifying Little's Law gives the mechanism in one line: the average number of requests in the system equals the arrival rate multiplied by the average time each spends there (L = λW). At 500 requests per second and 40 ms of latency, about 20 requests are in flight. Hold arrival rate constant and let latency rise to 2 seconds and 1,000 are in flight — a 50x increase in concurrency that nobody asked for. That concurrency is not free. Each in-flight request holds a thread or coroutine, a buffer, a slot in a connection pool, and some heap. As concurrency grows you get more context switching, more garbage-collection pressure, more lock and cache contention, and more competition for the same downstream connections. All of those raise latency. Higher latency raises concurrency again. The loop runs until something falls over — the heap, the file-descriptor limit, the thread pool, or the database's connection limit. ## The dead-request problem The second effect is the one that turns a slowdown into a collapse. Clients have timeouts. When queueing delay exceeds the client timeout, the caller gives up, but the server has no idea — it keeps the request, runs it, queries the database for it, serialises a response, and writes it to a socket nobody is reading. Capacity is being consumed at full rate on work with zero value. In a sustained overload with a FIFO discipline, this becomes the *normal* case rather than the exception: the requests at the head of the queue are always the oldest, therefore always the most likely to be expired. This is why the honest answer to "why is success rate near zero at 3x load?" is not "the machine is too slow" but "the machine is busy doing work that has already been discarded". ## Shedding restores goodput Load shedding means deciding, at admission time, that a request will not be served, and saying so immediately. Its purpose is not politeness to the client; it is protecting the requests you *are* going to serve, by keeping concurrency at the level where latency is still under the client's timeout. Two properties make shedding work: **Reject cheaply.** A rejection must cost dramatically less than a service. If a 503 costs 1% of a served request, you can absorb roughly a hundredfold overload before rejection alone saturates you; if it costs 10%, only about tenfold; if you shed *after* authenticating, deserialising and hitting the database, you have spent nearly the full cost and freed almost nothing. Push the check to the front of the request path. **Trigger on a local signal, not a static number.** A fixed "10,000 RPS" cap is wrong the moment request mix changes, a dependency slows down, or you deploy a heavier code path. Trigger on something that reflects real saturation — in-flight request count, queue wait time, or a measured concurrency limit — so the threshold moves with the actual cost of work. ```python MAX_IN_FLIGHT = 200 # measured at the knee, not guessed if in_flight >= MAX_IN_FLIGHT: return 503 # before parsing, auth, or any DB call ``` ## What the response should look like Say it fast and say it clearly: HTTP 503 for "I am out of capacity" or 429 for "you specifically are over your allowance". Keep the body tiny, and be careful not to log every rejection at full verbosity — a per-rejection log line with a stack trace is its own overload amplifier during an incident. On the client side, a shed response should be cheap to observe and should not look like a hard, permanent failure. ## The interview point The decision with a cost here is *where* in the request path you place the check. Early means cheap but ignorant — the front door does not know whether this request is a 2 ms cache hit or a 900 ms report. Late means informed but expensive. Most real systems do both: a cheap concurrency gate at the edge of the process, plus a finer decision once the request's cost class is known. A candidate who only says "add more capacity" has missed that you cannot add capacity in the sixty seconds you have, and that a service which sheds correctly survives an overload it cannot possibly serve.

  • How would you pick the concurrency limit that triggers shedding, given that request cost varies?
    Measure rather than guess: run load until latency leaves its flat region and record the in-flight count there, then set the limit just below it. Better still, use an adaptive limiter that adjusts the ceiling from observed latency, so the limit tracks the current request mix and dependency speed instead of a number that was true on the day you benchmarked.
  • If shedding protects the server, why can it still make the overall outage worse?
    Because a shed request usually comes back. Clients retry, users refresh, and upstream services fail over, so a rejection can return as more load a second later. Shedding only ends the incident when the offered load actually falls or the rejections are cheap enough that the loop is stable; otherwise you also need to stop the retries at their source.

A kitchen that accepts every ticket keeps cooking meals for diners who already walked out — full stoves, empty tables. Turning people away at the door is what keeps the seated diners fed.

saying these in an interview costs you the question

  • Assumes throughput plateaus at capacity rather than collapsing
  • Says the answer is simply to add capacity mid-incident
  • Rejects requests only after authentication and a database query
  • Treats every 503 as a failure to be avoided at all costs
  • Believes queueing the overflow is the same as handling it

context