A public site refuses new visitors while its internet link is far from full - what else ran out?
answer
- the pipe is not the only ceiling
- four resources, only one has an invoice
- what it holds, not what it weighs
- slots and table entries, not bytes
- one shared dependency caps everybody
basics
~20 sBandwidth is one of four exhaustible resources. A site also runs out of connection-table entries, worker-pool slots, or headroom in one slow dependency every request waits on - any of which fails while the link sits idle.
solid answer
~50 sFour things exhaust independently, and only one of them is bandwidth. The uplink fills when arriving bytes exceed its capacity. The connection table fills when open connections exceed what the host and every intermediary in front of it can hold, whether or not bytes flow over them. The worker or thread pool fills when in-flight requests exceed the concurrency the application can serve - and a request that is merely waiting still occupies a slot. The one slow dependency every request calls fills when concurrent calls exceed its fixed budget, which is often somebody else's to raise. An idle link with a dead site means one of the last three ran out. That is why 'we are being flooded, buy more bandwidth' is the wrong first sentence: ask what a request *holds*, not what it *weighs*.
go deeper
Be ready to name the four exhaustible resources and say which one an idle uplink rules out. The key recall is that a connection costs something just by existing, even if it carries almost no data.
Expect to explain how each resource is bounded and how you would tell them apart from the outside - link near capacity, connection count at its ceiling, or workers busy while the CPU idles.
Show you reach for the resource before the cause. Say out loud that a dependency slowdown and a deliberate hold exhaust a worker pool identically, and that naming the resource is what makes a remedy arguable.
Own the framing that availability is a portfolio of independent ceilings rather than one number. The organisational value is stopping a reflex spend on the resource with the clearest invoice attached.
## The symptom names nothing From the outside, every availability failure looks identical: a spinning tab, a connection that never completes, a timeout. 'The site is down' does not tell you which finite thing ran out, and until you know that, you cannot say whether more capacity would have helped, whether the traffic was hostile, or even whether the volume was large. The useful first question is not *how much traffic arrived* but *which resource was committed and never released*. A request travelling through a public web estate consumes four separable resources. They have different units, different ceilings, and they exhaust independently. ### 1. Link bandwidth The physical capacity of the uplink, in bits per second. It fills when the sum of arriving bytes exceeds what the pipe can pass. This is the resource everyone names first, because it is the one with a number on an invoice and the one the phrase 'denial of service' conjures. Its signature is collateral: when a shared link saturates, everything behind it degrades at once, including services that were never the target. If the link is not near its ceiling, bandwidth is not what ran out - and a great many outages are exactly that case. ### 2. The connection table Every accepted connection occupies an entry - addresses and ports, send and receive buffers, TLS session state - on the server and in every intermediary the traffic crosses, such as a load balancer or reverse proxy. The ceiling is set by memory and by configured limits. The important property is that this cost is paid for *existence*, not for *use*: a connection that has been open for ten minutes and carried forty bytes costs the same table entry as one serving a busy user. Counting bytes tells you nothing about how full this table is. ### 3. The worker or thread pool The concurrency the application itself can serve. In a thread-per-request server, one in-flight request pins one thread and its stack; in an event-driven or coroutine server, it pins a slot in a bounded concurrency budget, plus whatever state that request accumulated. The governing arithmetic is Little's Law: sustainable requests per second equals concurrency divided by average hold time. A pool of 400 workers serving 30 ms requests sustains roughly 13,000 requests per second; the same pool serving requests that hold for 300 seconds sustains about 1.3. Nothing about the pool changed - only how long each occupant stays. ### 4. One slow dependency every request waits on A shared database, an identity provider, a licence check, a third-party pricing or fraud API. If the hot path cannot complete without it, its concurrency budget is the estate's real ceiling regardless of how many replicas the front end runs. Two properties make it the sharpest of the four: the budget is usually fixed, and it is frequently not yours - you cannot buy more of somebody else's capacity, and your own scaling adds pressure to it rather than relieving it. ## Reading the failure backwards Because the four exhaust independently, what is *not* saturated is as informative as what is: - Link near capacity, everything else fine: a volumetric problem. - Link idle, connection count at its ceiling, workers mostly idle: connections are being held without being used. - Link idle, workers all busy but consuming almost no CPU: the workers are *waiting* - on a slow dependency, or on request bodies arriving a byte at a time. - Link idle, workers busy and CPU pinned: expensive computation per request, not a holding problem. ## Why this is the whole point of the leaf The wrong answer a competent engineer gives is to treat every availability attack as a volumetric one, because that is the only variety with a familiar picture attached. It leads directly to buying the wrong thing. An unauthenticated attacker holding only a published URL does not need to out-supply your uplink; they need to find whichever of the other three commits earliest and costs them least to keep committed. Ranking the four by *what the attacker must keep committed per unit of your capacity* is what turns 'the site is down' into a decision. One caution on direction: an idle link proves that bytes were not the limit. It does not prove hostility. Legitimate traffic, a dependency that slowed down on its own, or a bad deploy that doubled hold time exhausts a worker pool in exactly the same way. Naming the resource is the first step; naming the cause is a separate one.
- If the link is idle and every worker is busy but the CPU is nearly idle, what are those workers doing?Waiting, not computing. Either they are blocked on a downstream call that has slowed or saturated its own concurrency budget, or they are reading a request that is arriving a fragment at a time. Both pin a slot for the full duration while consuming almost no CPU, which is why worker-pool exhaustion and CPU exhaustion are different failures that need different answers.
- Does an outage caused by ordinary traffic look any different from the outside?No. Exhaustion presents identically whichever cause filled the resource - a launch-day surge, a dependency that got slower, a deploy that doubled hold time, or a deliberate hold. That is precisely why the resource has to be named before the cause: the phrase 'the site is down' is compatible with all four, and picking a remedy from the symptom alone is guesswork.
A restaurant can turn people away because the door is jammed, because every table is occupied by someone nursing a coffee, or because the single dishwasher cannot keep up. Only one of those is about how many people are on the pavement.
saying these in an interview costs you the question
- Calls every availability attack a volumetric flood
- Assumes an idle uplink means nothing hostile happened
- Confuses how big a request is with how much it costs
- Thinks a bigger instance always means more capacity
- Treats connection count and request rate as the same number