'Just buy more bandwidth' - which of a web estate's exhaustible resources does that actually relieve?
answer
- money moves exactly one of the four
- more slots, same price per slot
- the pool is the queue, not the service
- you cannot autoscale somebody else's contract
- change the ratio, not the ceiling
basics
~20 sOnly a saturated link. Bigger connection tables and worker pools add slots an attacker occupies almost free, and extra workers push more concurrent calls into the one slow dependency, whose capacity is usually not yours to buy.
solid answer
~50 sOnly one of the four resources, and only up to a point. More uplink relieves a saturated link until the attacker's supply exceeds the new number. It does nothing for the other three, and for two it is actively harmful. Raising connection-table limits and worker-pool size buys more slots the attacker must occupy - but their cost per slot is a socket and a trickle of bytes, so you are scaling linearly against something nearly free, and each extra slot costs you real memory. Worse, more concurrent workers means more concurrent calls into the one slow dependency every request waits on, whose budget is fixed and usually not yours to raise, so it fails sooner and takes legitimate traffic with it. The answer is not to buy more of what gets committed, but to stop committing it before anything is proven.
code
text · 7 linesresource attacker commits per unit does buying more help?
------------------ ----------------------------- ------------------------------------
link bandwidth sustained bits per second yes, until their supply exceeds yours
connection table one socket, held open barely - more slots, same near-zero cost each
worker pool one held or waiting request no - more waiters on the same bottleneck
slow dependency one held or waiting request cannot buy: the budget is someone else's
...go deeper
Know that more bandwidth only helps if bandwidth was the thing that ran out, and that a web tier has other ceilings that money does not move.
Be ready to explain why adding worker slots can make things worse: the extra waiters consume memory and push more concurrent calls into whatever was already the bottleneck.
Show the resource-by-resource reasoning and land on changing the attacker's cost per unit rather than raising your own ceiling. Name the control classes - commit late, bound the hold, make waiting cheap, shed load.
Own the conversation with the budget holder: state plainly which ceiling money can move, which it cannot, and what the engineering alternative costs. The organisational failure is a reflex spend on the resource with the clearest invoice.
## The question behind the question When an owner says 'buy more bandwidth', they are proposing to convert money into availability. That is a reasonable instinct and sometimes the right call - but it only works for one of the four exhaustible resources, and applied to the others it ranges from useless to counterproductive. Answering well means going resource by resource and saying what the purchase actually does to the *attacker's* cost. ### Link bandwidth: yes, genuinely A saturated uplink is the one case where buying capacity attacks the adversary's economics directly. Volume is the one currency an attacker must supply continuously and cannot fake: to keep a bigger pipe full they must find more of it, sustained, for the whole outage. Doubling the link doubles what they must supply. The purchase is a real defence with an honest limit - it is a race you can lose if their supply exceeds yours, and it does nothing at all if the link was never the constraint. ### The connection table: barely Raising connection limits adds slots. Each new slot costs you memory on the server and on every intermediary in the path, permanently, whether or not it is ever used. It costs the attacker one more socket and a few bytes a minute. You are scaling linearly against something priced near zero, so the ratio moves in their favour with every slot you add - and you have also enlarged the blast radius of memory exhaustion, converting a clean refusal of new connections into a host that swaps or dies. It buys time, measured in how long it takes them to open a few thousand more sockets. ### The worker pool: usually worse than useless More workers means more requests in flight simultaneously. If those requests are waiting rather than computing, you have added waiters, each with its own memory and its own open downstream connection. Two things get worse: the machine's memory footprint grows with a pool that spends its life blocked, and the concurrency arriving at everything downstream grows in proportion. This is the counter-intuitive part an interviewer is listening for - the pool is not the bottleneck, it is the *queue* in front of the bottleneck, and lengthening a queue never adds service capacity. ### The one slow dependency: you cannot buy it If every request on the hot path calls a shared database, an identity provider, or a third-party API, that call's concurrency budget is the estate's true ceiling. Frequently the budget belongs to someone else and is set by a contract, not by an autoscaler. Scaling your own front end raises the pressure on it. This is the resource where money genuinely does not convert into availability, and where an architect should say so plainly. ## What the money should buy instead The generalisation is that capacity purchases raise the *amount* the attacker must occupy, while the attacker's cost per unit occupied stays flat. Any linear defence loses that race. What changes the ratio is refusing to commit an expensive resource on the strength of an unproven request: - **Do not allocate an expensive slot until a request is complete.** A partial request should hold only cheap state, so a slow sender occupies a buffer rather than a worker. - **Bound the hold.** Every commitment gets a maximum duration; unbounded waits are how one slow peer converts into a permanently occupied slot. - **Make waiting cheap.** A suspended task waiting on a downstream answer costs a fraction of a pinned thread, which raises the number of concurrent waits that fit in the same memory by orders of magnitude. - **Cap concurrency per source and per dependency.** A bounded budget per caller stops one source monopolising the pool; a bounded budget into the dependency with a fast failure keeps its exhaustion from becoming everyone's exhaustion. - **Shed load deliberately.** Refusing a request in microseconds is dramatically cheaper than serving it slowly, and a system that degrades by refusing is one that has chosen which resource it is protecting. ## The sentence to say to the owner 'Bandwidth is the only one of our four ceilings that money moves in our favour, and it was not the ceiling we hit. For the other three, buying more raises the number of slots the attacker has to fill without raising what filling one costs them - and for the dependency, the capacity is not ours to buy. The spend that changes the outcome is the engineering that stops us committing a worker to a request nobody has finished making.' ## Direction of the claims Adding capacity is a preventive control against volume and nothing else. A larger pool that survives one attack does not prove the attack got harder - it may only prove the attacker did not bother to scale a nearly free input. And the fact that an estate has never been taken down this way is not evidence that its ratios are sound.
- Why is adding replicas in front of a fixed-capacity third-party API sometimes worse than doing nothing?Because it raises the concurrency arriving at the thing that was already the ceiling. The extra replicas cannot make the dependency answer faster, so they convert a partial outage into a total one: the dependency saturates sooner, its latency climbs for everybody, and requests that would have been refused quickly now wait and hold slots. You have lengthened the queue in front of an unchanged server.
- When is buying uplink capacity genuinely the right answer?When the link was demonstrably the saturated resource, when the volume needed to fill the new one is meaningfully harder to supply than the old, and when the alternative engineering cannot be delivered in the time available. It is a legitimate purchase with an honest limit, and the failure mode is buying it reflexively for an outage where the link was idle.
- How do you justify shedding load rather than serving it slowly?Refusing a request costs microseconds and no slot; serving it slowly costs a slot for its whole lifetime and degrades everyone else's. A system that refuses early keeps a defined fraction of users working, while one that queues indefinitely converts a partial shortage into a total outage. Choosing which resource to protect is a design decision, not an accident.
saying these in an interview costs you the question
- Answers every availability problem with more capacity
- Assumes a bigger worker pool always adds throughput
- Ignores that extra slots cost memory permanently
- Forgets a third party's capacity is not purchasable
- Treats surviving one attack as proof of hardening