Why can a few thousand trickle-fed requests exhaust a web tier at almost no cost to the attacker?
answer
- compare ledgers, not volumes
- one socket buys one worker
- the slowest traffic has the best ratio
- throughput equals concurrency over hold time
- not sending more, just not letting go
basics
~20 sRank attacks by what the attacker keeps committed per unit of your capacity. A request dribbled out over minutes costs them one socket and a few bytes; it costs you a connection entry and a pinned worker throughout.
solid answer
~50 sRank by committed cost per unit of the target's capacity. A request that arrives one header line every ten seconds costs the attacker a single socket and a few bytes per minute. It costs the target a connection-table entry, socket buffers, a negotiated TLS session, a worker or thread pinned for the entire duration, and frequently an open downstream connection as well. That ratio - near zero for them, a whole slot for you - is why slow traffic beats large traffic against anything with bounded concurrency. The governing arithmetic is Little's Law: sustainable requests per second equals concurrency divided by average hold time, so stretching hold time from 30 ms to 300 s cuts throughput by four orders of magnitude without a single extra byte per second. The attacker is not sending more. They are refusing to let go.
go deeper
Recall that an open request costs the server a slot for as long as it stays open, whether or not it is transferring anything. Sending slowly is a way of paying almost nothing to occupy something expensive.
Be ready to lay out both ledgers - socket and a trickle of bytes on one side, connection entry plus pinned worker on the other - and to state that sustainable throughput is concurrency divided by hold time.
Demonstrate that you attack the ratio rather than the volume: what stops committing a worker before the request is complete, what bounds the hold, what makes waiting cheap. Say why byte counts hide this failure entirely.
Own the position that concurrency budgets and hold times are capacity commitments the business is making implicitly. The strategic point is that availability spend aimed at bandwidth buys nothing against an adversary optimising this ratio.
## The metric: committed cost per unit of capacity Availability attacks are usually ranked by volume, which is the wrong axis. The axis that predicts what actually works is the ratio between two quantities: - what the **attacker** must keep committed - sockets, memory, bandwidth, patience - to hold one unit of the target hostile; - what the **target** commits in response, and for how long. When that ratio is heavily in the attacker's favour, they win with a laptop. When it is even close to parity, they need infrastructure. Slow, held traffic sits at the extreme favourable end, and this is the single most important thing to internalise about availability: **the cheapest attack per unit of capacity is usually the slowest traffic, not the largest.** ## Working the ratio Consider an unauthenticated attacker who holds nothing but a published URL. They open a connection and begin a request, then deliver it at a deliberately glacial pace - a header line every few seconds, or a body advertised as long and dribbled out a byte at a time - staying just inside whatever idle tolerance the server allows. Alternatively they issue a complete request to an endpoint known to be slow and simply wait for the answer. **Their side of the ledger, per held slot:** one socket on their host, a few hundred bytes of kernel state, and roughly one packet every few seconds. A single modest machine sustains tens of thousands of such slots, and the aggregate outbound rate is measured in kilobits per second. **Your side of the ledger, per held slot:** a connection-table entry on the server and on every intermediary the connection crosses; receive and send buffers; TLS session state established by a handshake you already paid CPU for; a worker or thread pinned for the full duration, along with any per-request memory it accumulated; and, if the endpoint has already called downstream, an open connection into that dependency too. So the attacker spends bytes and gets slots. That is the trade the whole technique class turns on. ## Little's Law is the reason it scales For any bounded concurrency, sustainable throughput obeys `throughput = concurrency / average hold time`. Three consequences follow, and they are what an interviewer is checking for: 1. **Hold time and concurrency are interchangeable weapons.** Filling half the pool and doubling everyone's hold time produce the same collapse. An attacker who cannot get more slots can instead make the slots they have stay longer. 2. **Throughput is not a function of arrival rate.** A pool serving 30 ms requests at 400 concurrency sustains around 13,000 requests per second; if hold time becomes 300 seconds, the same pool sustains a little over one. The attacker never raised their request rate. 3. **The ceiling is invisible in byte counts.** Nothing about this shows up as a large number of bytes anywhere, which is why an estate that is confident because its uplink is quiet is confident about the wrong thing. ## Why the attacker's stopping condition is cheap too A volumetric attack must be *sustained*: the moment the bytes stop, the link drains and service returns within seconds. A holding attack has hysteresis in its favour. Slots come back only as the server times them out, so the attacker pays for occupancy but gets the outage for free during the release period, and they can top up occupancy at a fraction of the rate they built it. Cost per minute of outage falls the longer they hold. ## The boundary of the technique Holding attacks are strongest against anything with a small bounded concurrency and a generous tolerance for slow peers, and weakest where a request commits nothing until it is complete. The attacker's cost rises sharply when the target refuses to allocate a worker before a full request exists, when the maximum hold is bounded, or when waiting is cheap because the server does not dedicate a thread to it. Those are control classes rather than products: each one attacks the *ratio* rather than the volume, which is the only thing that generalises. ## Two claims to get the right way round A held connection proves that a peer opened a socket and is sending slowly. It does not prove hostility - a mobile client on a failing link, a scraper with a tiny buffer, and a deliberate hold are indistinguishable one connection at a time; the population is what distinguishes them. And low byte volume is not evidence of low impact: the resource being drained is measured in slots and seconds, not in bits per second.
- Which is cheaper for the attacker: opening more slow connections, or making each one last longer?Extending the hold, almost always. New connections cost a handshake, possibly a TLS negotiation, and a fresh source socket, while extending an existing one costs one small packet before the idle timer expires. Little's Law treats the two as equivalent for the defender, but their prices differ by orders of magnitude for the attacker, which is why holds get stretched to just inside whatever tolerance exists.
- Where does this technique stop working?Where a request commits nothing until it is complete, and where waiting does not consume a dedicated slot. If the server buffers the full request before allocating a worker, a partial request holds only cheap state; if waiting is a suspended task rather than a pinned thread, tens of thousands of concurrent waits cost little. The attacker's ratio collapses in both cases.
- How is this different from an attack that makes each request expensive to compute?Both drain the same pool, but through different currencies. Expensive computation burns CPU, so it is visible as a saturated processor and it costs the attacker a completed request per unit of damage. A holding attack burns time on an otherwise idle machine, so the CPU stays low and the attacker pays almost nothing per unit. The resource is the same; the ratio is not.
Ten people who each buy one coffee and sit all afternoon close a cafe more cheaply than a coach party that eats and leaves. The bill is trivial; the tables are gone.
saying these in an interview costs you the question
- Measures an availability attack only in bits per second
- Assumes low traffic volume means low impact
- Believes a bigger connection limit removes the problem
- Thinks a busy pool must mean a busy CPU
- Calls one slow connection proof of hostility