skip to content

Why does a bits-per-second DDoS threshold never fire when an attacker's 400 search requests a minute exhaust the database pool behind a public API?

level: juniorimportance: must knowfreq 62%

answer

  1. the two ends measure different things
  2. cheap to send, costly to answer
  3. bits per second versus work per request
  4. border counters far below threshold
  5. the scarce resource sits downstream

basics

~20 s

Because the threshold and the damage are measured in different units. Four hundred small, well-formed requests are a trivial number of bits and packets; the cost lands as work per request inside the service, which the border never counts.

solid answer

~50 s

A volumetric mitigation device meters what it can see from the wire: bits per second, packets per second, new connections per second, connection state. An attacker who has found an expensive query - a wide date range, an unbounded export, a filter that defeats every index - spends almost nothing to send it. Four hundred such requests a minute is a few megabits and a handful of connections, so every border counter sits two or three orders of magnitude below its threshold and the appliance correctly reports zero mitigations. The scarce resource is downstream: the shared database pool, its workers, its temporary space. The mismatch is one of units - bits per second at the border against work per request at the service - and no tuning of a traffic-rate threshold closes it, because the signal is simply not present in that measurement.

code

text · 11 lines
text
border mitigation appliance -- 60 s window
  bits/s ingress ......... 6.2 Mbps    (mitigation threshold 4 Gbps)
  packets/s ingress ...... 1.9 kpps    (threshold 1.2 Mpps)
  new connections/s ...... 7           (threshold 20,000)
  mitigation events ...... 0

service tier -- same 60 s window
  requests received ...... 412
  p99 response time ...... 46 s        (7-day baseline 210 ms)
  shared DB pool ......... 200 / 200 busy, 3,900 queued
  ...

go deeper

for a junior

Know what a volumetric device actually counts - bits, packets and connections per second - and be able to say plainly that a few hundred costly requests move none of those numbers.

for a middle

Explain the unit mismatch precisely: the threshold is a traffic rate, the damage is work per request, and no threshold on the first quantity can detect the second.

for a senior

Show how you confirm it in production: put border counters and service-side concurrency and pool saturation in the same window, and resist the instinct to tune the appliance.

for a principal

Own the consequence that a control you fund and report on has a class of outage it can never see, and make that gap explicit before availability commitments are written on top of it.

## The two numbers that never meet A volumetric mitigation device sits in the path and makes decisions from quantities it can derive from packets alone: ingress bits per second, packets per second, new connections per second, connections in each TCP state, fragment ratios, per-source packet counts. Those are the right units for a flood. A reflected amplification attack or a packet flood is *defined* by moving more bits or more packets than the path can carry, so a threshold on bits or packets is a direct measurement of the damage. A low-volume expensive-request attack inverts that relationship. The attacker's whole strategy is to find an ask that is cheap for him to send and costly for you to answer, then to send just enough of them. The asymmetry does the work; the volume does not. A single well-formed search that scans an entire table, an export with no upper bound on its date range, a report that joins across a shared analytics store - each is a few hundred bytes on the wire and tens of seconds of a database worker on the other end. ## Why no threshold you could set would catch it Suppose the service normally takes 300 requests a minute and the attack adds 400. On the border counters that is a change of a few megabits per second on a link provisioned for gigabits. To fire on it, you would have to set a threshold below your own ordinary Tuesday afternoon, and far below the traffic of any legitimate launch, marketing email or client retry. The device would then mitigate your customers routinely and still not be measuring the thing that is hurting you: it would be firing on a proxy that correlates with load only by accident. This is the important generalisation, and it is the one interviewers are testing. **The measurement carries no information about the property you care about.** Increasing resolution, sampling more often, or adding more counters of the same kind does not help, because a request's cost to answer is not a function of its size on the wire. Two requests of identical length, identical protocol and identical source can differ by four orders of magnitude in what they cost to serve. ## What does move during the attack It is worth knowing which numbers *do* change, because that is how you confirm the diagnosis: | Where you look | What you see | | --- | --- | | Border bits/packets/connections per second | Flat, far below any threshold | | Mitigation event log | Zero events, and correctly so | | Concurrent established connections | Drifting up as responses stall | | Service request count | Barely changed | | Service p99 latency | Enormous | | Shared database pool | Fully busy, deep queue | | Client-side result | Timeouts, resets, an unusable service | The border-side symptoms - longer-lived connections, more concurrent sockets, eventual resets - are the *service failing* observed from the wire. They arrive after the damage, not before it, and they look identical to a slow backend for any other reason. They are not a volumetric signal. ## The direction of the claim Be careful about what zero mitigations proves. It proves that no counter the device meters crossed a configured threshold. It does not prove that no attack occurred, that no hostile traffic arrived, or that the estate is healthy. The absence of a detection is a statement about the detector's inputs. On this class of attack the detector was never given an input that could carry the evidence, so its silence is guaranteed in advance and tells you nothing at all about the adversary. ## What the defender is paying Two things, and naming both is what separates a real answer from a definition. First, the capacity: the shared database pool, held and funded for legitimate peak, is consumed by requests that cost the attacker almost nothing to issue - the asymmetry is the whole attack. Second, the visibility: the organisation is paying for and reporting on a mitigation control that has a class of outage it can never see, and unless someone says so out loud, availability commitments get written on the assumption that the control covers it. ## The answer an interviewer wants State the unit mismatch in one sentence - the threshold is a traffic rate, the damage is work per request - and then say where the control has to live instead: at a component that can see the request and know what it will cost before the expensive work begins. That is a different vantage from the border, and it comes with its own bill.

  • The appliance reports zero mitigations. Does that prove nothing malicious reached the service?
    No. It proves only that no counter the device meters crossed its threshold. The device measures traffic volume, not request cost, so an attack whose entire point is that it is small is invisible to it by construction. Absence of a mitigation event is evidence about the detector's inputs, never about the estate.
  • Would lowering the bits-per-second threshold or sampling more finely help?
    No, and lowering it actively hurts. The attack sits orders of magnitude below any threshold you could set without mitigating your own ordinary busy hour or a legitimate launch. Finer sampling measures the same quantity more precisely, and that quantity carries no information about what a request will cost to answer.
  • Which numbers do move at the border while this is happening?
    Connection duration rises as responses stall, concurrent established connections drift up, and eventually client timeouts appear as resets. Those are the service failing, seen from the wire: they arrive after the damage, look the same as any slow backend, and are not a volumetric signal you can threshold on in advance.

A weighbridge at the gate weighs every truck and waves through anything light. A single envelope holding a court order weighs nothing and ties up the whole legal department for a week.

saying these in an interview costs you the question

  • Assumes flat bandwidth graphs mean no attack is happening
  • Proposes lowering the bits-per-second threshold to catch it
  • Calls it a volumetric flood because the service is down
  • Reads zero mitigation events as proof the estate is healthy

context