skip to content

Many Sites, One Address

Spreading an attack over many sites means no site fails, and it also means no site's threshold fires. Interviewers use it because candidates argue capacity when the interesting loss is visibility.

on this pageshow

questions

4

One address announced from twelve sites absorbs a flood, yet no site's rate threshold fires — why?

level: juniorimportance: must knowfreq 62%

answer

  1. each site sees only its share
  2. the threshold is local, the flood is global
  3. nobody adds the twelve numbers up
  4. absorbing is not detecting
  5. no alert is not no attack

basics

~20 s

Each site sees only its own share of the flood. A per-site threshold is compared against roughly one twelfth of the total, so an attack far above your planning number can stay under every local trigger and never alert anyone.

solid answer

~50 s

Announcing one address from many sites means the attacker's sources are divided among them by routing, so no single site ever holds the whole flood. Thresholds are almost always set per site, from that site's own baseline, because that is where the counters live. If the estate receives twelve times what any one site receives, a flood that would have flattened a single-site deployment shows up as twelve unremarkable bumps. The absorption is real — that is the point of the design — but the *detection* is not, and nothing in the architecture computes the sum. So you find out later: a transit bill, a support ticket about latency, an upstream provider mentioning it. The fix is aggregation, not sensitivity: export flow records from every site to one collector and alert on the estate total and on each site's share of it.

go deeper

for a junior

Be ready to say plainly that a site can only measure its own traffic, and that a threshold set there is compared against a fraction of what the company received.

for a middle

An interviewer expects you to walk the arithmetic — total, share per site, trigger per site — and to explain why raising sensitivity locally trades a missed attack for false pages.

for a senior

Show the instrumentation change you would actually make: flow export from every edge into one collector, an alert on the estate total, and a second alert on each site's share of it.

for a principal

Own the claim being made to the business. 'We absorb attacks' and 'we detect and absorb attacks' are different promises, and the footprint only delivers the first until someone funds the aggregation.

## The design does two things, and only one of them is detection When a company announces the same address block from a dozen of its own sites, the internet's routing does the load spreading. Each source network on the internet picks one of your sites according to its own routing preferences, and its traffic goes there. An attacker who points a large botnet at that address does not get to choose which site each source lands on; the flood is divided among your sites the same way legitimate traffic is. That is genuinely an absorption strategy: no single uplink has to hold the whole thing. What it is not is a detection strategy, and conflating the two is the mistake this question is aimed at. ## Where thresholds actually live Thresholds are set where the counters are. Each site's edge routers and flow exporters see that site's traffic, a baseline is learned from that site's history, and an alert is defined as some multiple of it — say, three times the ninety-fifth percentile of the last month. Every site is instrumented identically and every site is blind to the other eleven. Now run the arithmetic. Suppose an attack delivers 240 Gbps to the estate and it splits roughly evenly. Each site receives about 20 Gbps. If a site normally peaks at 12 Gbps and its trigger sits at 36, nothing fires. Twelve sites report a busy evening. The estate absorbed a flood larger than the total capacity of any one site, and produced no alert at all. The attacker did not have to be clever about this. They did not size the flood to hide under your thresholds — the topology did that for free. Any attack under twelve times your smallest local trigger is invisible by construction. ## What it costs you The traffic was still paid for. It consumed transit at every site, ate the headroom you bought for a real growth curve, and if the flood pushed a site into congestion it degraded real users while nobody was looking. The bill arrives at the end of the month, or a customer opens a ticket about slow responses from one region, and only then does anyone go back through the graphs. And the write-up afterwards has to be honest about a claim you cannot make. "We absorbed it" is true. "We detected and absorbed it" is not, and the difference matters to anyone deciding whether the footprint is doing the job they funded. ## Why lowering thresholds is the wrong reflex The instinct is to make each site more sensitive so a twelfth of a big flood trips it. But a twelfth of a flood is not much above an ordinary busy hour. Push a local trigger down into the band where normal diurnal variation lives and you page someone for a sports final, a product launch, or a crawler. You have traded a missed attack for a stream of false positives, and the people carrying the pager will raise the thresholds again within a month. Sensitivity is the wrong axis. The problem is not that each site's number is too high; it is that the interesting number is never computed anywhere. ## What to build instead Export flow records — the five-tuple, byte and packet counts, and timestamps — from every site's edge to one collector, and define the alert over the sum. Then add the second signal the split gives you for free: each site's *share* of the total. Under normal traffic the shares are lumpy but reasonably stable. A flood that arrives from an unusual set of source networks will move those shares, and a share that doubles while the total climbs is a stronger signal than either number alone. Keep the per-site alerts as well. They still catch the case where the attack sources are concentrated in a few networks and most of the flood lands at one site — which does happen, and which then looks like a purely local event to everyone except the person holding the aggregate. ## The sentence to say out loud No alert fired is not evidence that nothing arrived. In a distributed footprint it is the expected outcome, and treating silence as an all-clear is how an estate absorbs an attack for a week without knowing it.

  • If nothing alerts, how do you usually find out the attack happened at all?
    Later, and from somewhere other than monitoring: a transit or bandwidth bill higher than the forecast, a support ticket about latency from one region, a peer or upstream mentioning they saw a burst toward your prefix, or an engineer reviewing capacity graphs for an unrelated reason. Sometimes never — a flood the estate swallowed without congestion leaves nothing but a cost line.
  • Would lowering every site's threshold fix it?
    No. To catch a twelfth of a large flood, each local trigger has to sit inside the band where ordinary daily variation already lives, so you buy false pages for busy evenings and launches. The pager owners will raise the thresholds back within weeks. The problem is the missing sum, not the sensitivity of the parts.
  • Does the split ever help detection rather than hurt it?
    Yes, when the attack sources are concentrated. If most of the flood comes from a handful of source networks that all route to the same site, that site takes a disproportionate share and does cross its threshold. It then looks like a local incident, which is its own trap: the responder investigates one site's link while the rest of the estate quietly carries the remainder.

Twelve tills each ring up a busy but unremarkable hour. Only when someone totals the day's takings does it become obvious that far more people came through the door than the shop was built for.

saying these in an interview costs you the question

  • Assuming an absorbed attack was therefore a detected attack
  • Believing the flood divides evenly and predictably across sites
  • Proposing to fix it by lowering every site's threshold
  • Treating 'no alert fired' as evidence nothing arrived
  • Assuming some component already computes the estate-wide total

context

open as a page

An attacker floods your origin's own address directly, ignoring your twelve-site footprint — how did they find it?

level: seniorimportance: must knowfreq 55%

basics

~20 s

From public history, usually: passive DNS showing the address your name used to resolve to, and Certificate Transparency logs listing hostnames such as the origin's own. The footprint records nothing, because no packet ever reached a site.

open as a page

What decides how much of a flood each site absorbs when one address is announced from many sites?

level: middleimportance: should knowfreq 46%

basics

~20 s

Other networks decide, not you. Each source network's own BGP best-path choice for your prefix picks the site it reaches, so the split follows routing policy and peering, not geography — uneven, and it moves.

open as a page

Your twelve sites each graph their own traffic — how do you produce one credible number for what the estate received?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Aggregate flow records from every site into one collector on a common clock, scale each site's counts by its own sampling rate, count each packet once, and report a peak per-second rate rather than a byte total.

open as a page