A blackhole stopped the flood in forty seconds but took eleven unrelated services with it, so what went wrong?
answer
- routing sees an address, not a service
- who else answers on that IP
- name-based hosting behind one front end
- the announcement may have been coarser than intended
- address plan is the real blast-radius control
basics
~20 sThe discard matched a destination address, not a service, and that address was shared. Eleven services answered on the same front-end IP, so killing the address killed all of them. The failure was the address plan and the missing inventory, not the routing decision.
solid answer
~50 sDestination-based blackholing has exactly one dial, the prefix length, and it knows nothing about what listens on the address. If eleven services share one front-end address through name-based virtual hosting or a single reverse-proxy VIP, the `/32` you announced was never one service. Two other variants of the same mistake: the on-call announced a covering `/24` because the upstream refused a longer prefix or the automation defaulted there, and the address doubled as the outbound source for other systems, so their return traffic died too. The fix is architectural rather than operational: make internet-facing addresses match the blast-radius unit you are willing to lose, keep a mapping from address to dependent services so the on-call can read the price before paying it, and reserve shared addresses for the granular upstream rule you negotiated instead of the coarse lever.
go deeper
Take away the core fact: a blackhole removes an address, and anything else answering on that address goes with it. Knowing which services share an IP is a real question with a real answer.
Explain the three routes to an oversized blast radius: shared front-end addresses, an announcement coarser than intended, and an address that also serves outbound or partner traffic. Connect each to what routing can and cannot see.
Diagnose the incident as an address-architecture failure and give a costed remediation order: build the address-to-service map, split shared front ends, cap the announced prefix length, and name the alternative for addresses that must never go dark.
Frame internet-facing addressing as a mitigation-granularity decision that has to be owned before an incident, and decide who pays for splitting shared front ends against buying a granular upstream capability.
## Read the postmortem correctly The tempting conclusion is that the on-call made a bad call. They did not. Forty seconds to end a volumetric flood is the control working exactly as designed. What failed was that nobody had ever answered the question **"what dies when we blackhole this address?"** before the night it mattered. A blackhole is a routing action. Routing sees a destination prefix. It cannot see hostnames, TLS SNI, ports, or which of your teams owns what. The unit of destruction is therefore the address, and if your architecture does not make address boundaries equal service boundaries, then your mitigation granularity is whatever your address plan happened to become over ten years. ## The three ways one host route kills eleven things **Shared front-end address.** The overwhelmingly common cause. Many services sit behind one reverse proxy or one load-balancer VIP, distinguished by hostname or SNI at layer 7. From the internet they are one address. The attacker floods one hostname, you blackhole the address that hostname resolves to, and every co-tenant of that address disappears. Note the asymmetry that makes this so easy to walk into: the attacker only needed to know one name, and you paid for all eleven. **Prefix length inflation.** The on-call intended a `/32` and announced something coarser. This happens when the upstream's inbound policy will not accept a prefix that long, when the automation was written against the allocation rather than the host, or when someone reasons that a wider announcement is safer. A `/24` takes two hundred and fifty six addresses off that path, and the postmortem then contains eleven services that were never even adjacent to the attack. **The address had a second job.** An internet-facing address is often also the source address for outbound connections, a NAT pool member, or the endpoint of a partner VPN or an API integration. Discarding traffic to it silently breaks return traffic for flows that had nothing to do with the flooded service, and those failures show up as unrelated timeouts in someone else's dashboard. ## What should have existed beforehand **An address-to-service inventory.** For each internet-facing address, what names resolve here, which teams own them, and what depends on it that is not a web request. This is a table, it is small, and it is the difference between an on-call making a priced decision and making a blind one. Reading it should take ten seconds, because the whole appeal of this lever is that it works in forty. **Address boundaries drawn as blast-radius boundaries.** Anything you might ever have to sacrifice should sit alone on its own address, and anything you cannot sacrifice should not share an address with a likely target. That costs addresses and it costs some proxy configuration, and it is the price of having a mitigation with a survivable blast radius. **A designated alternative for the shared addresses.** For the addresses that genuinely cannot be sacrificed, the coarse lever must be off the table, which means the runbook has to name what the on-call does instead. That is where a granular upstream rule earns its keep: matched on protocol, ports and packet characteristics rather than destination alone, it can shed the flood while leaving the address alive. It also has to exist in the transit contract *before* the incident, and that is a purchasing decision made in daylight, not at 03:00. ## Recovery is not symmetric Withdrawing the announcement restores routing in seconds, so the eleven services come back quickly once someone realises. But two clocks are longer than BGP. If anyone reacted by renumbering a service onto a new address, the recovery time is the cached DNS TTL of the old name, not convergence. And the attacker may still be there: withdrawing is a live re-test of the flood, so the sequence has to be planned as "withdraw, watch the link, be ready to re-announce", with a named person deciding, rather than a hopeful click. ## The answer an interviewer is listening for Say that the incident is evidence about your address architecture rather than about the responder. Then give a concrete remediation order: build the address-to-service map this week, split the shared front end so the sacrificeable services stand alone, cap what the on-call may announce at the longest prefix every upstream accepts, and negotiate a granular option for the addresses that must never go dark. That is a defensible plan with costs attached, which is what the question is really testing.
- How would you cap the damage the on-call can do without slowing them down?Set a maximum prefix length the runbook permits, chosen as the longest every upstream reliably accepts, and enforce it in the automation so a coarser announcement needs a named authoriser. Pair that with an address-to-service table the on-call reads in ten seconds. The aim is a priced decision at speed, not an approval queue.
- The flooded address is a shared front end that cannot be sacrificed. What is the plan instead?Either the address stops being shared, or you buy granularity. Splitting the sacrificeable services onto their own addresses restores the coarse lever's usefulness. Otherwise the runbook must point at a granular upstream rule that matches protocol, ports and packet characteristics rather than destination alone, which has to be negotiated into the transit arrangement in advance.
- Once the flood stops, what makes the recovery slower than the mitigation was?Withdrawing the announcement converges in seconds, so routing is not the constraint. The delays come from anything that reacted around it: a service renumbered onto a fresh address waits out the old name's cached DNS TTL, and clients or partners that pinned the old address need coordination. There is also a decision cost, because withdrawing re-tests whether the attacker is still there.
saying these in an interview costs you the question
- Blames the responder rather than the address plan
- Believes the announcement could have spared co-tenant services
- Assumes one IP means one application
- Forgets the address may be an outbound source too
- Treats withdrawal as instant recovery for everything