skip to content

Why does a BGP blackhole announced to only one of your two transits leave the flood arriving?

level: middleimportance: should knowfreq 44%

answer

  1. an announcement is per session, not per internet
  2. the other path never heard the request
  3. NO_EXPORT keeps it deliberately local
  4. relief is proportional, the outage is total
  5. per upstream: community, prefix length, filter exception

basics

~20 s

Because the request is per neighbour. The transit you asked discards the traffic; the second transit never saw the tagged route, still has a normal path to your covering prefix, and keeps delivering its share of the flood onto that link.

solid answer

~50 s

A blackhole is not a global fact about an address, it is a request made on one eBGP session. RFC 7999 even recommends carrying `NO_EXPORT` so it stays inside the AS you asked. Transit B's best path to your `203.0.113.0/24` is unchanged, so whatever share of the flood entered through B keeps arriving. You get relief in proportion to the traffic that happened to arrive via A, while paying the *full* cost immediately, because everyone reaching you through A is already cut off from that address. Partial coverage buys part of the benefit at all of the price. Getting it right means the arrangement exists on every upstream before the incident: their accepted community value, the longest prefix they will honour, and a tested announcement. Then verify with the upstream's route view and your own interface counters that the drop actually happened.

go deeper

for a junior

Remember that you announce to a neighbour, not to the internet, so anything you ask one provider to do has no effect on the path through another. Coverage means asking every upstream.

for a middle

Explain per-session announcement, the recommended NO_EXPORT scoping, and why a partial blackhole gives partial relief at full cost. Know that the upstream's inbound filter normally rejects a more specific and that accepting yours is a pre-arranged exception.

for a senior

Be able to prove the denial rather than assert it: upstream route view, your own interface counters, and external reachability per provider. Say plainly what fraction of your flood each upstream carries, because that fraction is your real mitigation capability.

for a principal

Treat per-upstream blackhole arrangements as a procurement and testing obligation, not a network task, and record the honest coverage number in the runbook instead of the word blackhole.

## BGP tells each neighbour something separately The mental model that produces this mistake is that a blackhole "marks an address as dead on the internet". It does not. A BGP speaker sends an announcement over a session to a specific neighbour, and each neighbour applies its own policy to what arrives. When you tag a more specific host route with the blackhole community and send it to transit A, you have asked **A** to install a discard route. Transit B was not asked, has no such route, and continues to resolve your address through your covering `/24` exactly as before. Its share of the flood keeps landing on your B-facing link. RFC 7999 reinforces this locality deliberately: it recommends the blackhole announcement carry `NO_EXPORT` (or `NO_ADVERTISE`) so that the discard does not leak onward into the wider internet from the AS you asked. That is a safety property, not a bug, but it means coverage is something you assemble one upstream at a time. ## The economics of a partial blackhole This is the part interviewers are actually probing. Attack traffic is destination routed, and the path it takes is chosen by each source's own network, not by you. So the flood arrives split, in whatever proportion the internet happens to route it. Blackholing at A removes A's share and does nothing to B's. If sixty percent came via A, your link relief is roughly sixty percent, and if your uplinks are still saturated by the remaining forty, you have gained nothing operationally. Meanwhile the cost landed in full the moment A honoured the request. Every user whose packets reached you via A now cannot reach that address at all. So a half-deployed blackhole is the worst cell in the table: **the complete outage, part of the relief**. If you are going to finish the denial for an address, do it everywhere at once, or do not do it. ## What has to be arranged in advance, per upstream Each transit has its own arrangements, and none of them are discoverable during an attack: | What differs per upstream | Why it bites you at 03:00 | |---|---| | The community value they honour | Some accept the well-known `65535:666`, others require a provider-specific community | | Longest prefix they accept | Their inbound filter normally rejects more specifics from you; the blackhole path is an explicit exception, often capped | | Whether they propagate it further | Some pass the discard to their own upstreams, some keep it local by design | | Automation or a phone call | A ticket-driven path can take longer than the attack lasts | The pattern to state in an interview: you announce a `/32` that your provider would normally filter, so this only works because someone made an exception for you in writing, on every session, and tested it on a quiet Tuesday with a harmless address. ## Prove the boundary actually denies An engineer who says "I announced it" has not answered the question of whether the packets stopped. Three checks, cheap and independent: - **Upstream route view.** Confirm each transit is actually carrying your more specific with the community intact, rather than silently filtering it. - **Your own interface counters.** Inbound bits per second on the link facing that transit should fall, and that is the number you actually care about. - **Reachability from outside.** The blackholed address should be unreachable from a source that transits each provider. Unreachable is the success criterion here, which is a strange thing to celebrate and worth saying plainly. ## The multi-homing twist worth knowing Because the covering prefix is still announced everywhere, the blackhole never affects your other addresses on either path. That is the design working: you have deaggregated one host out of your allocation for the purpose of killing it, and left the rest of the space routed normally. The risk in the other direction is an operator who reaches for the covering prefix instead, either because the automation defaulted there or because an upstream would not accept a longer one. That announcement is not a partial blackhole, it is a total one, and it takes every address you own off the internet on that path. ## What this means for the runbook The practical consequence is that "blackhole" is not one action but *n* actions, one per upstream, and they must fire together and be verified together. If your automation can only reach one provider, then your real mitigation capability is a fraction of what the runbook claims, and the honest thing to record is that number rather than the word.

  • Does blackholing at one transit push the attack traffic onto the other link?
    No, and claiming it does is a common error. Attack packets are destination routed, and the path each source takes is chosen by that source's network. Blackholing at A discards A's share where it arrives; it does not cause those packets to re-route to B. You lose A's portion of the flood and keep B's exactly as it was.
  • How would you verify that the discard is actually in effect?
    Check the upstream's route view to confirm they carry your more specific with the community intact rather than filtering it, watch the inbound bits per second on the link facing that transit fall, and test reachability to the blackholed address from outside via each provider. Unreachable from everywhere is the success criterion.
  • Why might an upstream reject the /32 you announce during an incident?
    Because their normal inbound policy filters more specifics from customers, to stop deaggregation and mis-announcement. Accepting a host route from you is a deliberate exception in their prefix filter, usually with a maximum length and often tied to a specific community. If it was never arranged and tested, the announcement is silently dropped and you spend the outage on the phone.

saying these in an interview costs you the question

  • Thinks a blackhole is a global property of the address
  • Claims the flood re-routes onto the surviving transit
  • Never verifies the upstream actually installed the route
  • Assumes every provider uses the same community value
  • Reaches for the covering prefix when a /32 is refused

context