skip to content

Why may a BGP speaker not pass a route learned over iBGP to another iBGP peer, and why does that force a full mesh?

level: middleimportance: must knowfreq 42%

answer

  1. nothing marks an internal hop
  2. AS_PATH loop check needs a change
  3. every speaker hears the border directly
  4. n(n-1)/2
  5. not the distance-vector rule

basics

~20 s

Inside one AS the AS_PATH is never changed, so BGP's loop check cannot catch a route circulating between internal peers. RFC 4271 therefore forbids relaying iBGP-learned routes to iBGP peers, so in base BGP every speaker must peer with every other one: n(n-1)/2 sessions.

solid answer

~50 s

BGP prevents loops with `AS_PATH`: each eBGP hop prepends an AS number, and a speaker rejects a route that already contains its own. On iBGP sessions `AS_PATH` is not modified, so a route relayed from internal peer to internal peer would carry no record of where it had been and could loop. RFC 4271 section 9.2 closes this by saying a speaker SHALL NOT re-distribute routes learned from an internal peer to other internal peers. The consequence is that each speaker can only learn an external route from the border router that learned it, so every BGP speaker in the AS needs a session with every other: a **full mesh** of `n(n-1)/2` sessions. With 30 routers that is 435 sessions and 29 per router. Route reflectors (RFC 4456) and confederations (RFC 5065) are the two standard ways out.

go deeper

for a junior

Remember that a BGP router does not pass routes heard from one iBGP neighbour on to another iBGP neighbour, and that this is why iBGP is configured as a full mesh.

for a middle

Derive the rule from the AS_PATH loop check, compute n(n-1)/2 for a given size, and separate the iBGP rule clearly from distance-vector split horizon.

for a senior

Show the operational cost: each new router touches every other router, a missing session silently strands prefixes, and at some size you must move to reflectors or a confederation.

for a principal

Decide when the mesh stops being acceptable for your AS, weighing session count and change risk against the path-diversity and migration costs of the alternatives.

## The rule RFC 4271 section 9.2: "When a BGP speaker receives an UPDATE message from an internal peer, the receiving BGP speaker SHALL NOT re-distribute the routing information contained in that UPDATE message to other internal peers (unless the speaker acts as a BGP Route Reflector)." Routes learned over iBGP may still be sent to **external** peers; it is only the internal-to-internal relay that is forbidden. Operators often call this **iBGP split horizon**. The name is borrowed, and the borrowing misleads: distance-vector split horizon withholds a route only from the interface it was learned on, while the iBGP rule withholds it from **every** internal peer, including ones that never saw it. ## Why the rule is needed BGP's loop prevention is the **`AS_PATH` check**: a speaker that finds its own AS number in a received `AS_PATH` excludes the route (RFC 4271 section 9.1.2). That works because every eBGP hop prepends an AS number. Inside an AS: - an iBGP speaker SHALL NOT modify `AS_PATH` when advertising to an internal peer; - a route originated inside the AS is sent to internal peers with an **empty** `AS_PATH`; - so every internal copy of a route looks the same however many internal hops it has made. If internal relaying were allowed, router A could pass a route to B, B to C and C back to A, and nothing in the attributes would reveal it. When the border router later withdraws the route, stale copies could keep circulating. Rather than add an in-AS marker to every route, base BGP forbids the relay. ## Why that forces a full mesh With relaying forbidden, an external route can reach an internal speaker in exactly one way: directly from a border router that learned it over eBGP. Since any router can be a border router for some prefix, every BGP speaker must hold an iBGP session with every other BGP speaker. RFC 4271 section 3 states the assumption, and RFC 5065's introduction states the cost: "for n BGP speakers within an AS, n*(n-1)/2 unique Internal BGP (IBGP) sessions are required." The mesh is **logical**, not physical: the sessions are TCP connections, usually between loopback addresses that the IGP makes reachable, so a session can cross many routers. ## The arithmetic at scale | BGP speakers (n) | sessions n(n-1)/2 | sessions per router (n-1) | |---|---|---| | 5 | 10 | 4 | | 10 | 45 | 9 | | 30 | 435 | 29 | | 31 | 465 | 30 | Growth is quadratic. Adding the thirty-first router means configuring **30 new sessions**, one on every existing router, and each session carries its own copy of the routes. The per-router cost is memory and CPU for every peer's view; the operational cost is that forgetting one session silently leaves a router without some external routes. ## The two standard exits 1. **Route reflection** (RFC 4456): a reflector is allowed to relay iBGP-learned routes to its clients and adds `ORIGINATOR_ID` and `CLUSTER_LIST` to restore loop detection. With two reflectors and 28 clients peering with both, 30 routers need 28 x 2 + 1 = **57** sessions instead of 435. 2. **Confederations** (RFC 5065): the AS is split into member-ASes with a full mesh inside each and eBGP-like sessions between them; the member-AS numbers recorded in `AS_CONFED_SEQUENCE` restore loop detection. Both work by putting back the information whose absence made the rule necessary. ## What the rule does not say - It does not stop a border router advertising iBGP-learned routes to its eBGP neighbours; policy decides that. - It does not apply to eBGP-learned routes, which a speaker sends to all internal peers. - It does not require a physical full mesh of links. - It does not make routers that do not run BGP able to forward to external destinations; they still need a way to forward that traffic, which is a separate design question.

  • In an AS of 30 fully meshed iBGP speakers, one session between two interior routers fails. What breaks?
    Those two routers stop hearing each other's eBGP-learned routes, and nobody else will relay them, because the no-relay rule applies everywhere. If either one is the only exit for some prefix, the other loses that prefix, while the IGP and every other session look healthy. That silent partial loss is why mesh completeness needs monitoring.
  • Is the iBGP rule the same mechanism as split horizon in a distance-vector protocol?
    No. Distance-vector split horizon withholds a route only from the interface it was learned on, to stop two neighbours counting to infinity. The iBGP rule withholds an iBGP-learned route from all internal peers, because `AS_PATH` carries no in-AS marker. They share a nickname, not a mechanism.

saying these in an interview costs you the question

  • The iBGP rule is the same as distance-vector split horizon on an interface
  • A full iBGP mesh means every router needs a physical link to every other
  • Thirty iBGP routers need 870 sessions, one in each direction
  • iBGP-learned routes may not be sent to eBGP peers either
  • iBGP prepends the local AS number, so loops are caught anyway
  • The IGP floods external BGP prefixes, so the mesh is only a backup