skip to content

When an AS of 30 iBGP routers moves to BGP route reflectors, where does a reflector send each route, and what stops reflection loops?

level: seniorimportance: should knowfreq 24%

answer

  1. clients and non-clients
  2. from a client: everyone
  3. from a non-client: clients only
  4. two attributes, types 9 and 10
  5. only the best path is reflected

basics

~20 s

A route reflector (RFC 4456) relays a client's route to all clients and non-clients, and a non-client's route to clients only. ORIGINATOR_ID drops a route returning to its originator; CLUSTER_LIST drops one returning to a cluster that already reflected it.

solid answer

~50 s

RFC 4456 lets a **route reflector** break the iBGP no-relay rule for its **clients**; the reflector and its clients form a **cluster**. A route from a client is reflected to every other client and to every non-client; a route from a **non-client** goes only to clients, which is why non-clients must still be fully meshed. Clients need no special support. Two optional non-transitive attributes replace the missing `AS_PATH` evidence: `ORIGINATOR_ID` (type 9), the BGP Identifier of the router that introduced the route into the AS, so that router ignores the route if it comes back; and `CLUSTER_LIST` (type 10), to which each reflector prepends its `CLUSTER_ID`, ignoring routes that already contain it. With two reflectors and 28 clients, 30 routers need 57 sessions instead of 435. The price: a reflector advertises only its own best path, so clients can lose path diversity.

go deeper

for a junior

Know that a route reflector lets iBGP routers peer with one or two central routers instead of with every other router.

for a middle

State the three reflection rules by source of the route, explain why non-clients stay meshed, and compute the session count before and after.

for a senior

Explain ORIGINATOR_ID and CLUSTER_LIST precisely, choose shared or distinct cluster IDs for redundant reflectors, and diagnose suboptimal exits caused by hidden paths.

for a principal

Design the reflection hierarchy: where reflectors sit relative to the topology, how many per cluster, and whether multipath advertisement is worth its memory to recover path diversity.

## What a reflector is allowed to do Base BGP forbids relaying a route learned from one internal peer to another. **Route reflection** (RFC 4456, which obsoletes RFCs 2796 and 1966) designates some speakers as **route reflectors (RRs)** that may do exactly that. Each RR divides its internal peers into two groups: - **Clients**: peers the RR reflects for. An RR and its clients form a **cluster**. - **Non-clients**: ordinary iBGP peers, typically other RRs. RFC 4456: "The Non-Client peer must be fully meshed but the Client peers need not be." Clients run ordinary iBGP and do not know they are clients; RFC 4456 lets *conventional* speakers sit in either group, which is what makes gradual migration possible. ## The reflection rules After selecting its best path, an RR sends it according to where it came from (RFC 4456 section 6): | Best path learned from | Reflected to | |---|---| | a client | all other clients and all non-clients | | a non-client | all clients | | an eBGP peer | all clients and non-clients (normal BGP) | The non-client row is the reason non-clients still form a mesh: a non-client's route is not reflected to other non-clients. ## Applying it to 30 routers AS 64500 has 30 BGP speakers. A full mesh needs 30 x 29 / 2 = **435** sessions. Make two of them RRs in one cluster, the other **28** clients of both, and peer the RRs as non-clients: 1. client sessions: 28 x 2 = 56; 2. RR-to-RR: 1; 3. total: **57**, and adding a router means two new sessions instead of thirty. Two RRs rather than one because a single RR is a single point of failure for route distribution. ## What prevents loops Reflection removes the guarantee the no-relay rule gave, and `AS_PATH` still records nothing inside the AS, so RFC 4456 section 8 adds two optional non-transitive attributes: - **`ORIGINATOR_ID`** (type code 9, 4 bytes): the BGP Identifier of the route's originator in the local AS, created by the RR that first reflects the route; a speaker SHOULD NOT create one if one already exists. A router SHOULD ignore a route carrying its own BGP Identifier as `ORIGINATOR_ID`. - **`CLUSTER_LIST`** (type code 10): the sequence of `CLUSTER_ID` values the route has passed. An RR MUST prepend its own `CLUSTER_ID` when reflecting and SHOULD ignore a route whose list already contains it. A cluster with one RR is identified by that RR's BGP Identifier. Redundant RRs serving the same clients can be configured with a shared 4-byte `CLUSTER_ID` so that "an RR can discard routes from other RRs in the same cluster". Giving each RR its own ID instead means each keeps the other's reflected copies too: more state, and an extra copy if one RR loses a client session. RFC 4456 section 10 also says an RR SHOULD NOT modify `NEXT_HOP`, `AS_PATH`, `LOCAL_PREF` or `MED` when reflecting, because rewriting them could create loops. ## The cost: hidden paths An RR advertises only **its own best path** for each prefix. Consequences: - Clients choose among what the RR chose, so a client may exit through a border router that is near the RR but far from the client. RFC 4456 section 11 warns that route selection may differ from a full mesh and advises placing reflectors congruent with the physical topology. - Alternate exits are invisible to clients, so after a failure they wait for the RR to select and send a replacement. - Advertising more than one path per prefix needs an extension such as RFC 7911, which adds a Path Identifier so multiple paths for one prefix can coexist. RFC 5065 also notes that reflectors, like confederations, can produce persistent oscillation with some MED and tie-breaking combinations, so MED policy needs care in reflected designs.

  • Why must non-clients still be fully meshed if reflectors exist?
    Because a route learned from a non-client is reflected only to clients. If RR1 and RR2 were not peered, a route arriving at RR1 from a non-client would never reach RR2 or its clients. In the usual design the non-clients are the reflectors themselves, so the mesh is small: two reflectors need one session between them.
  • Two reflectors in one cluster share a CLUSTER_ID. A client's session to RR1 fails. Do RR1's other clients lose that client's routes?
    No, as long as the client's session to RR2 is up. RR2 reflects the routes to every client, including RR1's. RR1 discards RR2's copy because its own `CLUSTER_ID` is in the `CLUSTER_LIST`, but RR1's clients still hear the routes from RR2, since every client peers with both reflectors.

saying these in an interview costs you the question

  • A route reflector rewrites NEXT_HOP to itself so traffic flows through it
  • Clients must support ORIGINATOR_ID and CLUSTER_LIST before reflection can start
  • A reflector passes a non-client's route on to the other non-clients
  • AS_PATH prevents reflection loops because the reflector prepends its AS
  • Clients see every path the reflector learned for a prefix
  • ORIGINATOR_ID and CLUSTER_LIST are well-known mandatory attributes