skip to content

When an iBGP router receives an eBGP-learned route whose NEXT_HOP its IGP cannot reach, what happens, and how do next-hop-self and IGP advertisement fix it?

level: middleimportance: should knowfreq 33%

answer

  1. NEXT_HOP survives the iBGP hop
  2. points at the provider's address
  3. unresolvable means excluded
  4. rewrite at the border, or advertise the link

basics

~20 s

iBGP leaves NEXT_HOP pointing at the external neighbour, so an interior router without an IGP route to that address must exclude the route from selection. Fix it at the border with next-hop-self, or by carrying the external link's subnet in the IGP.

solid answer

~50 s

Over eBGP the provider sets `NEXT_HOP` to its own interface, say `198.51.100.1`. RFC 4271 section 5.1.3 says the border router SHOULD NOT change it when relaying to internal peers unless configured to, so interior routers receive a route whose next hop is an address on the border's external link. If the IGP does not carry that link, the next hop is **unresolvable**, and section 9.1.2 says the route MUST be excluded from selection: it sits in `Adj-RIB-In` and is never used. Two fixes: configure the border router to put its own loopback in `NEXT_HOP` toward internal peers (**next-hop-self**), or advertise the external link's subnet into the IGP. The first keeps external addresses out of the IGP; the second lets interior routers see the actual exit and drop routes as soon as the IGP removes a failed link.

go deeper

for a junior

Recall that a BGP route carries a NEXT_HOP and that a router can only use the route if it can reach that address.

for a middle

Walk the attribute across both hops: rewritten to the sender on eBGP, left alone on iBGP, then excluded as unresolvable; name both fixes and where each is configured.

for a senior

Diagnose it from symptoms: route received but not best, session healthy, nothing installed. Choose between the fixes on failure behaviour and IGP hygiene, and keep reflectors out of it.

for a principal

Set the AS-wide convention, next-hop-self everywhere or edge subnets in the IGP, by weighing convergence on edge failures against IGP size and boundary exposure.

## The setup AS 64500 has a border router **R1** (loopback `10.255.0.1`) with an eBGP session to its provider in AS 64501 over the link `198.51.100.0/30`: the provider is `198.51.100.1`, R1 is `198.51.100.2`. Interior router **R2** (loopback `10.255.0.2`) has an iBGP session with R1 between loopbacks, and the IGP carries the AS's internal links and loopbacks only. 1. The provider advertises `203.0.113.0/24` to R1. Over eBGP it sets `NEXT_HOP` to its own session address, `198.51.100.1` (RFC 4271 section 5.1.3, the default case). 2. R1 relays the route to R2 over iBGP. Section 5.1.3 rule 1: when sending to an internal peer a route it did not originate, the speaker "SHOULD NOT modify the NEXT_HOP attribute unless it has been explicitly configured to announce its own IP address as the NEXT_HOP". 3. R2 now holds a route whose `NEXT_HOP` is `198.51.100.1`. R2's routing table has no entry covering that address. ## What the receiver must do RFC 4271 section 9.1.2: if `NEXT_HOP` "depicts an address that is not resolvable", the route **MUST be excluded** from the Phase 2 decision. Such routes are removed from `Loc-RIB` but **SHOULD be kept in `Adj-RIBs-In`** in case they become resolvable. Symptoms: R2's BGP table shows the prefix as received but not best and not installed, R2 advertises nothing for it, and traffic falls to a default route or is dropped. The session is up and no error is raised, which is why this is a classic interview diagnosis. ## Fix 1: next-hop-self at the border R1 is configured to put its own address, normally its loopback `10.255.0.1`, in `NEXT_HOP` on routes it sends to internal peers. The RFC permits exactly this exception. R2 resolves `10.255.0.1` through the IGP, and the forwarding chain becomes: prefix, then R1's loopback, then the IGP path to R1. - External link subnets stay out of the IGP. - The setting belongs on the **border router**. A route reflector SHOULD NOT modify `NEXT_HOP` when reflecting (RFC 4456 section 10), so setting it there is the wrong place. - After a provider failure, interior routers keep forwarding to R1 until R1's BGP processes the loss and withdraws the route. ## Fix 2: carry the external link in the IGP The `198.51.100.0/30` link is advertised into the IGP (usually as a passive interface, so no IGP adjacency forms with the provider). R2 now resolves `198.51.100.1` directly. - If the link goes down, the IGP withdraws the subnet and every route using `198.51.100.1` becomes unresolvable as soon as the IGP converges, without waiting for BGP. - Interior routers can compare the IGP cost to each real exit. - The IGP now carries addresses that belong to the boundary, which some operators dislike. ## Confirming the diagnosis 1. Check that R2 received the prefix from R1: the session is up and the UPDATE arrived. 2. Read the route's `NEXT_HOP`: an address on the border's external link, not R1's loopback. 3. Look that address up in R2's routing table: nothing covers it. 4. Apply one of the two fixes and watch the route become best and reach the forwarding table. ## Comparing the fixes | | next-hop-self on the border | external link in the IGP | |---|---|---| | Interior routers' next hop | border loopback | provider's interface address | | IGP contents | internal only | internal plus edge subnets | | Reaction to the edge link failing | waits for BGP withdrawal | IGP removes the next hop | | Where configured | each border router, per internal peer | each border router's IGP | ## Where the eBGP side fits On eBGP the rewrite is the default: when R1 advertises its own prefixes to the provider, `NEXT_HOP` is the address R1 uses for that session, `198.51.100.2`. Exceptions are the **third-party next hop**, allowed when the peer shares a subnet with a better router, and **multihop eBGP**, where the speaker MAY propagate `NEXT_HOP` unchanged. The asymmetry between the two session types is deliberate: the eBGP neighbour must be told how to reach you, while internal routers can be told to rely on the IGP.

  • Why not set next-hop-self on a route reflector instead of on every border router?
    RFC 4456 section 10 says a reflector SHOULD NOT modify `NEXT_HOP` when reflecting, because changing it can create loops and pulls all traffic through the reflector, which often is not in the forwarding path at all. The border router is where the external next hop originates, so it is the one place the rewrite is both correct and harmless.
  • With next-hop-self in place, the provider link fails. How long do interior routers keep sending traffic toward R1?
    Until R1 notices the failure and sends a withdrawal over iBGP. If the interface goes down, that can be fast; if the link stays up but the provider stops responding, R1 waits for its hold timer to expire. Meanwhile R1's loopback stays reachable, so interior routers keep forwarding to it and R1 drops or redirects the traffic.

saying these in an interview costs you the question

  • A border router rewrites NEXT_HOP for iBGP peers by default, just as on eBGP
  • An unresolvable NEXT_HOP makes the receiving router tear down the iBGP session
  • Next-hop-self should be configured on the route reflector
  • Raising LOCAL_PREF fixes a route that is not installed
  • The route is dropped from the BGP table entirely and must be re-sent