skip to content

When a router's best path to a prefix fails, how does EIGRP's feasible-successor failover differ from OSPF rerunning SPF, and when does EIGRP lose its advantage?

level: seniorimportance: should knowfreq 25%

answer

  1. detection is the same problem
  2. local switch versus area-wide recompute
  3. route stays passive
  4. no backup means queries
  5. query scope and stuck-in-active

basics

~20 s

With a feasible successor, an EIGRP router switches locally and the route stays passive; OSPF floods a new LSA and every router in the area reruns SPF. Without one, EIGRP must query its neighbours and can converge more slowly.

solid answer

~50 s

Both protocols first have to **detect** the failure — loss of carrier, hellos timing out, or BFD (RFC 5880). After that, an EIGRP router that has a **feasible successor** (a neighbour whose reported distance is strictly below its feasible distance) installs it immediately; the route stays **passive** and only neighbours learn the new metric. OSPF instead originates a new LSA, floods it through the area, and **every router in that area reruns SPF**; routers in other areas process changed summary-LSAs incrementally, and an ABR's area range can hide the change completely. EIGRP loses its edge when **no feasible successor exists**: the route goes **active**, queries diffuse until replies return, and in a large unbounded network that takes longer than an OSPF recompute — or ends in **stuck-in-active**. Summaries and stub-style query boundaries are what keep EIGRP's worst case short.

go deeper

for a junior

Recall that EIGRP may hold a ready loop-free backup called a feasible successor, while OSPF recomputes its shortest-path tree after flooding the change.

for a middle

Explain the feasibility condition, why the route stays passive with a feasible successor, and why OSPF's recomputation is confined to one area.

for a senior

Diagnose when EIGRP's advantage disappears: no feasible successor, unbounded queries, stuck-in-active; and show how summaries, query boundaries and BFD change the outcome.

for a principal

Judge whether predictable area-wide recomputation or topology-dependent local failover better suits the estate's failure patterns and the team's ability to design query boundaries.

## Detection comes first, and it is the same problem Convergence after a failure has two parts: **noticing** the failure and **computing** the new path. Neither protocol can do the second before the first. - **Link down:** if the interface loses carrier, both protocols react at once. - **Hello timeout:** when the link stays up but the neighbour dies (a failure behind a switch, a hung control plane), EIGRP waits for its **hold time** and OSPF for its **RouterDeadInterval**. The values are configurable; RFC 7868 describes a 5-second hello and a hold time of three hellos as defaults. - **BFD:** Bidirectional Forwarding Detection (RFC 5880, with RFC 5882 describing how routing protocols use it) can give both protocols sub-second detection, so detection speed does not separate them. The interesting difference is what happens next. ## EIGRP with a feasible successor: a local decision EIGRP's DUAL algorithm keeps, for every destination, a **feasible distance (FD)** and the **reported distance (RD)** each neighbour advertises. A neighbour whose RD is strictly lower than the FD satisfies the **feasibility condition** and is a **feasible successor**: its path provably does not loop back through this router. 1. The successor's link fails. 2. The router checks its topology table and finds a feasible successor. 3. It installs that neighbour as the new successor at once. The route stays **passive** — RFC 7868 says a router may keep a route passive when it knows other paths meeting the feasibility condition. 4. It sends an UPDATE so neighbours learn the new metric. No query, no wait for replies, no recomputation elsewhere. The forwarding change is bounded by local processing time, which is why EIGRP's failover is described as fast "without tuning". ## EIGRP without a feasible successor: a diffusing computation If no neighbour meets the feasibility condition, the router cannot prove any alternative is loop-free. The route goes **active** and the router sends **QUERY** packets to its neighbours (except, under split horizon, towards the failed successor). Each neighbour either replies from its own table or, if it too depends on the lost path, queries further. The route returns to passive only when every outstanding **REPLY** has arrived. - In a large network with no boundaries, queries can travel to every router. - RFC 7868 §4.4 handles slow answers: after half of the active-time interval (an implementation-set value) the router sends an **SIA-QUERY**, which must be answered with a REPLY or an **SIA-REPLY**. A neighbour that answers neither is deemed **stuck in active**, and DUAL either deletes the routes through it or resets the whole adjacency; which one is an implementation choice. - Unaffected routers reply at once, so the computation shrinks as fast as it grows; the risk is the routers that do not. ## OSPF: flood, then recompute in the area 1. A router attached to the failure originates a new **router-LSA** (and, if a router on a broadcast segment is lost, the DR updates the **network-LSA**). 2. The LSA is flooded reliably to every router in the **area**. 3. Every router in that area reruns the **SPF** calculation for it, because each holds an identical copy of the area's link-state database. 4. ABRs advertise changed **summary-LSAs** into other areas; routers there update incrementally (RFC 2328 §16.5) rather than rerunning intra-area SPF. If an ABR's **area address range** still covers the prefix, other areas see no change at all. How quickly SPF runs after the first LSA, and how runs are throttled during a flap, are implementation choices, not RFC 2328 constants. ## Side by side | Situation | EIGRP | OSPF | |---|---|---| | Backup known in advance | Local switch to the feasible successor | Flood plus SPF in the area | | No proven backup | Queries and replies, possibly network-wide | Flood plus SPF in the area | | Worst case | Stuck-in-active: routes dropped or adjacency reset | Repeated SPF during a flap, damped by throttling | | What bounds the work | Summaries and query boundaries | Area size and ABR ranges | ## Design implications - EIGRP's speed depends on **topology**: a feasible successor exists only where a second path's reported distance happens to be below the FD. Ring and hub-and-spoke designs often lack one for some prefixes. - Bound EIGRP queries with **summaries** at aggregation points and the implementation's **stub routing** feature at the edge. - Keep OSPF areas small enough that an SPF run is cheap, and summarise at ABRs so churn stays local. - Add BFD to whichever protocol you choose; detection is usually the larger share of outage time.

  • Why can a network have a perfectly good second path that is still not a feasible successor in EIGRP?
    The feasibility condition is conservative: the neighbour's reported distance must be strictly below this router's feasible distance. A neighbour whose own path is longer than ours may still route correctly, but EIGRP cannot prove it does not loop back through us, so it is not a feasible successor and the route goes active instead.
  • How does summarisation shorten EIGRP convergence for a lost prefix?
    A router that knows the lost prefix only through a summary has no specific route for it, so when queried it replies straight away instead of joining the search. Summaries at aggregation points therefore stop the diffusing computation from spreading past them, keeping the active period short and the stuck-in-active risk low.
  • Does BFD make OSPF and EIGRP converge equally fast?
    It equalises detection, not computation. After BFD reports the failure, EIGRP with a feasible successor switches locally, while OSPF still floods and reruns SPF in the area. With fast SPF scheduling the gap is small; without a feasible successor, EIGRP can be the slower one.

saying these in an interview costs you the question

  • OSPF routers also switch to a feasible successor without recalculating.
  • When EIGRP loses its successor, it always sends queries before switching.
  • A failure inside one OSPF area makes every router in all areas rerun SPF.
  • EIGRP always converges faster than OSPF, whatever the topology.
  • BFD can be used by OSPF but not by EIGRP.