skip to content

When a link fails in an OSPF network, what steps make up the time until traffic takes a new path, and which usually dominates?

level: seniorimportance: must knowfreq 34%

answer

  1. detect, originate, flood, compute, install
  2. carrier loss versus silence
  3. RouterDeadInterval as a multiple of Hello
  4. MinLSInterval limits re-origination
  5. forwarding-table install grows with prefixes

basics

~20 s

OSPF convergence is detection, LSA origination, flooding, SPF and route install. Detection usually dominates: lost carrier is noticed at once, but a failure hidden behind a switch waits for the Dead interval, 40 s with RFC 2328's sample timers.

solid answer

~40 s

I split it into five steps. **Detection**: if the interface loses carrier, the lower layer signals it at once; if the failure is invisible locally — a switch in between, a hung neighbour — only the `RouterDeadInterval` expiring notices it, 40 s with RFC 2328's sample 10 s Hello and a multiple of 4. **Origination**: the router reoriginates its router-LSA, but never two instances of one LSA within `MinLSInterval` (5 s), plus any delay the implementation adds. **Flooding**: hop by hop, with retransmission after `RxmtInterval` if an update is lost. **SPF** after its scheduling delay, then **route install**, which grows with the number of changed prefixes. Detection dominates unless carrier drops; shorten it with direct links, lower timers or a dedicated detection protocol.

go deeper

for a junior

Recall that OSPF must first notice a failure, then tell other routers, then recompute; noticing is often the slow part.

for a middle

Explain the chain step by step and the difference between carrier loss, which is immediate, and the Dead interval, which waits for missed Hellos.

for a senior

Diagnose which step dominates in a given topology and pick the right lever for each, weighing aggressive timers against false adjacency loss on a busy control plane.

for a principal

Treat convergence as a budget across detection, flooding, computation and install, and decide which failures the design must survive within it.

## The five steps When a link inside an OSPF area fails, traffic keeps going into the dead link until every router on the path has done all of these: 1. **Detect** that the neighbour or interface is gone. 2. **Originate** a new router-LSA (and a new network-LSA if it is the Designated Router of an affected segment) that no longer lists the link. 3. **Flood** that LSA hop by hop across the area. 4. **Compute** a new shortest-path tree (SPF). 5. **Install** the changed routes into the forwarding table. **Convergence time** is the sum along the slowest chain of these steps. Knowing which one dominates tells you where tuning pays off. ## Detection: carrier loss versus silence There are two very different ways a router learns of a failure. - **Carrier loss.** On a direct point-to-point cable, the lower layer reports that the interface is down. RFC 2328 models this as the `InterfaceDown` event (and `LLDown` for a neighbour), which forces the interface or neighbour to Down at once. Detection is as fast as the lower layer reports it. - **Silence.** If the failure is not visible locally — the routers connect through a Layer 2 switch whose port stays up, or the far router's control plane hangs — the only signal is that Hellos stop arriving. Each received Hello restarts the neighbour's **Inactivity Timer**, whose length is `RouterDeadInterval`. Only when it fires does the neighbour go Down. RFC 2328 gives a *sample* HelloInterval of 10 s on a LAN and says the dead interval should be "some multiple" of it ("say 4"), so the familiar 40 s is that multiple, not a mandated value. Because the last Hello may have arrived just before the failure, detection by silence takes roughly 30-40 s with those values. Everything after detection is usually far shorter, which is why **detection dominates** whenever carrier does not drop. ## Origination and flooding - A router reoriginates an LSA when its contents change, but RFC 2328 forbids originating **two instances of the same LSA within `MinLSInterval` (5 s)**. A second change soon after the first is delayed, not dropped. - Implementations commonly add their own configurable delay before the first origination so that several simultaneous changes go into one LSA; that delay is an implementation choice. - Flooding is reliable: each router installs the new LSA and sends it out its other interfaces, and an unacknowledged update is retransmitted after `RxmtInterval` (sample value 5 s). Receivers also ignore a new instance arriving within `MinLSArrival` (1 s) of the previous one. - **Only one end has to report the failure.** During SPF a router uses a link between two routers only if *both* routers' LSAs list each other. As soon as either endpoint floods an LSA without the link, every router drops it from its tree. ## SPF and route install Each router schedules an SPF run after receiving the change; how that run is delayed and batched is a subject of the SPF computation itself. The final step, writing changed routes into the forwarding table, scales with the **number of prefixes** that changed, so in a large routing table install time can be a visible share of the total. ## A worked timeline | Step | Driven by | Typical lever | |---|---|---| | Detection, carrier loss | lower-layer signal | use direct links where possible | | Detection, silence | `RouterDeadInterval` | lower Hello/Dead, or a dedicated detection protocol | | Origination | change in LSA content, `MinLSInterval` | shorter implementation generation delay | | Flooding | hop count, link and CPU load | keep the area's diameter and load sane | | SPF | scheduling delay, area size | the SPF scheduling settings | | Install | number of changed prefixes | fewer prefixes, prioritised install | ## Shortening each step - Prefer topologies where a failure drops carrier on both ends. - Lower `HelloInterval` and `RouterDeadInterval` (they must match on a segment), accepting more Hello processing and a higher risk that a busy control plane misses Hellos and tears down a healthy adjacency. - For detection well below a second, use a dedicated fast-detection protocol such as BFD, which is its own subject. - Keep the area small enough that flooding, SPF and install stay quick.

  • Why not set the OSPF Hello interval to 1 s and the Dead interval to 3 s on every interface?
    The protocol allows it — both are whole-second fields in the Hello packet and must match on a segment — but it costs something. Every interface processes more Hellos, and a congested link or busy control plane that loses three Hellos tears down a healthy adjacency, which itself forces a reconvergence. RFC 4222, a Best Current Practice, recommends prioritising Hellos and acknowledgements for exactly this reason.
  • If only one end of a failed OSPF point-to-point link notices the failure, does the area still route around it?
    Yes. In SPF a router uses a link between two routers only if each router's LSA lists the other, the back-link check in RFC 2328 section 16.1. Once the detecting router floods a router-LSA without the link, every router, including the end that has not noticed, drops the link from its tree. That end still keeps the adjacency until its own Dead interval fires.

saying these in an interview costs you the question

  • OSPF convergence time is simply the 40-second Dead interval the RFC mandates.
  • Even when a direct interface loses carrier, OSPF waits for the Dead interval.
  • Both ends must detect a failure before any router stops using the link.
  • SPF computation is always the slowest step of OSPF convergence.
  • Lowering Hello and Dead timers is free apart from a little bandwidth.