skip to content

For a routed network whose application tolerates at most one second of loss on a link failure, how do you build an end-to-end convergence budget?

level: seniorimportance: should knowfreq 22%

answer

  1. outage is a sum of phases
  2. detect, tell, compute, install
  3. the slowest term dominates
  4. precomputed backups skip steps

basics

~20 s

Convergence time is the sum of failure detection, propagation to other routers, path computation and forwarding-table install. Give each phase a measured or assumed value, add them with a margin, and compare against the one-second limit.

solid answer

~40 s

Break the outage into phases that run one after another: detecting the failure, telling the other routers, recomputing paths, and installing the result into the forwarding table on every line card. Detection ranges from milliseconds for loss of light, to `Detect Mult × interval` with BFD (RFC 5880), to many seconds for protocol hold timers — RFC 4271 suggests a 90-second BGP HoldTime. Propagation depends on flooding or BGP update pacing; computation on SPF delay and run time; install time grows with the number of affected prefixes. Sum measured values, keep a margin, and attack the largest term. A precomputed backup — a remaining equal-cost path, a loop-free alternate or MPLS fast reroute — lets the detecting router switch locally, so the budget shrinks to detection plus local repair.

go deeper

for a junior

Recall that routing recovery has phases: the failure must be detected, announced, recomputed and installed before traffic flows again.

for a middle

Explain each phase's main driver: link or BFD detection, flooding or update pacing, SPF timing, and install time growing with prefixes.

for a senior

Build the budget with measured values, show where the time goes, and use precomputed backups and faster detection to bring the largest term down.

for a principal

Set the tolerance with application owners, decide which failures must be covered sub-second and at what cost, and require failure testing that verifies the budget.

## Why a budget "The network converges fast" is not a design statement. An application that drops sessions after one second of loss sets a hard number, and the network either fits inside it or does not. A **convergence budget** breaks the outage after a failure into phases, assigns each a value, and adds them up. It turns a vague worry into arithmetic you can check, and it shows which phase to spend effort on. ## The phases The interval between a link failing and traffic flowing again on a new path is roughly the sum of: 1. **Detection** — the router adjacent to the failure learns of it. Loss of light on a point-to-point fibre is noticed in milliseconds. BFD in asynchronous mode declares a session down after the remote system's `Detect Mult × the agreed interval` (RFC 5880). A routing protocol's own keepalives are far slower: RFC 4271 suggests a 90-second BGP HoldTime, and RFC 7938 notes that many implementations cannot set a hold timer below three seconds. 2. **Propagation** — the news reaches every router that must change its path: a link-state flood hop by hop, or BGP UPDATE and withdraw messages. BGP's `MinRouteAdvertisementIntervalTimer` has suggested defaults of 30 s on eBGP and 5 s on iBGP sessions (RFC 4271), though RFC 7938 observes that the initial withdrawals after an event are commonly not delayed by it. 3. **Computation** — each router reruns its path selection. For a link-state protocol that is an SPF delay plus the run itself. RFC 8405 standardises an SPF back-off algorithm, requires its delays to be configurable and suggests a 50 ms `INITIAL_SPF_DELAY` where the algorithm is on by default; the values actually used are an operator choice. 4. **Install** — the new next hops are written into the routing table, distributed to line cards and programmed into the forwarding hardware. RFC 8405 lists these steps — flooding, SPF wait, SPF computation, FIB distribution across line cards, FIB update — and the time grows with the number of prefixes that change. ## A worked budget Target: under 1 second, with margin. The values below are illustrative assumptions for one design; real figures come from measuring the platforms in a lab. | Phase | What sets it | Budget | |---|---|---| | Detection | BFD, Detect Mult 3 × 100 ms | 300 ms | | Propagation | flooding across about four hops | 40 ms | | Computation | SPF initial delay of 50 ms plus a 20 ms run | 70 ms | | Install | affected prefixes on all line cards | 200 ms | | **Total** | | **610 ms** | | **Margin to 1 s** | | **390 ms** | Replace BFD with a protocol hold timer measured in seconds and detection alone consumes the entire budget several times over. The budget makes that visible before the first outage does. ## Shrinking the budget - **Detect faster.** Rely on loss of light where links are point to point, and add BFD where failures can be silent. RFC 7938 argues sub-second convergence is achievable when sessions are torn down promptly on link failure and RIB and FIB updates are timely. - **Precompute the backup.** If the router already holds a usable alternative — another member of an equal-cost group, a loop-free alternate, a backup path in the BGP Loc-RIB, or an MPLS fast reroute detour or bypass, which RFC 4090 describes as redirecting traffic in "10s of milliseconds" — it can switch locally, skipping network-wide propagation and computation. The common 50 ms target is an operator convention, not an RFC requirement. - **Install less.** Fewer prefixes to reprogram, or forwarding structures that change one shared next hop instead of every prefix, cut install time. ## What the budget misses - **Micro-loops.** Routers install new paths at slightly different times, so for a moment two may point at each other; RFC 8405 describes this effect. - **The application's own timers.** A transport retransmission or a health check may add delay beyond the network's convergence, so the user-visible outage can exceed the routing sum. - **Recovery.** When the link comes back, traffic moves again; an unstable link that flaps repeatedly can cause more loss than the original failure, which is why implementations commonly hold a flapping link down for a while before trusting it again. Measure end to end with continuous probes during failure tests, and compare the result with the budget line by line.

  • Why does a precomputed backup path reduce the budget more than faster SPF does?
    Faster SPF only shortens one phase, while a precomputed backup lets the router that detects the failure switch traffic locally at once, skipping propagation and network-wide computation entirely. The outage shrinks to detection plus the local switch; the network then reconverges in the background with little or no further loss for the protected traffic.
  • Why can the measured outage be longer than the sum of the routing phases?
    Routers install new paths at different moments, so transient micro-loops can drop or loop packets after the budgeted install. On top of that, the application's own retransmission and health-check timers add delay before users see recovery, so the end-to-end outage includes time the routing budget never counted.

saying these in an interview costs you the question

  • Treating forwarding-table install as instant once the route is computed
  • Relying on routing-protocol keepalives alone to meet a sub-second target
  • Assuming every BGP update waits the full 30-second eBGP MRAI
  • Quoting 50 ms failover as an RFC requirement rather than an operator convention
  • Ignoring application timers when comparing network convergence with the tolerance