For a routed network whose application tolerates at most one second of loss on a link failure, how do you build an end-to-end convergence budget?
answer
- outage is a sum of phases
- detect, tell, compute, install
- the slowest term dominates
- precomputed backups skip steps
basics
~20 sConvergence time is the sum of failure detection, propagation to other routers, path computation and forwarding-table install. Give each phase a measured or assumed value, add them with a margin, and compare against the one-second limit.
solid answer
~40 sBreak the outage into phases that run one after another: detecting the failure, telling the other routers, recomputing paths, and installing the result into the forwarding table on every line card. Detection ranges from milliseconds for loss of light, to `Detect Mult × interval` with BFD (RFC 5880), to many seconds for protocol hold timers — RFC 4271 suggests a 90-second BGP HoldTime. Propagation depends on flooding or BGP update pacing; computation on SPF delay and run time; install time grows with the number of affected prefixes. Sum measured values, keep a margin, and attack the largest term. A precomputed backup — a remaining equal-cost path, a loop-free alternate or MPLS fast reroute — lets the detecting router switch locally, so the budget shrinks to detection plus local repair.
go deeper
Recall that routing recovery has phases: the failure must be detected, announced, recomputed and installed before traffic flows again.
Explain each phase's main driver: link or BFD detection, flooding or update pacing, SPF timing, and install time growing with prefixes.
Build the budget with measured values, show where the time goes, and use precomputed backups and faster detection to bring the largest term down.
Set the tolerance with application owners, decide which failures must be covered sub-second and at what cost, and require failure testing that verifies the budget.
## Why a budget "The network converges fast" is not a design statement. An application that drops sessions after one second of loss sets a hard number, and the network either fits inside it or does not. A **convergence budget** breaks the outage after a failure into phases, assigns each a value, and adds them up. It turns a vague worry into arithmetic you can check, and it shows which phase to spend effort on. ## The phases The interval between a link failing and traffic flowing again on a new path is roughly the sum of: 1. **Detection** — the router adjacent to the failure learns of it. Loss of light on a point-to-point fibre is noticed in milliseconds. BFD in asynchronous mode declares a session down after the remote system's `Detect Mult × the agreed interval` (RFC 5880). A routing protocol's own keepalives are far slower: RFC 4271 suggests a 90-second BGP HoldTime, and RFC 7938 notes that many implementations cannot set a hold timer below three seconds. 2. **Propagation** — the news reaches every router that must change its path: a link-state flood hop by hop, or BGP UPDATE and withdraw messages. BGP's `MinRouteAdvertisementIntervalTimer` has suggested defaults of 30 s on eBGP and 5 s on iBGP sessions (RFC 4271), though RFC 7938 observes that the initial withdrawals after an event are commonly not delayed by it. 3. **Computation** — each router reruns its path selection. For a link-state protocol that is an SPF delay plus the run itself. RFC 8405 standardises an SPF back-off algorithm, requires its delays to be configurable and suggests a 50 ms `INITIAL_SPF_DELAY` where the algorithm is on by default; the values actually used are an operator choice. 4. **Install** — the new next hops are written into the routing table, distributed to line cards and programmed into the forwarding hardware. RFC 8405 lists these steps — flooding, SPF wait, SPF computation, FIB distribution across line cards, FIB update — and the time grows with the number of prefixes that change. ## A worked budget Target: under 1 second, with margin. The values below are illustrative assumptions for one design; real figures come from measuring the platforms in a lab. | Phase | What sets it | Budget | |---|---|---| | Detection | BFD, Detect Mult 3 × 100 ms | 300 ms | | Propagation | flooding across about four hops | 40 ms | | Computation | SPF initial delay of 50 ms plus a 20 ms run | 70 ms | | Install | affected prefixes on all line cards | 200 ms | | **Total** | | **610 ms** | | **Margin to 1 s** | | **390 ms** | Replace BFD with a protocol hold timer measured in seconds and detection alone consumes the entire budget several times over. The budget makes that visible before the first outage does. ## Shrinking the budget - **Detect faster.** Rely on loss of light where links are point to point, and add BFD where failures can be silent. RFC 7938 argues sub-second convergence is achievable when sessions are torn down promptly on link failure and RIB and FIB updates are timely. - **Precompute the backup.** If the router already holds a usable alternative — another member of an equal-cost group, a loop-free alternate, a backup path in the BGP Loc-RIB, or an MPLS fast reroute detour or bypass, which RFC 4090 describes as redirecting traffic in "10s of milliseconds" — it can switch locally, skipping network-wide propagation and computation. The common 50 ms target is an operator convention, not an RFC requirement. - **Install less.** Fewer prefixes to reprogram, or forwarding structures that change one shared next hop instead of every prefix, cut install time. ## What the budget misses - **Micro-loops.** Routers install new paths at slightly different times, so for a moment two may point at each other; RFC 8405 describes this effect. - **The application's own timers.** A transport retransmission or a health check may add delay beyond the network's convergence, so the user-visible outage can exceed the routing sum. - **Recovery.** When the link comes back, traffic moves again; an unstable link that flaps repeatedly can cause more loss than the original failure, which is why implementations commonly hold a flapping link down for a while before trusting it again. Measure end to end with continuous probes during failure tests, and compare the result with the budget line by line.
- Why does a precomputed backup path reduce the budget more than faster SPF does?Faster SPF only shortens one phase, while a precomputed backup lets the router that detects the failure switch traffic locally at once, skipping propagation and network-wide computation entirely. The outage shrinks to detection plus the local switch; the network then reconverges in the background with little or no further loss for the protected traffic.
- Why can the measured outage be longer than the sum of the routing phases?Routers install new paths at different moments, so transient micro-loops can drop or loop packets after the budgeted install. On top of that, the application's own retransmission and health-check timers add delay before users see recovery, so the end-to-end outage includes time the routing budget never counted.
saying these in an interview costs you the question
- Treating forwarding-table install as instant once the route is computed
- Relying on routing-protocol keepalives alone to meet a sub-second target
- Assuming every BGP update waits the full 30-second eBGP MRAI
- Quoting 50 ms failover as an RFC requirement rather than an operator convention
- Ignoring application timers when comparing network convergence with the tolerance