skip to content

Why do OSPF routers delay and throttle SPF runs, and how does an initial delay with back-off batch a flapping link's updates?

level: seniorimportance: should knowfreq 24%

answer

  1. one failure, many LSAs
  2. CPU, churn and micro-loops
  3. short first wait, longer under churn
  4. a running timer absorbs new events
  5. QUIET, SHORT_WAIT, LONG_WAIT

basics

~20 s

One failure produces several LSAs and a flapping link a stream of them. SPF throttling waits briefly so related changes go into one run, then backs off while churn continues. RFC 8405 standardises one such back-off for link-state IGPs.

solid answer

~40 s

Any changed LSA can call for an SPF run, one failure produces several LSAs, and a flapping link produces a stream of them. Running SPF for each one burns CPU and triggers churn downstream. It also has routers computing on different views at different moments, which causes **micro-loops**. Throttling waits a short **initial delay** so that related LSAs arrive together, then **backs off** while changes keep coming, so a burst costs one run. RFC 2328 defines no SPF timer, and implementations each had their own until **RFC 8405** standardised a three-state machine: `QUIET`, `SHORT_WAIT` and `LONG_WAIT`. If an implementation enables it by default, RFC 8405 says the delays SHOULD default to 50 ms, 200 ms and 5000 ms. The same values should be used across the whole area.

go deeper

for a junior

Recall that routers wait a moment before recomputing routes after a change, so several related changes are handled in one calculation instead of many.

for a middle

Explain initial delay versus back-off, and why a running timer that absorbs new events is what batches a flapping link into fewer SPF runs.

for a senior

Walk a flap through RFC 8405's states and timers, and argue the trade-off between convergence time, CPU and micro-loops when you set the delays.

for a principal

Treat SPF delay as one term in a convergence budget and as an area-wide policy: uniform values, a migration plan, and how much instability is worth absorbing.

## The problem: one failure, many LSAs Any change in an OSPF area's link-state database can call for a new SPF run, and one failure rarely produces just one change. When a point-to-point link fails, the routers at both ends each originate a new router-LSA. When a router fails, every one of its neighbours does. A link that flaps down and up repeats the whole set each time. Running SPF the moment each LSA arrives has three costs: - **CPU and churn.** Every run is followed by everything that reacts to route changes: new summary-LSAs at area border routers, BGP re-evaluating its next hops, label distribution, FIB updates on the line cards. - **Work done on half a picture.** The first LSA of a node failure describes only part of it. A run on that view installs an intermediate state that the next LSA immediately changes. - **Routers out of step.** Routers receive LSAs at different moments. If each computes at a different time, some forward on the new topology while others still use the old one, and packets can bounce between them. These are **micro-loops**. **SPF throttling** accepts a small, deliberate delay in exchange for fewer and better-informed runs. It waits briefly after a change so that related LSAs arrive, and waits longer while changes keep coming. ## What RFC 2328 does and does not say RFC 2328 defines **no SPF delay or hold timer**. Its rate limits apply to LSAs: `MinLSInterval` (5 seconds) between originations of one LSA, and `MinLSArrival` (1 second) between accepted instances of one LSA during flooding. Those limit how fast the database can change. They do not schedule the computation. Scheduling SPF was left to implementations, and many adopted an exponential back-off: a start delay, then a wait that grows (often doubling) with each run during churn, capped at a maximum. The parameters, the defaults and the growth rule differed between implementations. ## RFC 8405: one standard back-off RFC 8405 (Standards Track, 2018) standardises a single SPF back-off algorithm for link-state IGPs (IS-IS, OSPF and OSPFv3), so that routers from different implementations delay SPF by the same amount. It is a three-state machine: | State | Entered when | Delay used for a new run | |---|---|---| | `QUIET` | no IGP event for `HOLDDOWN_INTERVAL` | `INITIAL_SPF_DELAY` | | `SHORT_WAIT` | an event arrives in `QUIET` | `SHORT_SPF_DELAY` | | `LONG_WAIT` | `TIME_TO_LEARN_INTERVAL` passes in `SHORT_WAIT` | `LONG_SPF_DELAY` | Three timers drive it: - **`SPF_TIMER`** runs the computation when it expires. An event starts it only if it is **not already running**, and that is what batches events. - **`LEARN_TIMER`** starts on the first event in `QUIET` and moves the machine to `LONG_WAIT` when it expires. - **`HOLDDOWN_TIMER`** restarts on every event and returns the machine to `QUIET` once events have stopped for `HOLDDOWN_INTERVAL`. If an implementation enables the algorithm by default, RFC 8405 says it SHOULD use these defaults: `INITIAL_SPF_DELAY` 50 ms, `SHORT_SPF_DELAY` 200 ms, `LONG_SPF_DELAY` 5000 ms, `TIME_TO_LEARN_INTERVAL` 500 ms and `HOLDDOWN_INTERVAL` 10000 ms. `HOLDDOWN_INTERVAL` must be longer than `TIME_TO_LEARN_INTERVAL`, and every value must be configurable. ## Batching a flapping link Here is a link flapping under the RFC 8405 defaults. Times are counted from the first event, and an event means a change to the database. | Time | What happens | State | Effect | |---|---|---|---| | 0 ms | link down: event E1 | `QUIET` to `SHORT_WAIT` | SPF in 50 ms; learn 500 ms; holddown 10000 ms | | 50 ms | `SPF_TIMER` expires | `SHORT_WAIT` | run 1 | | 200 ms | link up: E2 | `SHORT_WAIT` | holddown to 10200 ms; SPF in 200 ms | | 400 ms | `SPF_TIMER` expires | `SHORT_WAIT` | run 2 | | 500 ms | `LEARN_TIMER` expires | to `LONG_WAIT` | no run | | 900 ms | link down: E3 | `LONG_WAIT` | holddown to 10900 ms; SPF in 5000 ms | | 1400 ms | link up: E4 | `LONG_WAIT` | holddown to 11400 ms; SPF timer already running | | 5900 ms | `SPF_TIMER` expires | `LONG_WAIT` | run 3, covering E3 and E4 | | 11400 ms | `HOLDDOWN_TIMER` expires | to `QUIET` | the next event gets 50 ms again | The first change converges almost at once. Once the link has shown itself unstable, E3 and E4 cost one run instead of two. The router goes back to reacting at full speed only after ten quiet seconds. ## Choosing the values - **Short delays** converge fastest after a single failure. Under churn they spend CPU, and they may compute before all of a node failure's LSAs have arrived. - **Long delays** protect the control plane, but traffic keeps following a dead path for longer. The SPF wait adds directly to convergence time. - **Uniformity matters more than the exact numbers.** RFC 8405 recommends the same values on every router in an area, or better in the whole domain, and moving every router to the algorithm at about the same time, because mismatched delays raise the chance and duration of micro-loops. The algorithm reduces micro-loops; it does not eliminate them. - RFC 8405 also notes that `LONG_SPF_DELAY` helps contain an attacker who can generate many IGP events. Throttling does not slow down failure detection. Hello and Dead timers, or a dedicated detection protocol, find the failure. The SPF delay starts only after the database has changed.

  • How does SPF throttling differ from OSPF's MinLSInterval and MinLSArrival?
    MinLSInterval (5 seconds) limits how often a router originates a new instance of one LSA, and MinLSArrival (1 second) limits how often a new instance of one LSA is accepted during flooding. Both are RFC 2328 constants that limit how fast the database changes. SPF throttling schedules the route computation over all changes, and RFC 2328 does not define it.
  • What does an initial SPF delay of zero buy, and what does it cost?
    A single link failure is computed immediately, the fastest possible reaction. But a router failure produces several LSAs. A zero-delay run sees only the first, so another run follows as soon as the rest arrive, which means more CPU and more intermediate states. RFC 8405 gives 0 ms as an example value and recommends 50 ms as the default.
  • Does throttling SPF make OSPF slower to detect a failure?
    No. Hello and Dead timers, or a dedicated fast-detection protocol, detect the failure, and the adjacent routers then originate new LSAs. The SPF delay starts only after the database changes. It adds to convergence time as one term in the budget, after detection, origination and flooding.

saying these in an interview costs you the question

  • RFC 2328 defines OSPF's SPF delay and hold timers.
  • Throttling SPF makes OSPF routers detect link failures more slowly.
  • With throttling, a burst of LSAs still gets one SPF run per LSA, just later.
  • SPF back-off values only matter locally, so mixing them across an area is harmless.
  • MinLSArrival is the OSPF timer that spaces out SPF runs.