skip to content

Why do many RIP implementations add a holddown timer that RFC 2453 never defines, and what does holddown cost when a real backup path exists?

level: seniorimportance: should knowfreq 16%

answer

  1. distrust good news after bad
  2. a period of refusing alternatives
  3. stale offers cannot reinstall it
  4. the real backup waits too

basics

~20 s

Holddown is an implementation mechanism: after a route fails, the router ignores offers no better than the lost route, so typical stale echoes cannot restart a loop. The cost: a genuine but worse backup path waits too.

solid answer

~50 s

RFC 2453 relies on split horizon, poisoned reverse, triggered updates and a small infinity; it defines no holddown timer. Many implementations add one. When a route becomes unreachable, the router starts a holddown period during which it advertises the route at 16 and ignores offers for that prefix that are no better than the route it lost; the exact acceptance rule and timer value are the implementation's. That targets the race the RFC admits it cannot close: a neighbour that has not yet heard the bad news offers the old route, and without holddown the router takes it and counting to infinity begins. The trade is convergence. If `R1` loses a metric-2 path to `10.20.30.0/24` while a real backup via `R4` offers metric 5, holddown rejects the backup too, and the prefix stays unreachable until the timer ends. Holddown buys safety with outage time.

go deeper

for a junior

Recall that holddown is a waiting period after a route fails during which a router refuses offers for it, and that it comes from implementations, not the RIP standard.

for a middle

Explain which offers holddown ignores, why that stops a stale echo from restarting a loop, and how it differs from the RFC's timeout and garbage-collection timers.

for a senior

Show the race holddown closes, quantify the outage it adds when a worse backup exists, and keep implementation timer defaults separate from protocol rules.

for a principal

Weigh holddown as a policy choice between loop safety and failover time, and decide how long it can be for a network whose design depends on backup paths.

## What the specification contains RFC 2453 (RIPv2) defines two per-route timers: a **timeout** of 180 seconds without a refresh, after which a route is invalid, and a **garbage-collection** timer of 120 seconds, during which the dead route is advertised at metric 16 before it is removed. Its loop defences are split horizon, split horizon with poisoned reverse, triggered updates and the small infinity of 16. It has **no holddown timer**, and neither does RFC 1058 (RIPv1) or RFC 2080 (RIPng). Holddown is an **implementation mechanism**, widely added and widely taught. ## How holddown typically behaves The details belong to each implementation, but the common shape is: 1. A route becomes unreachable — its timeout expires or its next hop poisons it. 2. The router starts a holddown period for that prefix and keeps advertising it at 16, as the RFC's deletion process already does. 3. While holddown runs, it ignores advertisements for the prefix that are no better than the route it lost. Many implementations still accept a strictly better offer; the exact rule varies. 4. When holddown ends, normal processing resumes and the best offer wins. The idea: right after bad news, distrust good-looking news, because it is probably an echo of the route that just died. ## The race it closes RFC 2453 admits that triggered updates leave a window open. While the poison spreads, a router that has not yet heard it can send a regular update still listing the old route, and a router that has already poisoned the prefix sees an offer better than 16 and takes it. That is how counting to infinity starts — including the three-router loop that split horizon cannot stop. Holddown targets exactly that offer. If `R1` lost a metric-2 route to `10.20.30.0/24` and neighbour `R2`, not yet informed, offers it at 3, holddown rejects the 3 — no better than the lost 2 — and the loop never forms. It is not a guarantee: a stale offer from a router whose own path never ran through this one can be better than the route just lost, and a holddown that accepts better offers lets it in. Holddown narrows the window; the cap of 16 still bounds what slips through. ## What it costs Holddown cannot tell an echo from a real backup. Suppose `R1` also has a genuine path to the prefix through `R4`, at metric 5: | Moment after the failure | Without holddown | With holddown | |---|---|---| | `R4`'s next update arrives | `R1` installs the path via `R4` at 5 (and stays exposed to stale echoes) | `R1` ignores the 5, no better than 2; the prefix stays unreachable | | for the rest of the holddown period | traffic flows via `R4` | the prefix is still unreachable | | holddown ends | unchanged | `R1` installs the path via `R4` at 5 | The backup was there all along but is held out of service for the whole period. The price of safety is outage time on every failure whose surviving path is worse than the one that failed — which, for a backup, is the usual case. ## Timer values are not protocol constants Implementations often show timer sets such as invalid 180 s, holddown 180 s and flush 240 s. Those are implementation defaults, not RIP: - the 180-second invalid value matches the RFC's timeout; - holddown does not exist in the RFC at all; - the RFC removes a route 120 seconds after it times out; a 240-second flush is the implementation's own figure. State the RFC's two timers as protocol and everything else as configuration. ## Judging it - With triggered updates and poisoned reverse in place, holddown guards a narrow race; weigh it against how often the path you need after a failure is a worse backup. - A holddown period adds directly to every outage that has a backup path, so a long one is expensive in a network built for redundancy. - A router without holddown is still fully compliant with RFC 2453; never present holddown as RIPv2 behaviour in a design review.

  • If RIP already has triggered updates and poisoned reverse, why would an implementation still want holddown?
    Because neither closes the race between a triggered poison and a regular update from a router that has not heard it yet; RFC 2453 itself says counting to infinity is still possible. Holddown narrows that window by refusing not-better offers for a while after a loss, which is what a stale echo of the dead route usually is.
  • How would you judge whether a RIP holddown period is too long for a network?
    Compare it with the outage the network can tolerate. Holddown is pure downtime for any prefix whose surviving path is worse than the one that failed, so a period of minutes means minutes of loss on every such failure. Where backups are the norm and stale echoes rare, a short holddown or none, with triggered updates, converges faster; where stale updates are likely, a longer one is the safer price.

saying these in an interview costs you the question

  • Holddown is defined in RFC 2453 alongside the update and timeout timers.
  • During holddown a RIP router keeps forwarding on the failed route.
  • Holddown speeds up convergence because routers stop exchanging bad news.
  • With holddown on, a RIP router switches to any backup path immediately.
  • Timer values of 180, 180 and 240 seconds are constants of the RIP protocol.