skip to content

A provider-edge router runs BFD at 50 ms x 3 on hundreds of eBGP sessions, and they flap during CPU spikes; how should BFD timers be sized and flaps contained?

level: seniorimportance: should knowfreq 17%

answer

  1. detection time versus worst stall
  2. sessions x rate x two directions
  3. who processes BFD: CPU or forwarding plane
  4. hysteresis and holding sessions down
  5. lowest layer subscribes

basics

~20 s

Size BFD so the detection time exceeds the worst stall of whatever processes BFD, budget packets as sessions x rate x two, attach BFD at the lowest routing layer only, and contain flaps with hysteresis or hold-down.

solid answer

~50 s

At 50 ms x 3 a session dies after 150 ms of silence, so any control-plane stall longer than that (route churn, a busy CPU) looks like a failure. Each false Down tears down a BGP session, which causes more churn and more stalls, so the failure feeds itself. Budget the load: 300 sessions at 50 ms is 20 packets per second each way, about 12,000 per second in total, against 2,000 at 300 ms. Size the detection time above the worst stall of whatever processes BFD: forwarding-plane BFD (the `C` bit set) can run far faster than a software process. Keep the multiplier at 3 or more, use intervals both sides support (RFC 7419), and do not run BFD on iBGP where the IGP already has it (RFC 5882). For flaps, RFC 5882 allows hiding a quick Up-Down-Up, RFC 5880 allows holding a session down, and implementations add backoff dampening.

go deeper

for a junior

Recall that a shorter BFD interval means more packets and less tolerance for a busy router, so faster is not automatically better.

for a middle

Compute a detection time and a packet budget for many sessions, and explain why a stalled CPU looks the same as a dead peer to BFD.

for a senior

Size timers from measured stalls, use the C bit and RFC 7419 intervals, keep BFD off iBGP above a protected IGP, and contain flaps with hysteresis or hold-down.

for a principal

Own the trade-off across the estate: detection speed against false-failure cascades, where BFD runs in hardware or software, and how it fits the convergence budget.

## Why aggressive timers produce false failures BFD's detection time is the remote `Detect Mult` times the agreed interval (RFC 5880). At **50 ms x 3** a session goes Down after **150 ms** without a Control packet. BFD cannot tell a dead peer from a peer, or a local receiver, that was simply too busy to send or process packets for 150 ms. Typical causes of such a stall: - a burst of routing updates, such as a full-table BGP re-exchange after a session reset; - a control-plane CPU spike from management traffic, logging or a software upgrade; - congestion on the link or in the router's path to its control plane, if BFD packets are not protected there. Every false Down makes each client act: **BGP tears down the session**, withdraws its routes and later re-learns them. That re-learning is itself CPU-heavy, which stalls BFD on other sessions, which tears down *those* sessions. The aggressive timer has turned a brief stall into a cascade. ## The packet budget A BFD session sends at its interval in each direction, and the router must also receive and process the peer's packets. | Sessions | Interval | Packets per second per direction per session | Total sent + received | |---|---|---|---| | 300 | 50 ms | 20 | 300 x 20 x 2 = 12,000 | | 300 | 100 ms | 10 | 300 x 10 x 2 = 6,000 | | 300 | 300 ms | about 3.3 | about 2,000 | The multiplier changes how many losses are tolerated, not the packet rate. RFC 5881 and RFC 5883 both say operators must provision BFD rates to avoid congestion of the link, I/O or CPU, and false detection. ## Sizing rules 1. **Detection time above the worst stall.** Find out who processes BFD. If it runs in the forwarding plane and does not share fate with the control plane, the implementation sets the **`C` (Control Plane Independent) bit**, and short intervals are realistic. If a software process on the control-plane CPU handles it, the detection time must exceed that CPU's worst observed scheduling stall. 2. **Multiplier of 3 or more.** `Detect Mult` 1 means one lost packet kills the session; RFC 5880 has to squeeze the jitter to 75-90% of the interval just to keep such a session alive. A larger multiplier tolerates isolated loss at no extra packet rate. 3. **Intervals both sides support.** RFC 7419 defines a common set (3.3, 10, 20, 50, 100 ms and 1 s) so two different platforms do not fall back to a much slower shared value. 4. **Detect only as fast as the client can use it.** A 150 ms detection buys little if the client then needs seconds to reconverge; the end-to-end convergence budget belongs to the design's availability numbers, and BFD is one term in it. 5. **Subscribe at the lowest layer.** RFC 5882 section 4.4: where iBGP depends on OSPF, faster failure detection relayed to iBGP may be detrimental, because a BGP peer transition is expensive while OSPF heals the path on its own. Let the IGP use BFD and leave iBGP alone; on eBGP, BFD is the useful choice. 6. **Share sessions.** One session per path for all clients (RFC 5882) keeps the packet budget proportional to neighbours, not to protocols. ## Containing flaps | Mechanism | Source | What it does | |---|---|---| | Session state hysteresis | RFC 5882 section 3.1 | May hide an Up-Down-Up that returns within a reasonable time; clients are told only if it stays Down | | Holding sessions down | RFC 5880 section 6.8.18 | A system may keep a session in Down or AdminDown to limit how fast sessions come up, and may advertise a large Required Min RX meanwhile | | Session dampening with backoff | Implementation feature | Penalises each flap and keeps a repeatedly flapping session down for growing periods | | BGP route flap damping | RFC 2439 | Suppresses unstable *prefixes*; it does not touch the BFD session | The standards give the hooks; the backoff algorithm is each implementation's choice, so describe it as such in a design. ## Graceful restart interaction If BFD shares fate with the control plane (the `C` bit is clear), a planned control-plane restart kills BFD too, and a naive client would abort its graceful restart. RFC 5882 section 4.3 says that in this case it is best not to abort the restart; when BFD is independent of the control plane, a BFD failure means forwarding really stopped, so the restart SHOULD be aborted to avoid black holes. ## Applying it to the flapping edge router - Measure the CPU stall that triggers the flaps and size the detection time above it, for example 300 ms x 3 for software BFD. - Keep fast timers only where BFD runs in the forwarding plane. - Remove BFD from iBGP sessions that ride an IGP already using it. - Enable hysteresis or session dampening so one bad minute does not repeat every few seconds.

  • Two sessions both detect in 300 ms, one at 100 ms x 3 and one at 50 ms x 6. Which is better?
    Neither is better everywhere. 50 ms x 6 tolerates five consecutive lost packets instead of two, so it rides out isolated loss better, but it sends twice the packets and still dies after the same 300 ms stall. When stalls rather than random loss cause the flaps, the cheaper 100 ms x 3 gives the same protection at half the load.
  • Why might BFD Down abort a BGP graceful restart on one router but not on another?
    It depends on the C bit. If BFD runs independently of the control plane, a BFD failure means forwarding stopped, so RFC 5882 says the restart SHOULD be aborted to avoid black holes. If BFD shares fate with the control plane, the restart itself kills BFD, and aborting would defeat graceful restart, so it is best not to abort.

saying these in an interview costs you the question

  • The fastest BFD timer the hardware allows is always the best choice.
  • Raising Detect Mult increases the BFD packet rate.
  • BFD should run on every BGP session, iBGP included, whatever the IGP does.
  • BFD dampening is the same thing as BGP route flap damping.
  • A Detect Mult of 1 is safe because jitter keeps packets early.