What is Bidirectional Forwarding Detection (BFD), and why do routers run it beside BGP or OSPF instead of relying on their own keepalives?
answer
- a detector, not a router
- Ethernet may stay up when the peer dies
- hold timer versus milliseconds
- the client acts, BFD only reports
- RFC 5880 and RFC 5882
basics
~20 sBFD (RFC 5880) is a lightweight hello protocol that only checks, in milliseconds, whether the forwarding path to a neighbour works; it carries no routes and tells clients such as BGP, OSPF or a static route to react.
solid answer
~50 sMany links, Ethernet above all, give no loss-of-signal when the router on the far side dies, so the local port stays up and a routing protocol must notice through its own hellos. Those are coarse: RFC 5882 notes that OSPF cannot detect a failure in under two seconds, and a BGP session waits for its hold time (RFC 4271 suggests 90 s). BFD sends small control packets at a negotiated interval and declares the session Down after `Detect Mult` intervals pass with nothing received, so `300 ms x 3` detects in 900 ms. It has no discovery and learns no topology: the client (BGP, OSPF, a static route) bootstraps the session with the neighbour address, and on Down the client acts with its own machinery: BGP tears the session down, OSPF drops the neighbour, the static route is withdrawn. One session per path serves every client.
go deeper
Recall that BFD only detects that a forwarding path died; it carries no routes. Name the clients that act on it: BGP, OSPF and static routes.
Explain why Ethernet hides a dead peer, how slow the protocol timers are by comparison, and what each client does when the BFD session goes Down.
Show you know where BFD must not trigger action: AdminDown, peers without BFD, and why one shared session per path beats one per client.
Frame BFD as a shared detection service: decide which layer subscribes to it and weigh faster detection against the cost of false failures across the network.
## The gap BFD fills A router needs to know quickly when the next router on a path stops forwarding. Some media tell it for free: a SONET alarm or a dark fibre drops the interface. **Ethernet often does not.** When a carrier hands off a service on an Ethernet port, the local port's link partner is the carrier's switch, so when the customer's or the provider's router dies behind that switch, the local interface stays up and nothing at layer 1 or 2 reports the failure. The routing protocols do have liveness checks, but they are slow: - **OSPF** detects a dead neighbour through its Hello and dead intervals, which are counted in seconds; RFC 5882 states OSPF's minimum detection time as two seconds. - **BGP** waits for its **hold timer** to expire. RFC 4271 suggests a hold time of 90 seconds, with keepalives at roughly a third of it; the actual value is negotiated per session and implementations choose their own defaults. - **Static routes** have no liveness check of their own: without some tracking mechanism, the route stays as long as its interface is up. RFC 5880's introduction names the problem plainly: hello mechanisms in existing protocols detect failures in no better than a second, which is a great deal of lost data at gigabit rates. ## What BFD is, and what it is not **Bidirectional Forwarding Detection** (RFC 5880, Standards Track) is a single, protocol-independent liveness check between two forwarding engines. In its mandatory **asynchronous mode** each side sends small **BFD Control packets** (a 24-byte mandatory section) at a negotiated interval; if `Detect Mult` intervals pass with none received, the session goes **Down**. | Property | Routing protocol hellos | BFD | |---|---|---| | Purpose | Neighbour discovery plus liveness | Path liveness only | | Carries routes or topology | Yes | No | | Discovery | Built in | None: the client supplies the neighbour | | Typical detection | Seconds to minutes | Tens to hundreds of milliseconds | | Shared across protocols | No, each protocol runs its own | Yes, one session per path | BFD is **advisory**. RFC 5882 compares it to hardware signalling loss of light on a fibre: it reports a fact, and the client uses its existing mechanisms to react. It carries no application information and computes no paths. ## How a client uses a BFD session RFC 5882 describes the interaction: 1. The client (say, BGP) learns a neighbour by its own means (configuration or discovery) and asks BFD for a session to that address, with timing parameters. 2. BFD runs its three-way handshake through the **Down**, **Init** and **Up** states and reports Up to the client. 3. If the session leaves Up because packets stopped arriving, BFD notifies every client bound to it. 4. Each client reacts with its own mechanism, as if its own timer had expired. RFC 5882 also says an implementation SHOULD run **one BFD session per data-protocol path** no matter how many clients use it, and that clients of the same data protocol to the same neighbour MUST share one session. IPv4 and IPv6 over the same link are two paths, so they get two sessions. ## Client by client | Client | Reaction to BFD Down (RFC 5882 section 10 and section 5) | |---|---| | eBGP | Tear down the BGP session, unless a graceful restart is in progress | | OSPFv2 / OSPFv3 | Tear down the corresponding OSPF neighbour | | Static route | Withdraw the route (and from any protocol redistributing it) | | RIP | Simulate expiry of the timeout timer for routes learned from that neighbour | The reaction is what turns a 900 ms detection into a reroute: BFD alone moves no traffic. ## When the client should not react - **AdminDown.** If either side moves the session to AdminDown, the path is not known to be broken; BFD was disabled administratively. Clients that have their own liveness check SHOULD NOT take action. A static route, which has no other check, SHOULD treat AdminDown as Down. - **A peer that does not run BFD.** Adjacency establishment SHOULD NOT be blocked if the neighbour is believed not to support BFD; it SHOULD be blocked when both sides want BFD but the session cannot come up, since the data path may be broken while the control protocol still talks. - **Not a replacement.** BFD does not retire the client's own keepalive: a control protocol may depend on things BFD does not check, multicast Hellos for example, so its timers keep running. ## Worked example A customer router 192.0.2.1 in AS 64500 peers over eBGP with a provider router 192.0.2.2 in AS 64510 across a carrier Ethernet hand-off. The provider router loses power; the customer's port stays up. Without BFD, the customer keeps forwarding into the dead path until the BGP hold time runs out, tens of seconds. With a single-hop BFD session at 300 ms x 3, the customer declares the session Down after about 900 ms, BGP tears down the session, the routes learned from 192.0.2.2 are withdrawn, and traffic moves to the backup path.
- Does BFD replace BGP's hold timer or OSPF's dead interval?No. RFC 5882 calls BFD advisory: it adds a much faster signal that the forwarding path failed, but the client keeps its own liveness mechanism, because the client may depend on things BFD does not verify, such as multicast Hellos. The protocol timers still run and still catch failures that BFD cannot see, for example a control process that hangs while forwarding continues.
- A BGP speaker sees its peer move the BFD session to AdminDown. Should it drop the BGP session?No. AdminDown means BFD was disabled administratively, not that the path failed. RFC 5882 says a client with its own liveness check SHOULD NOT treat this as a connectivity failure, so BGP keeps the session and falls back on its hold timer. A static route has no other check, so for it AdminDown SHOULD be treated as Down.
BFD is a smoke detector wired to a building's sprinkler controller: it never puts out a fire or decides where water goes, it only raises the alarm quickly, and the controller (BGP, OSPF or the static route) does the acting.
saying these in an interview costs you the question
- BFD is a fast routing protocol that exchanges reachability between neighbours.
- Once BFD runs, BGP no longer needs its hold timer or keepalives.
- An Ethernet port always goes down when the router behind it fails.
- Every client protocol needs its own BFD session to the same neighbour.
- A peer moving BFD to AdminDown should make BGP tear the session down immediately.