skip to content

Network Design & High Availability

Designing networks rather than reciting protocols: topologies, redundancy at every layer, traffic priority and the WAN. Interviewers run it as the network engineer's system-design round.

on this pageshow

explore

questions

page 1 of 2

What is Bidirectional Forwarding Detection (BFD), and why do routers run it beside BGP or OSPF instead of relying on their own keepalives?

level: juniorimportance: must knowfreq 34%

answer

  1. a detector, not a router
  2. Ethernet may stay up when the peer dies
  3. hold timer versus milliseconds
  4. the client acts, BFD only reports
  5. RFC 5880 and RFC 5882

basics

~20 s

BFD (RFC 5880) is a lightweight hello protocol that only checks, in milliseconds, whether the forwarding path to a neighbour works; it carries no routes and tells clients such as BGP, OSPF or a static route to react.

solid answer

~50 s

Many links, Ethernet above all, give no loss-of-signal when the router on the far side dies, so the local port stays up and a routing protocol must notice through its own hellos. Those are coarse: RFC 5882 notes that OSPF cannot detect a failure in under two seconds, and a BGP session waits for its hold time (RFC 4271 suggests 90 s). BFD sends small control packets at a negotiated interval and declares the session Down after `Detect Mult` intervals pass with nothing received, so `300 ms x 3` detects in 900 ms. It has no discovery and learns no topology: the client (BGP, OSPF, a static route) bootstraps the session with the neighbour address, and on Down the client acts with its own machinery: BGP tears the session down, OSPF drops the neighbour, the static route is withdrawn. One session per path serves every client.

go deeper

for a junior

Recall that BFD only detects that a forwarding path died; it carries no routes. Name the clients that act on it: BGP, OSPF and static routes.

for a middle

Explain why Ethernet hides a dead peer, how slow the protocol timers are by comparison, and what each client does when the BFD session goes Down.

for a senior

Show you know where BFD must not trigger action: AdminDown, peers without BFD, and why one shared session per path beats one per client.

for a principal

Frame BFD as a shared detection service: decide which layer subscribes to it and weigh faster detection against the cost of false failures across the network.

## The gap BFD fills A router needs to know quickly when the next router on a path stops forwarding. Some media tell it for free: a SONET alarm or a dark fibre drops the interface. **Ethernet often does not.** When a carrier hands off a service on an Ethernet port, the local port's link partner is the carrier's switch, so when the customer's or the provider's router dies behind that switch, the local interface stays up and nothing at layer 1 or 2 reports the failure. The routing protocols do have liveness checks, but they are slow: - **OSPF** detects a dead neighbour through its Hello and dead intervals, which are counted in seconds; RFC 5882 states OSPF's minimum detection time as two seconds. - **BGP** waits for its **hold timer** to expire. RFC 4271 suggests a hold time of 90 seconds, with keepalives at roughly a third of it; the actual value is negotiated per session and implementations choose their own defaults. - **Static routes** have no liveness check of their own: without some tracking mechanism, the route stays as long as its interface is up. RFC 5880's introduction names the problem plainly: hello mechanisms in existing protocols detect failures in no better than a second, which is a great deal of lost data at gigabit rates. ## What BFD is, and what it is not **Bidirectional Forwarding Detection** (RFC 5880, Standards Track) is a single, protocol-independent liveness check between two forwarding engines. In its mandatory **asynchronous mode** each side sends small **BFD Control packets** (a 24-byte mandatory section) at a negotiated interval; if `Detect Mult` intervals pass with none received, the session goes **Down**. | Property | Routing protocol hellos | BFD | |---|---|---| | Purpose | Neighbour discovery plus liveness | Path liveness only | | Carries routes or topology | Yes | No | | Discovery | Built in | None: the client supplies the neighbour | | Typical detection | Seconds to minutes | Tens to hundreds of milliseconds | | Shared across protocols | No, each protocol runs its own | Yes, one session per path | BFD is **advisory**. RFC 5882 compares it to hardware signalling loss of light on a fibre: it reports a fact, and the client uses its existing mechanisms to react. It carries no application information and computes no paths. ## How a client uses a BFD session RFC 5882 describes the interaction: 1. The client (say, BGP) learns a neighbour by its own means (configuration or discovery) and asks BFD for a session to that address, with timing parameters. 2. BFD runs its three-way handshake through the **Down**, **Init** and **Up** states and reports Up to the client. 3. If the session leaves Up because packets stopped arriving, BFD notifies every client bound to it. 4. Each client reacts with its own mechanism, as if its own timer had expired. RFC 5882 also says an implementation SHOULD run **one BFD session per data-protocol path** no matter how many clients use it, and that clients of the same data protocol to the same neighbour MUST share one session. IPv4 and IPv6 over the same link are two paths, so they get two sessions. ## Client by client | Client | Reaction to BFD Down (RFC 5882 section 10 and section 5) | |---|---| | eBGP | Tear down the BGP session, unless a graceful restart is in progress | | OSPFv2 / OSPFv3 | Tear down the corresponding OSPF neighbour | | Static route | Withdraw the route (and from any protocol redistributing it) | | RIP | Simulate expiry of the timeout timer for routes learned from that neighbour | The reaction is what turns a 900 ms detection into a reroute: BFD alone moves no traffic. ## When the client should not react - **AdminDown.** If either side moves the session to AdminDown, the path is not known to be broken; BFD was disabled administratively. Clients that have their own liveness check SHOULD NOT take action. A static route, which has no other check, SHOULD treat AdminDown as Down. - **A peer that does not run BFD.** Adjacency establishment SHOULD NOT be blocked if the neighbour is believed not to support BFD; it SHOULD be blocked when both sides want BFD but the session cannot come up, since the data path may be broken while the control protocol still talks. - **Not a replacement.** BFD does not retire the client's own keepalive: a control protocol may depend on things BFD does not check, multicast Hellos for example, so its timers keep running. ## Worked example A customer router 192.0.2.1 in AS 64500 peers over eBGP with a provider router 192.0.2.2 in AS 64510 across a carrier Ethernet hand-off. The provider router loses power; the customer's port stays up. Without BFD, the customer keeps forwarding into the dead path until the BGP hold time runs out, tens of seconds. With a single-hop BFD session at 300 ms x 3, the customer declares the session Down after about 900 ms, BGP tears down the session, the routes learned from 192.0.2.2 are withdrawn, and traffic moves to the backup path.

  • Does BFD replace BGP's hold timer or OSPF's dead interval?
    No. RFC 5882 calls BFD advisory: it adds a much faster signal that the forwarding path failed, but the client keeps its own liveness mechanism, because the client may depend on things BFD does not verify, such as multicast Hellos. The protocol timers still run and still catch failures that BFD cannot see, for example a control process that hangs while forwarding continues.
  • A BGP speaker sees its peer move the BFD session to AdminDown. Should it drop the BGP session?
    No. AdminDown means BFD was disabled administratively, not that the path failed. RFC 5882 says a client with its own liveness check SHOULD NOT treat this as a connectivity failure, so BGP keeps the session and falls back on its hold timer. A static route has no other check, so for it AdminDown SHOULD be treated as Down.

BFD is a smoke detector wired to a building's sprinkler controller: it never puts out a fire or decides where water goes, it only raises the alarm quickly, and the controller (BGP, OSPF or the static route) does the acting.

saying these in an interview costs you the question

  • BFD is a fast routing protocol that exchanges reachability between neighbours.
  • Once BFD runs, BGP no longer needs its hold timer or keepalives.
  • An Ethernet port always goes down when the router behind it fails.
  • Every client protocol needs its own BFD session to the same neighbour.
  • A peer moving BFD to AdminDown should make BGP tear the session down immediately.
open as a page

In network design, what does an availability of 99.99% mean in downtime per year, and how is availability derived from MTBF and MTTR?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Availability is the fraction of time a system is in service: MTBF / (MTBF + MTTR). At 99.99% the allowed downtime is 0.01% of a year, about 52.6 minutes; 99.9% allows about 8.76 hours and 99.999% about 5.26 minutes.

open as a page

How does a first-hop redundancy protocol such as VRRP keep a LAN's static default gateway working when one router fails?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Routers share one virtual gateway IP and one virtual MAC; an election makes one Active, which answers ARP and forwards. If its advertisements stop, a Backup claims the same addresses, so hosts keep their configured gateway unchanged.

open as a page

In SD-WAN, what is the difference between the overlay and the underlay, and why does that split let a branch use almost any transport?

level: juniorimportance: must knowfreq 40%

basics

~20 s

The underlay is the rented transports, such as MPLS, broadband or LTE, that only carry packets between edge addresses; the overlay is the encrypted tunnels, routes and policy built over them, so sites see one private network whatever the transport.

open as a page

In BFD asynchronous mode, how do Desired Min TX, Required Min RX and Detect Mult set each side's transmit rate and detection time?

level: middleimportance: must knowfreq 29%

basics

~20 s

Each BFD router transmits at the larger of its own Desired Min TX and the peer's Required Min RX; its detection time is the peer's Detect Mult times that peer's agreed interval, so 300 ms x 3 gives 900 ms.

open as a page

How does a spine-leaf data-centre fabric differ from a three-tier access, distribution and core design, and why do data centres prefer it?

level: middleimportance: must knowfreq 55%

basics

~20 s

A three-tier design is a tree that grows by buying bigger upper switches; spine-leaf connects every leaf to every spine, so any two racks are one spine apart, with one equal-cost path per spine and growth by adding switches.

open as a page

In network availability math, how do you combine components in series versus redundant components in parallel, and what does the parallel formula assume?

level: middleimportance: must knowfreq 45%

basics

~20 s

Components in series must all work, so their availabilities multiply. Redundant components in parallel fail only when all fail, giving 1 − (1 − A)^n for n identical units — but only if their failures are independent.

open as a page

In VRRP (RFC 9568), how is the Active Router elected, and what do priority 255, priority 0 and preemption change?

level: middleimportance: must knowfreq 38%

basics

~20 s

The highest-priority VRRP router becomes Active; if two Active Routers tie, the higher primary IP wins. Priority 255 marks the address owner, which always preempts; priority 0 signals a graceful exit. Preemption, on by default, lets a strictly higher-priority Backup reclaim the role.

open as a page

In MPLS, how does a label switching router forward a packet by pushing, swapping and popping labels instead of doing an IP lookup?

level: middleimportance: must knowfreq 32%

basics

~20 s

The ingress router classifies a packet once and pushes a label; each transit router swaps the top label using an exact-match table, and the last or second-to-last router pops it, so the core never re-reads the IP header.

open as a page

In DiffServ QoS, how should voice, video and bulk traffic be marked, and where should the network's trust boundary sit?

level: middleimportance: must knowfreq 42%

basics

~20 s

Mark by service class at the edge: voice EF (46), interactive video AF41, signalling CS5, bulk AF11 or Lower Effort, everything else DF. Place the trust boundary as close to the source as possible, re-marking untrusted hosts and policing what you trust.

open as a page

What is VRF-lite, and how does a campus carry separate per-zone routing tables across several routers without MPLS?

level: middleimportance: must knowfreq 35%

basics

~20 s

VRF-lite is VRFs without MPLS or MP-BGP: each router keeps a routing table per zone, and every inter-router link carries one 802.1Q subinterface per VRF with its own routing adjacency, so the arriving interface tells each hop which table to use.

open as a page

A data-centre leaf switch has 48 x 25G server ports and 6 x 100G spine uplinks; what is its oversubscription ratio, and when does that ratio hurt?

level: seniorimportance: must knowfreq 35%

basics

~20 s

Server-facing bandwidth is 48 x 25 = 1,200 Gb/s and uplink bandwidth is 6 x 100 = 600 Gb/s, so the leaf is 2:1 oversubscribed. It hurts when many servers send off-rack at once: replication, shuffles, backups, incast bursts.

open as a page

A provider's MPLS L3VPN carries two customers that both use 10.0.0.0/8; how do VRFs, route distinguishers and route targets keep their routes and traffic apart?

level: seniorimportance: must knowfreq 26%

basics

~20 s

Each PE holds a VRF per customer; a route distinguisher turns each customer's 10.0.0.0/8 into a distinct VPN-IPv4 route for MP-BGP, route targets decide which remote VRFs import it, and an inner VPN label selects the VRF at the egress PE.

open as a page

In QoS design, how do traffic policing and traffic shaping differ, and where does each belong when a branch's 1 Gb/s port connects to a 100 Mb/s WAN contract?

level: seniorimportance: must knowfreq 36%

basics

~20 s

Policing drops or re-marks traffic above a rate at once, adding no delay; shaping buffers the excess and sends it later, adding delay instead of loss. The branch shapes egress to 100 Mb/s with queues inside; the provider polices.

open as a page

How does SD-WAN application-aware routing keep a store's voice calls on a path that meets their SLA, and what happens when that path degrades?

level: seniorimportance: must knowfreq 28%

basics

~20 s

Edges probe every tunnel for loss, latency and jitter, compare the results with each application's SLA thresholds, and send voice only over compliant paths; when a path breaches, new flows, and often existing ones, move to another compliant path.

open as a page

In a VRF-lite campus, how do you force all traffic between the corporate and guest VRFs through a firewall, and what quietly breaks that?

level: seniorimportance: must knowfreq 28%

basics

~20 s

Give the firewall an interface in each VRF, keep no leaks between the zones, and point each zone VRF's routes for other zones at the firewall. A leaked more-specific route or an asymmetric return path silently bypasses or breaks it.

open as a page

What distinguishes north-south from east-west traffic in a data centre, and why does the mix shape the network's topology?

level: juniorimportance: should knowfreq 45%

basics

~20 s

North-south traffic crosses the data centre's edge, between outside clients and inside servers; east-west traffic stays inside, server to server. Tree designs were sized for north-south; heavy east-west traffic favours a spine-leaf fabric with many equal paths.

open as a page

In IP networks, what problem does quality of service (QoS) solve, and why does it change nothing on a link that is never congested?

level: juniorimportance: should knowfreq 46%

basics

~20 s

QoS decides which packets wait and which are dropped when traffic reaches an output faster than the link can send it. It trades delay, jitter and loss between classes but adds no bandwidth, so without a queue it has nothing to decide.

open as a page

In campus network design, how does separating guest traffic with a VLAN differ from separating it with a VRF?

level: juniorimportance: should knowfreq 42%

basics

~20 s

A VLAN separates guest and corporate hosts only at Layer 2; once both VLANs reach a router with one shared routing table, it routes between them. A VRF gives the guest interfaces their own routing table, so corporate routes are simply absent.

open as a page

How do single-hop BFD (RFC 5881) and multihop BFD (RFC 5883) sessions differ, and why does single-hop insist on a TTL of 255?

level: middleimportance: should knowfreq 16%

basics

~20 s

Single-hop BFD runs over one link on UDP 3784 and sends and checks TTL 255, so only an on-link neighbour can inject packets; multihop BFD runs over routed paths on UDP 4784, relies on authentication instead, and forbids echo.

open as a page

In campus network design, what is a collapsed core, and when should a campus keep a separate core layer instead?

level: middleimportance: should knowfreq 30%

basics

~20 s

A collapsed core merges the distribution and core layers into one redundant pair that every access switch uplinks to. It fits a single building; a campus with several distribution blocks keeps a core so blocks avoid a full mesh.

open as a page

In network and data-centre design, what is the difference between N+1 and 2N redundancy, and when is 2N worth its cost?

level: middleimportance: should knowfreq 28%

basics

~20 s

N+1 adds one spare unit to the N a load needs, usually inside one shared system; 2N builds two complete, independent systems, each able to carry the full load. 2N costs more but survives losing an entire side.

open as a page

How do VRRP, HSRP and GLBP differ in standards status, election messages, timers and the way they share load across gateways?

level: middleimportance: should knowfreq 34%

basics

~20 s

VRRP is the IETF standard (RFC 9568); HSRP is a vendor protocol in Informational RFC 2281; GLBP is a vendor protocol without an RFC. VRRP and HSRP forward through one router per group; GLBP spreads one virtual IP's hosts across several forwarders.

open as a page

In an MPLS core, how does LDP build label-switched paths, and why do those paths follow the IGP's best route?

level: middleimportance: should knowfreq 20%

basics

~20 s

LDP routers find neighbours with UDP hellos, open a TCP session on port 646 and advertise a label for each prefix they route; each router forwards with the label from its IGP next hop, so every LSP copies the IGP's path.

open as a page

In router QoS queuing, why pair a strict-priority queue for voice with class-based weighted fair queues, and why must the priority queue be policed?

level: middleimportance: should knowfreq 30%

basics

~20 s

A strict-priority queue sends waiting voice first, giving the low delay and jitter Expedited Forwarding needs; weighted class queues guarantee every other class a minimum share. The priority queue must be rate-limited, or excess priority traffic starves every other class.

open as a page

In an SD-WAN, what do the central controllers do, and how does a new branch edge provision itself through zero-touch provisioning?

level: middleimportance: should knowfreq 18%

basics

~20 s

Controllers hold policy and distribute routes, tunnel endpoints and keys to edges, but user traffic normally flows edge to edge; zero-touch provisioning lets an unconfigured edge get an address, prove a pre-registered identity and download its site configuration.

open as a page

A provider-edge router runs BFD at 50 ms x 3 on hundreds of eBGP sessions, and they flap during CPU spikes; how should BFD timers be sized and flaps contained?

level: seniorimportance: should knowfreq 17%

basics

~20 s

Size BFD so the detection time exceeds the worst stall of whatever processes BFD, budget packets as sessions x rate x two, attach BFD at the lowest routing layer only, and contain flaps with hysteresis or hold-down.

open as a page

showing 1–30 of 49