Network Design & High Availability
Designing networks rather than reciting protocols: topologies, redundancy at every layer, traffic priority and the WAN. Interviewers run it as the network engineer's system-design round.
on this pageshowhide
explore
- Campus and Data Center Topologies6 questions
- First-Hop Redundancy (HSRP/VRRP/GLBP)6 questions
- Link Aggregation and MLAG5 questions
- Bidirectional Forwarding Detection5 questions
- QoS and Traffic Shaping6 questions
- MPLS and L3VPN5 questions
- SD-WAN5 questions
- VRFs and Network Segmentation5 questions
- Capacity Planning and Availability6 questions
questions
page 2 of 2In a spine-leaf fabric where each leaf has one 100G uplink to each of six spines, what does losing one spine cost, and what keeps that failure small?
basics
~20 sEach leaf loses one of six uplinks: 600 Gb/s becomes 500 Gb/s, about a sixth of fabric capacity, and a 2:1 leaf becomes 2.4:1. No server loses reachability; width, routed links and ECMP keep the failure small.
For a routed network whose application tolerates at most one second of loss on a link failure, how do you build an end-to-end convergence budget?
basics
~20 sConvergence time is the sum of failure detection, propagation to other routers, path computation and forwarding-table install. Give each phase a measured or assumed value, add them with a margin, and compare against the one-second limit.
Two uplinks in a redundant pair each peak at 65% utilisation; why is this not truly redundant, and how should such links be sized against forecast growth?
basics
~20 sAfter one link fails, the survivor must carry both loads, 2 × 65% = 130% of its capacity, so it congests and drops traffic. Keep each member of a pair below about 50% at peak and order upgrades ahead of lead times.
In a looped Layer 2 access design, why should a VLAN's VRRP Active Router sit on the distribution switch that is that VLAN's spanning-tree root?
basics
~20 sSpanning tree leaves each access switch forwarding toward the VLAN's root. If VRRP's Active Router is the other distribution switch, every upstream frame crosses the inter-distribution link; co-locating root and Active keeps the path one hop.
A VRRP Active Router loses its only uplink but stays Active and black-holes traffic; how do priority tracking and preemption fix it, and how can they flap?
basics
~20 sVRRP sees only LAN advertisements, so a router with a dead uplink stays Active. Tracking, an implementation feature, cuts its priority below the Backup's; with preemption on, the Backup takes over. A flapping uplink then moves the role back and forth.
When one member of a four-member 10G LACP bundle fails, what happens to the flows it carried, and why would you set a minimum-links threshold?
basics
~20 sOnce detected, the member leaves the bundle, its in-flight frames are lost and its flows rehash onto the survivors, leaving 30G. A minimum-links threshold takes the whole bundle down below N members so routing moves traffic to a healthier path.
How does a multi-chassis link aggregation group (MLAG) let a server bundle links to two switches, and what happens when the switches' peer link fails?
basics
~20 sTwo switches present one shared LACP System ID, so the server bundles links to both as to one partner. If their peer link fails, a separate keepalive shows the peer is alive and the secondary disables its MLAG ports, avoiding split-brain.
On a congested router queue carrying many TCP flows, why does tail drop cause global synchronisation, and how do RED and weighted RED avoid it?
basics
~20 sTail drop discards only when a queue is full, so many flows lose packets at once, slow down together and leave the link idle. RED drops randomly and early as the average queue grows; weighted RED sets thresholds per drop precedence.
A 300-store retailer's SD-WAN backhauls all SaaS traffic to its data centre; what does local internet breakout gain, and what has to move with it?
basics
~20 sLocal breakout sends chosen applications straight from each store to the internet, cutting the detour, data-centre bandwidth and a shared choke point; inspection, logging, address translation and policy that lived at the data centre must now exist per store.
Designing a 2,000-server data centre, would you stretch Layer 2 across a spine-leaf fabric or route at every leaf, and what does each choice trade?
basics
~20 sRoute at every leaf with ECMP across the spines, adding a VXLAN overlay with BGP EVPN only where workloads need Layer 2 reach. Stretched Layer 2 eases VM mobility but makes the broadcast domain the failure domain.
A design review claims a branch site reaches four nines because it has two WAN circuits at 99.9% each; how do you evaluate that claim and decide what to change?
basics
~20 sTwo independent 99.9% circuits do reach 99.9999% as a pair, but the site is the whole path: a single edge router, power feed or shared duct in series caps it far lower. Find those elements, quantify each, fix the largest.
When does a provider's MPLS L3VPN still beat encrypted tunnels over internet links for an enterprise connecting forty branches?
basics
~20 sMPLS L3VPN wins where predictability is the requirement: one provider engineers the whole path, honours traffic classes under a contract and gives any-to-any reachability; internet tunnels win on price, speed of delivery, provider diversity and direct cloud access.
For QoS on a branch's 20 Mb/s WAN link carrying voice, video calls and backups, where calls break up at busy hours, what end-to-end design would you build?
basics
~20 sFind the congestion points, mark at the trust boundary, shape to the contracted rate, give voice a policed priority queue sized from calls times per-call rate, give other classes minimum shares, and cap calls with admission control so the policer never drops.
Replacing a 300-store retailer's MPLS-only WAN with SD-WAN, how do you choose each store's transport mix, and what guarantee do you give up?
basics
~20 sTier stores by outage cost and give each physically diverse transports, such as dual broadband plus LTE, or MPLS plus internet at critical sites; you give up a contracted end-to-end bound on loss and latency for measurement and steering.
Should a campus card-payment zone be a VRF on the shared routers or a physically separate network, and how does each choice change PCI DSS scope?
basics
~20 sBoth can shrink PCI DSS scope. A payment VRF on shared routers is cheap but puts every router and switch carrying it into scope and relies on configuration discipline; separate hardware costs more but keeps scope and blast radius small.
What does the BFD echo function test that asynchronous BFD Control packets do not, and why can't every link use it?
basics
~20 sBFD echo packets are addressed so the peer simply forwards them back, testing its forwarding path without its BFD process answering; echo needs the peer's consent, works only on single-hop sessions and breaks under ingress filtering.
How long does VRRPv3 take to declare a silent Active Router down, and what do sub-second advertisement intervals buy and risk?
basics
~20 sA VRRPv3 Backup declares the Active down after 3 x interval + Skew_Time, about 3.61 s at the 1-second default for priority 100. Centisecond intervals cut that below 40 ms, but risk false failovers under queueing delay and complicate mixed VRRPv2 operation.
In an MPLS core, how does RSVP-TE fast reroute keep traffic flowing within tens of milliseconds when a protected link fails?
basics
~20 sRSVP-TE fast reroute pre-signals a backup path around each protected link or node; when the failure is detected, the adjacent router redirects traffic onto that backup immediately, with no path computation or signalling, while the head-end later re-optimises.
showing 31–49 of 49