skip to content

Network Design & High Availability

Designing networks rather than reciting protocols: topologies, redundancy at every layer, traffic priority and the WAN. Interviewers run it as the network engineer's system-design round.

on this pageshow

explore

questions

page 2 of 2

In a spine-leaf fabric where each leaf has one 100G uplink to each of six spines, what does losing one spine cost, and what keeps that failure small?

level: seniorimportance: should knowfreq 25%

basics

~20 s

Each leaf loses one of six uplinks: 600 Gb/s becomes 500 Gb/s, about a sixth of fabric capacity, and a 2:1 leaf becomes 2.4:1. No server loses reachability; width, routed links and ECMP keep the failure small.

open as a page

For a routed network whose application tolerates at most one second of loss on a link failure, how do you build an end-to-end convergence budget?

level: seniorimportance: should knowfreq 22%

basics

~20 s

Convergence time is the sum of failure detection, propagation to other routers, path computation and forwarding-table install. Give each phase a measured or assumed value, add them with a margin, and compare against the one-second limit.

open as a page

In a looped Layer 2 access design, why should a VLAN's VRRP Active Router sit on the distribution switch that is that VLAN's spanning-tree root?

level: seniorimportance: should knowfreq 22%

basics

~20 s

Spanning tree leaves each access switch forwarding toward the VLAN's root. If VRRP's Active Router is the other distribution switch, every upstream frame crosses the inter-distribution link; co-locating root and Active keeps the path one hop.

open as a page

A VRRP Active Router loses its only uplink but stays Active and black-holes traffic; how do priority tracking and preemption fix it, and how can they flap?

level: seniorimportance: should knowfreq 30%

basics

~20 s

VRRP sees only LAN advertisements, so a router with a dead uplink stays Active. Tracking, an implementation feature, cuts its priority below the Backup's; with preemption on, the Backup takes over. A flapping uplink then moves the role back and forth.

open as a page

On a congested router queue carrying many TCP flows, why does tail drop cause global synchronisation, and how do RED and weighted RED avoid it?

level: seniorimportance: should knowfreq 19%

basics

~20 s

Tail drop discards only when a queue is full, so many flows lose packets at once, slow down together and leave the link idle. RED drops randomly and early as the average queue grows; weighted RED sets thresholds per drop precedence.

open as a page

A 300-store retailer's SD-WAN backhauls all SaaS traffic to its data centre; what does local internet breakout gain, and what has to move with it?

level: seniorimportance: should knowfreq 22%

basics

~20 s

Local breakout sends chosen applications straight from each store to the internet, cutting the detour, data-centre bandwidth and a shared choke point; inspection, logging, address translation and policy that lived at the data centre must now exist per store.

open as a page

In a VRF-lite campus, how do you leak a shared-services VRF's DNS and DHCP routes into each zone VRF without merging the zones?

level: seniorimportance: should knowfreq 22%

basics

~20 s

Copy only the shared-services prefix into each zone VRF and only each zone's own prefix back into the shared VRF, filtered so nothing passes onward. Overlapping zone prefixes need translation or renumbering, and leaked paths bypass any firewall.

open as a page

Designing a 2,000-server data centre, would you stretch Layer 2 across a spine-leaf fabric or route at every leaf, and what does each choice trade?

level: principalimportance: should knowfreq 15%

basics

~20 s

Route at every leaf with ECMP across the spines, adding a VXLAN overlay with BGP EVPN only where workloads need Layer 2 reach. Stretched Layer 2 eases VM mobility but makes the broadcast domain the failure domain.

open as a page

A design review claims a branch site reaches four nines because it has two WAN circuits at 99.9% each; how do you evaluate that claim and decide what to change?

level: principalimportance: should knowfreq 18%

basics

~20 s

Two independent 99.9% circuits do reach 99.9999% as a pair, but the site is the whole path: a single edge router, power feed or shared duct in series caps it far lower. Find those elements, quantify each, fix the largest.

open as a page

When does a provider's MPLS L3VPN still beat encrypted tunnels over internet links for an enterprise connecting forty branches?

level: principalimportance: should knowfreq 18%

basics

~20 s

MPLS L3VPN wins where predictability is the requirement: one provider engineers the whole path, honours traffic classes under a contract and gives any-to-any reachability; internet tunnels win on price, speed of delivery, provider diversity and direct cloud access.

open as a page

For QoS on a branch's 20 Mb/s WAN link carrying voice, video calls and backups, where calls break up at busy hours, what end-to-end design would you build?

level: principalimportance: should knowfreq 14%

basics

~20 s

Find the congestion points, mark at the trust boundary, shape to the contracted rate, give voice a policed priority queue sized from calls times per-call rate, give other classes minimum shares, and cap calls with admission control so the policer never drops.

open as a page

Replacing a 300-store retailer's MPLS-only WAN with SD-WAN, how do you choose each store's transport mix, and what guarantee do you give up?

level: principalimportance: should knowfreq 14%

basics

~20 s

Tier stores by outage cost and give each physically diverse transports, such as dual broadband plus LTE, or MPLS plus internet at critical sites; you give up a contracted end-to-end bound on loss and latency for measurement and steering.

open as a page

Should a campus card-payment zone be a VRF on the shared routers or a physically separate network, and how does each choice change PCI DSS scope?

level: principalimportance: should knowfreq 14%

basics

~20 s

Both can shrink PCI DSS scope. A payment VRF on shared routers is cheap but puts every router and switch carrying it into scope and relies on configuration discipline; separate hardware costs more but keeps scope and blast radius small.

open as a page

What does the BFD echo function test that asynchronous BFD Control packets do not, and why can't every link use it?

level: seniorimportance: nice to knowfreq 9%

basics

~20 s

BFD echo packets are addressed so the peer simply forwards them back, testing its forwarding path without its BFD process answering; echo needs the peer's consent, works only on single-hop sessions and breaks under ingress filtering.

open as a page

How long does VRRPv3 take to declare a silent Active Router down, and what do sub-second advertisement intervals buy and risk?

level: seniorimportance: nice to knowfreq 14%

basics

~20 s

A VRRPv3 Backup declares the Active down after 3 x interval + Skew_Time, about 3.61 s at the 1-second default for priority 100. Centisecond intervals cut that below 40 ms, but risk false failovers under queueing delay and complicate mixed VRRPv2 operation.

open as a page

In an MPLS core, how does RSVP-TE fast reroute keep traffic flowing within tens of milliseconds when a protected link fails?

level: seniorimportance: nice to knowfreq 11%

basics

~20 s

RSVP-TE fast reroute pre-signals a backup path around each protected link or node; when the failure is detected, the adjacent router redirects traffic onto that backup immediately, with no path computation or signalling, while the head-end later re-optimises.

open as a page

showing 31–49 of 49