When an IP router holds several equal-cost next hops for one prefix, how should ECMP split traffic, and why per flow rather than per packet?
answer
- keep each flow on one path
- reordering and needless retransmits
- hash header fields to a region
- how many flows move on change
- the same hash at every tier
basics
~20 sECMP should hash the header fields that identify a flow and map the result to one next hop, so each flow keeps a single path. Spraying per packet reorders segments, varies the path MTU and confuses diagnostics (RFC 2991).
solid answer
~50 sEqual-cost multipath applies when several routes for the same prefix survive selection with the same source and an equal metric; RFC 1812 lets a router split load across them. Splitting **per packet** (round-robin) sends one flow's packets over paths with different latencies, so they arrive out of order and TCP treats the gap as loss and retransmits; RFC 2991 adds a path MTU that varies packet by packet and unreliable ping and traceroute. So routers split **per flow**: hash the fields identifying a flow (addresses, protocol, often ports) and map the result to a next hop. RFC 2991 compares three mappings by how many flows move when a next hop is added or removed: modulo-N moves (N-1)/N, hash-threshold between 1/4 and 1/2 (analysed in RFC 2992), highest random weight only 1/N. The catches: one large flow still uses one path, and identical hashing on consecutive tiers can polarise traffic onto a subset of links.
code
pseudocode · 8 lines# hash-threshold next-hop choice (RFC 2992), N equal-cost next hops
key = crc16(src_addr, dst_addr, protocol) # 0 .. 65535
region_size = 65536 / N # N = 4 gives 16384
index = floor(key / region_size) # 0 .. N-1
return next_hops[index]
# example: N = 4, key = 40000
# 40000 / 16384 = 2.44 -> index 2 -> next_hops[2]go deeper
Know that ECMP lets a router use several equally good paths to one prefix at once, while keeping each conversation on one path.
Explain per-flow hashing, which header fields feed it, and why per-packet spraying reorders TCP segments and triggers needless retransmissions.
Reason about real imbalance: large flows, few flows, flows moved when a next hop changes, and polarisation across tiers, and say which hashing choice limits each.
Judge whether hashing across equal paths meets a network's needs or whether traffic engineering or flow-aware balancing is warranted, given the cost of reshuffled flows.
## When ECMP applies Route selection normally ends with one route per prefix. RFC 1812 notes that it can end with several: routes for the same prefix, from the same source, with equal metrics. A router may then keep one arbitrarily or **split load** across all of them, and RFC 1812 advises that an implementation offering load-splitting also give the operator a way to disable it. Using several equally good next hops at once is **equal-cost multipath (ECMP)**. It is how a leaf switch in a data-centre fabric uses all its uplinks, and how a router with two equal links to the same provider uses both. The routing protocol only decides that the paths are equal; how packets are spread across them is a forwarding decision, and RFC 2991 and RFC 2992, both Informational, describe it. ## Per packet: why spraying hurts The naive method is **round-robin**: each packet for the prefix goes to the next next-hop in turn. Load balances perfectly, but RFC 2991 lists the costs: - **Reordering.** Paths have different latencies, so packets of one flow arrive out of order. When three or more later packets arrive before a late one, TCP enters fast retransmit and resends data that was never lost, wasting bandwidth and cutting throughput. - **Variable path MTU.** If paths have different MTUs, the usable packet size changes packet by packet, defeating path MTU discovery. - **Debugging.** Ping and traceroute become unreliable, and can show a path no flow actually takes. ## Per flow: hashing The fix is to keep **every packet of a flow on one path**. The router hashes the header fields that identify the flow and maps the hash to a next hop. What counts as a flow is the implementation's choice: RFC 2991 mentions the destination alone, or source, destination and protocol. Adding transport ports spreads many connections between two hosts across paths, but RFC 2991 warns that fragments other than the first carry no ports, so they may hash differently from the first fragment. ## Three mappings, three disruption costs When a next hop is added or removed, the mapping changes and some flows move to another path, risking reordering or loss for them. RFC 2991 compares three methods: | Method | How it maps a flow | Flows moved when a next hop is added or removed | |---|---|---| | Modulo-N | hash mod N | (N-1)/N | | Hash-threshold | hash falls in one of N equal regions | between 1/4 and 1/2 | | Highest random weight | hash of flow plus next hop; highest wins | 1/N, at roughly N times the cost | RFC 2992 analyses hash-threshold: with a 16-bit key and 4 next hops, each owns a region of 65536 / 4 = 16384 values. Removing a next hop resizes the neighbouring regions, moving only flows near the boundaries; it recommends adding new regions in the middle rather than at the ends to minimise disruption. RFC 2991 recommends hash-threshold for forwarders that keep no per-flow state, and highest random weight where state is kept. ## What per-flow hashing cannot fix 1. **Large flows.** One big transfer hashes to one path, so a single flow never gets more than one link's bandwidth, and a few large flows can leave paths unevenly loaded. 2. **Few flows.** Hashing balances statistically; with a handful of flows the split can be badly uneven. 3. **Polarisation.** If consecutive tiers of routers use the same hash function over the same fields, the flows reaching a second-tier router already share a hash outcome, and hashing them again identically maps them onto a subset of its links, leaving others idle. Implementations avoid this by mixing a per-router seed into the hash or varying the fields hashed. 4. **Asymmetry.** Each router hashes independently, so a flow's return traffic may take a different path; stateful devices on only one path see half a conversation. ## Where answers go wrong - Calling round-robin the best method because it balances packets perfectly. - Expecting one flow to use the bandwidth of every path. - Quoting 1/N disruption for hash-threshold; that figure belongs to highest random weight. - Expecting ECMP across sources: an OSPF route and a static route are not equal-cost peers in the usual case, because preference picks one source before metrics are compared.
- Why might a router leave transport ports out of its ECMP hash?RFC 2991 notes that fragments other than the first carry no transport header, so a port-based hash can place them on a different path from the first fragment; it also notes that path-dependent caches such as path MTU are less useful when the path depends on ports. Including ports spreads many connections between the same two hosts better, so implementations choose per deployment.
- How does hash polarisation arise across two tiers of routers, and how is it avoided?If both tiers hash the same fields with the same function, a second-tier router receives only flows that produced one particular result at the first tier; hashing them identically maps them onto the same subset of its next hops, leaving other equal-cost links idle. Mixing a per-router seed into the hash, or hashing different fields at each tier, breaks the correlation.
saying these in an interview costs you the question
- Round-robin per packet is the best ECMP method because it balances load perfectly.
- Per-flow hashing guarantees every equal-cost link carries the same traffic.
- A single large flow is spread across all equal-cost paths.
- Hash-threshold moves only 1/N of flows when a next hop is removed.
- ECMP needs per-flow state on the router to keep each flow on one path.