skip to content

In an Ethernet link aggregation group, how does the sender choose a member link for each frame, and which header fields should feed that hash?

level: middleimportance: must knowfreq 38%

answer

  1. same flow, same member
  2. ordering beats balance
  3. MAC, IP, then ports
  4. the standard leaves the algorithm open
  5. two routers share one MAC pair

basics

~20 s

The sender hashes header fields such as MAC addresses, IP addresses and TCP or UDP ports and maps the result to one member, so a flow's frames share a link and stay in order. Richer inputs spread more flows; no flow exceeds one member.

solid answer

~50 s

Each end of a bundle picks the transmit member for its own frames by **hashing** a set of header fields and mapping the result onto the active members. All frames with the same inputs land on the same member, which keeps each flow in order; the price is that one flow never exceeds one member's speed. IEEE 802.1AX does **not** standardise the algorithm (RFC 7209 states this): the only requirement is in-order delivery per flow, so each implementation offers its own choice of fields, typically **Layer 2** (source and destination MAC, VLAN), **Layer 3** (source and destination IP) and **Layer 4** (TCP or UDP ports). Pick the richest inputs the traffic offers: between two routers every frame carries the same MAC pair, so a MAC-only hash puts the whole bundle's traffic on one member, while adding IP addresses and ports spreads connections. The two directions are hashed independently and may use different members.

go deeper

for a junior

Recall that the bundle hashes header fields to pick one member per flow, so frames stay in order and one flow is limited to one member.

for a middle

Explain the Layer 2, 3 and 4 hash inputs, why a MAC-only hash fails between routers, and that each end hashes its own transmit direction.

for a senior

Diagnose uneven member load from the flow mix and hash inputs, recognise polarisation between tiers, and fix it with richer fields or per-device seeds.

for a principal

Judge when a bundle's per-flow ceiling and imbalance are acceptable versus moving to faster single links, given the flow sizes the workload actually produces.

## Why a bundle hashes instead of taking turns A link aggregation group (LAG) offers several member links between two devices. The simplest way to use them, sending frame 1 on member 1, frame 2 on member 2 and so on, would balance perfectly and break things. Members can differ slightly in latency and queue depth, so frames of one stream would arrive out of order. RFC 2991, written about multipath forwarding in general, spells out the cost: when three or more packets arrive ahead of a late one, TCP assumes loss and enters fast retransmit, wasting bandwidth and cutting throughput. So a bundle keeps every **flow** on one member. RFC 7209 (section 4.1) states the rule for Ethernet bundles plainly: the load-balancing algorithm is flexible, and the only requirement is that it ensures in-order frame delivery for a given flow. ## How the choice is made 1. The sender extracts the configured **keys** from the frame. 2. It runs them through a hash function, producing a number. 3. It maps that number onto the list of members in the Distributing state, for example by taking it modulo the member count or indexing a table. 4. The frame leaves on the chosen member. Every later frame with the same keys gets the same result. The IEEE standard does not define this function. RFC 7209 notes that different implementations behave differently, and a bundle works correctly even when the two ends hash differently, because each end only decides its own **transmit** side. A TCP connection's data and its acknowledgements may travel on different members. ## Which fields to hash RFC 7209 lists the usual inputs: | Layer | Fields | Spreads traffic between | |---|---|---| | Layer 2 | Source MAC, destination MAC, VLAN | Different pairs of hosts on one segment | | Layer 3 | Source IP, destination IP | Different pairs of IP endpoints | | Layer 4 | TCP or UDP source and destination port | Different connections between the same endpoints | RFC 6790 (section 1) adds the trade-off: too conservative a choice maps many flows to one hash value and balances poorly; too aggressive a choice can split one flow across members. Where each choice fails: - **MAC only between two routers or a router and a firewall.** Every frame carries the same two MAC addresses, so the hash returns one value and the whole bundle collapses onto one member. - **IP only between two busy hosts.** A server talking to one storage array or one proxy keeps one address pair, so all its connections share a member. - **IP plus ports.** Each TCP or UDP connection has its own source port, so many connections between the same two hosts spread across members. This is the usual best choice for routed or server traffic. - **Fragments.** Only the first fragment of a fragmented IP packet carries the TCP or UDP header, a problem RFC 2991 (section 3) raises for transport-layer inputs. Implementations typically fall back to the IP fields for fragments so all pieces of one packet stay together. ## What the hash cannot fix - **One flow, one member.** A single backup stream over a 4 x 10G bundle tops out at 10G however good the hash is. - **Few large flows.** With four members and four large flows, the chance that the hash places them on four different members is 4/4 x 3/4 x 2/4 x 1/4 = 24/256, under 10 per cent; usually two of them share a member while another member idles. - **Hash polarisation.** When two tiers of switches run the same hash over the same fields, the second tier re-sorts already-sorted traffic. If the first tier sends to switch X only flows whose hash is even, and X picks among its own two members by the same hash modulo 2, every flow X receives lands on the same member and the other stays idle. RFC 7938 (section 6.1) warns that stacking bundles beneath multipath routing raises this risk, and RFC 6790 (section 9) notes that many implementations add a random factor to the hash precisely to avoid polarisation. The fix is to vary the hash seed or the input fields per device. ```pseudocode // one transmit decision on a LAG members = active members in Distributing state // e.g. [m0, m1, m2, m3] key = (src_ip, dst_ip, protocol, src_port, dst_port) if packet is a non-first fragment: key = (src_ip, dst_ip, protocol) // no ports present h = hash(key, seed) // seed varies per device send frame on members[h mod len(members)] ``` ## Reading a utilisation graph Uneven member utilisation is normal with few flows and expected with poor inputs. Before blaming the bundle, check which fields the hash uses, how many concurrent flows cross it, and whether an upstream device uses the same hash. More, smaller flows, richer inputs and per-device seeds are what move a bundle toward its aggregate capacity.

  • Two routers are joined by a four-member bundle using a MAC-based hash; why is one member saturated while three sit idle?
    All routed traffic between two routers carries the same pair of router MAC addresses, so a MAC-only hash returns one value for every frame and maps it all to one member. Changing the hash inputs to include IP addresses and TCP or UDP ports spreads the many flows inside that traffic across members.
  • What is hash polarisation in a two-tier design with link bundles, and how do you avoid it?
    If the upper tier sends a switch only flows whose hash satisfies one condition, and that switch hashes the same fields with the same function over its own members, the already-filtered flows all map to the same subset of members. Using a different seed or different fields at each tier breaks the correlation, which is why many implementations add a per-device random factor.
  • Must both ends of a link bundle use the same hash algorithm?
    No. Each end only chooses the member for frames it transmits, and the IEEE standard defines no common algorithm. The bundle works with different algorithms at each end; it only means the two directions of one connection may use different members.

saying these in an interview costs you the question

  • The two ends must negotiate one shared hash algorithm
  • A MAC-based hash balances traffic well between two routers
  • The bundle sends frames round-robin to balance load evenly
  • IEEE 802.1AX standardises which header fields are hashed
  • Adding more members makes one TCP connection faster
  • Hashing on ports is always safe, even for IP fragments