When one member of a four-member 10G LACP bundle fails, what happens to the flows it carried, and why would you set a minimum-links threshold?
answer
- detection: link loss or LACP timeout
- survivors absorb the flows
- rehash moves more than you think
- routing still sees one interface
- fewer than N members, bundle down
basics
~20 sOnce detected, the member leaves the bundle, its in-flight frames are lost and its flows rehash onto the survivors, leaving 30G. A minimum-links threshold takes the whole bundle down below N members so routing moves traffic to a healthier path.
solid answer
~50 sThe bundle first has to **detect** the failure: loss of physical link is near-instant, while a member that keeps link but stops passing frames is caught only when LACP's partner information times out (3 s or 90 s with the IEEE's fast or slow timers). The member then leaves the Distributing state, frames already queued on it are lost, and the hash re-maps its flows onto the three survivors; Capacity drops from 40G to 30G, so anything above 75 % load now congests. Depending on how the implementation maps hashes to members, flows on healthy members may move too. Above the bundle the **topology does not change**: RFC 7130 notes that routing protocols treat a bundle as one bigger interface. A **minimum-links** threshold, an implementation feature, fixes that blindness by declaring the whole bundle down below N members, so routing or a redundancy protocol moves traffic to a path that still has full capacity.
go deeper
Recall that a bundle survives a member failure with less capacity, and that the failed member's flows move to the remaining members.
Explain how link loss and LACP timeouts detect a member failure, and why a modulo-based hash moves flows from healthy members too.
Size bundles for a member failure, alert on member count rather than interface state, and place minimum-links thresholds only where an alternative path exists.
Decide how much degraded capacity a design tolerates before failing over, weighing an overloaded bundle against the cost of shifting all its traffic to another path.
## The scenario A server is bundled to its top-of-rack switching with four 10G members using LACP, for 40G of aggregate capacity. One member's optic dies, or its cable is pulled, or it starts dropping frames in one direction. What the bundle does next happens in three steps: detect, remove, redistribute. ## Step 1: detection | Failure | Who notices | How fast | |---|---|---| | Physical link loss | Both ends' interfaces | Effectively immediately | | One-way failure, link still up | The end that stops hearing LACPDUs | After LACP's timeout: 3 s with the IEEE's fast timers, 90 s with the slow ones | | Forwarding fault invisible to LACP | Only an active per-member check | RFC 7130 runs BFD on each member for this, a separate subject | A **static bundle** has only the first row: a member with link but no forwarding keeps its share of traffic indefinitely. ## Step 2: removal The member leaves the **Distributing** state, so the transmit hash no longer selects it. Frames already queued on it, or in flight on the wire, are lost. TCP flows that lose segments retransmit them; UDP flows simply see a gap. ## Step 3: redistribution The hash now maps flows over three members instead of four. How many flows move depends on how the implementation maps hash values to members: - **A plain modulo of the member count** reshuffles far more than the failed member's flows. Take twelve equally likely hash values 0-11. With four members, value *h* uses member *h mod 4*; with three, *h mod 3*. The two agree only for values 0, 1 and 2. So 9 of 12 hash values change member: the 3 that had to move (those on the failed member, values 3, 7 and 11) and 6 that sat on healthy members. - **A table-based or 'resilient' mapping**, offered by some implementations, reassigns only the failed member's buckets, so flows on healthy members stay where they were. Moving a flow does not break it, since its frames still go in order on the new member, but a move mid-burst can reorder a few frames at the switch-over and briefly upsets the receiver. The same reshuffle happens again in reverse when the member returns, which is why implementations often wait before re-adding a recovered member. ## What the layers above see No topology change. The logical interface stays up, so spanning tree sees the same port, and routing sees the same interface. RFC 7130 (section 4) puts it directly: protocols such as OSPF have no insight into the members and treat the bundle as one bigger interface. Whether a routing cost is recomputed from the reduced bandwidth is an implementation choice, and most designs should not count on it. That blindness is the problem. Capacity is now 30G, and if the bundle ran at 32G it is now overloaded: - 40G x 0.75 = 30G, so any load above **75 %** of the original capacity no longer fits. - After a second failure, 20G remains, and anything above **50 %** of the original no longer fits. ## Minimum links A **minimum-links** threshold, an implementation feature, says: if fewer than N members are active, declare the whole bundle down. The logic is that a bundle too thin to carry its load is worse than no bundle, provided an **alternative path** exists. 1. Set N to the smallest member count that still carries the expected load. For a bundle carrying 25G across four 10G members, N = 3. 2. When the active count drops below N, the logical interface goes down. 3. Routing withdraws the routes over it, a gateway-redundancy protocol tracking it lowers its priority, or spanning tree picks another port, and traffic moves to a path with full capacity. The trade-offs: - **No alternative path, no minimum links.** A server with one bundle and nowhere else to go is better served by a degraded 20G than by a dead link. Setting N there turns a partial failure into an outage. - **N too high** makes the bundle flap down on every single-member event. - **N too low** leaves the congestion problem in place. - **Both ends should agree.** If only one end enforces the threshold, the other may keep hashing traffic onto members of a bundle its partner has abandoned; with LACP, the downed end stops Distributing and its partner follows, while a static bundle has no such signal. ## Operational checklist - Alert on member count, not only on the logical interface state, because the logical state hides partial failures. - Size each bundle so the load still fits after one member fails, or set minimum links where a better path exists. - Prefer LACP's fast timers on bundles where a one-way member fault must clear within seconds.
- Why can a single member failure move flows that were on healthy members?If the implementation picks a member by taking the hash modulo the number of active members, changing that number from four to three changes the result for most hash values. Of twelve equally likely values, only three map to the same member under both moduli, so nine move, six of them needlessly. Table-based mappings that reassign only the failed member's buckets avoid this.
- When is setting a minimum-links threshold a mistake?When the bundle is the only path. Minimum links turns a degraded bundle into a down one so traffic moves elsewhere; with nowhere else to go, it turns a capacity loss into a full outage. It is also a mistake when set so high that every single-member blip takes the bundle down.
saying these in an interview costs you the question
- Routing protocols reroute as soon as one bundle member fails
- A member failure only ever moves that member's own flows
- A static bundle detects a member that keeps link but drops frames
- Minimum links should be set on every bundle, even a server's only uplink
- Losing one of four members cannot overload a bundle running below 100 %