How does a multi-chassis link aggregation group (MLAG) let a server bundle links to two switches, and what happens when the switches' peer link fails?
answer
- two switches, one LACP identity
- the server thinks it has one partner
- peer link carries sync and spillover
- keepalive on a separate path
- secondary shuts its bundle ports
basics
~20 sTwo switches present one shared LACP System ID, so the server bundles links to both as to one partner. If their peer link fails, a separate keepalive shows the peer is alive and the secondary disables its MLAG ports, avoiding split-brain.
solid answer
~50 sLACP only bundles ports that lead to one partner System ID. An **MLAG** pair (one vendor calls its version vPC) makes two switches advertise the **same** System ID and key on the server-facing ports, so the server's LACP sees one partner and bundles all four links: two to each switch. The two switches stay separate control planes and keep each other consistent over a **peer link**, syncing learned MACs and port state and carrying traffic for devices attached to only one of them. A **keepalive** on a separate path tells a dead peer apart from a dead peer link. If the peer link fails while both switches live, the secondary disables its MLAG ports; the server's LACP drops those members and runs on the primary. Lose the keepalive too and both act alone: **split-brain**. None of this is standardised.
go deeper
Recall that MLAG lets one server bundle its links to two switches, so the server survives a whole switch failing, not just a cable.
Explain how a shared LACP System ID makes two switches look like one partner, and what the peer link carries between them.
Walk through peer-link loss, peer death and split-brain, and show how the keepalive and the secondary's shutdown keep forwarding consistent.
Weigh MLAG against stacking and against standards-based EVPN multihoming, considering vendor lock-in, upgrade independence and the two-switch scaling limit.
## The problem MLAG solves A server bundled to one top-of-rack switch survives a cable failure but not the switch failing. Bundling to **two** switches would survive both, but LACP, part of IEEE 802.1AX, assumes a bundle has a single partner: a port is aggregated only with others that report the same partner **System ID** and key. Two independent switches have different System IDs, so the server would refuse to bundle across them. A **multi-chassis link aggregation group** (MLAG, also written M-LAG; one vendor's version is called vPC) makes two switches cooperate so they look like one partner to LACP. RFC 7938 (section 4.1) describes it as an enhancement to link aggregation that allows active-active Layer 2 paths, and lists its downsides: in most implementations it cannot scale past two switches, there is no standards-based implementation, and syncing state between the devices adds failure-domain risk. ## How the pair presents one identity The scenario: a server with four 10G members, two to switch A and two to switch B. 1. A and B are configured as MLAG peers and agree on a shared **System ID** and key for the server-facing bundle. 2. Both send LACPDUs with that shared identity on their members. 3. The server sees one partner on all four ports and aggregates them into one 40G bundle; it hashes each flow onto one member, without knowing two chassis are involved. 4. To spanning tree, on the switches and on anything attached, the bundle is still one logical port, so none of its members is blocked. The standards-based counterpart does the same trick: EVPN multihoming (RFC 7432, section 5) can derive its Ethernet Segment Identifier from the CE's LACP System ID and port key, and states that the CE treats the multiple PEs it connects to as the same switch. That is a separate, EVPN-based design. ## The peer link and the keepalive The two switches keep independent control planes, so they must coordinate: | Element | Carries | Why it matters | |---|---|---| | Peer link | MAC and port-state sync, plus data for single-attached devices and failure cases | Keeps both switches forwarding consistently | | Peer keepalive | A small heartbeat over a separate path, often the management network | Distinguishes a dead peer from a dead peer link | | Loop-avoidance rule | Frames received over the peer link are not sent out an MLAG port whose local member is up | Prevents duplicates and loops, since the server is reachable directly | One switch is elected **primary** and the other **secondary**, by a priority set on each; the election and its tie-breaks are implementation-specific. ## Failure cases - **A server member fails.** The server's LACP drops it and rehashes; the switches carry on. - **Switch B fails.** Its members go dark, the server loses half its capacity (40G to 20G), and traffic continues on A. The keepalive also goes silent, so A knows B is dead, not just unreachable. - **The peer link fails, both switches alive.** The keepalive still arrives, so each knows the other is up but cannot sync. Continuing on both sides would let MAC tables and port state diverge. The usual rule: the **secondary disables its MLAG member ports** (often its routed interfaces too). Its LACPDUs stop, the server drops those two members, and everything runs through the primary at 20G. Devices attached only to the secondary may be cut off, which is why single-attached devices on an MLAG pair are a design smell. - **Peer link and keepalive both fail.** Neither switch can tell whether the other is alive. Both may stay active with the shared identity: this is **split-brain**. Learned addresses diverge, flooding can duplicate frames, and the server keeps hashing flows onto switches that no longer coordinate, risking loss and loops. Spanning tree, which RFC 7938 notes MLAG designs keep as the backup loop guard, may be the last line of defence. ## Design guidance - Build the **peer link** from at least two members on different line cards. - Run the **keepalive** over a path that shares nothing with the peer link. - Keep **single-attached** devices off MLAG pairs, or accept that the secondary's shutdown may isolate them. - Upgrade the two switches one at a time; separate control planes are what make that possible. ## MLAG versus stacking A switch **stack** joins several chassis into one logical switch with **one** control plane, so a bundle across members is just an ordinary local bundle. That avoids the peer-sync problem but shares fate: a control-plane fault or a software upgrade affects every member at once. MLAG keeps two control planes, which allows one switch to be upgraded while the other forwards, at the price of the sync machinery and its split-brain failure mode. Both are implementation features; neither is an IETF or IEEE standard.
- Why does MLAG need a keepalive on a separate path when it already has a peer link?Because a silent peer link has two explanations: the peer switch died, or only the link between them did. If the peer died, the survivor should keep all its ports up. If only the link died, both are alive and one must stand down to avoid inconsistent forwarding. The keepalive, on an independent path, tells the two cases apart.
- How does a switch stack differ from an MLAG pair as a way to bundle a server to two chassis?A stack makes the chassis one logical switch with one control plane, so the bundle is local and needs no peer sync, but a control-plane fault or upgrade hits every member. MLAG keeps two independent switches that share an LACP identity, which allows upgrading one at a time but adds a peer link, a keepalive and a split-brain failure mode.
- Why does the MLAG pair not forward a frame received over the peer link out of a bundle port whose local member is up?The server attached to that bundle is reachable directly from both switches. If a frame flooded by one switch crossed the peer link and the other switch sent it down its own member too, the server would receive it twice, and broadcasts could loop. The rule lets peer-link traffic reach a bundle only when the local member there is down.
saying these in an interview costs you the question
- MLAG is an IEEE standard that every switch vendor implements alike
- The server must run special software to bundle across two switches
- When the peer link fails, both switches should keep forwarding
- The peer link alone can tell a dead peer from a broken link
- Spanning tree blocks one switch's members of the MLAG bundle