In a switched Ethernet LAN, a whole floor loses connectivity with line-rate broadcast on the uplinks, pegged switch CPUs and one MAC flapping between two ports; how do you confirm a Layer 2 loop and stop it?
answer
- symptoms that arrive together
- the same frame, over and over
- management plane may be unreachable
- cut a link on the loop
- then ask why nothing blocked it
basics
~20 sLine-rate broadcast plus a flapping MAC means a Layer 2 loop. Confirm it from flap logs and identical repeated frames, cut a link on the loop to stop it at once, then find why spanning tree did not block it.
solid answer
~50 sThose three symptoms together are a Layer 2 loop. Confirm it: broadcast and multicast counters at line rate on many ports at once, MAC-move logs naming the ports the loop passes through, and a capture showing the **same** frame (an identical ARP request, say) repeating thousands of times. Expect in-band management to be slow or unreachable, so use a console or out-of-band path. To stop it, shut a port that lies **on** the loop: an inter-switch or recently patched port that both receives and sends the storm, never the flapping host's own access port, which is a victim. The storm drains almost at once. Then find the root cause: where spanning tree was off or defeated, such as a stray cable through a device that does not pass BPDUs, a host bridging two interfaces, a one-way link, or a CPU too busy to process BPDUs. Broadcast rate limiting caps the damage but never removes a loop.
go deeper
Recognise the signature: line-rate broadcast, pegged switch CPUs and a MAC moving between ports all at once mean a switching loop. The fix is to cut a link in the loop.
Explain how each symptom follows from flooding and source learning, and why the flapping host's own port is the wrong one to shut.
Show the operational sequence: out-of-band access, flap logs, counters, a short capture, cutting a link that is on the loop, then the root cause, including the storm starving the CPU of BPDUs.
Turn the incident into controls: edge-port protections, broadcast limits, MAC-move alerting, out-of-band management and smaller Layer 2 domains, each matched to a root cause actually seen.
## Recognise the signature A Layer 2 loop has a recognisable pattern, because its symptoms come from the same cause and arrive **together and suddenly**: | Observation | Why a loop produces it | |---|---| | Uplinks at line rate, mostly broadcast and multicast | Flooded frames circulate with no hop count to expire them, and every new broadcast adds more. | | Many access ports transmitting the same high broadcast rate | Every pass of the loop floods each copy to every host port in the VLAN. | | Logs of one MAC moving between two ports, many times a second | The switch records each frame's source against its arrival port, so looped copies of one host's frames keep overwriting the entry. | | Switch management CPUs pegged | Broadcasts addressed to the switch itself (ARP for its management address, for example) and constant MAC moves land on the control plane. | | Hosts sluggish or disconnected across the whole VLAN | Their network stacks process every broadcast copy, and unicast for flapping hosts is misdirected. | A routing loop looks different: it hits particular destinations, the traffic is unicast, and each packet dies when its `TTL` runs out. A failing uplink or a congested server does not make MAC entries flap. ## Confirm it 1. **Get a working management path.** In-band access through the looped VLAN may time out because the CPU is drowning. Use a console or out-of-band management network. 2. **Read the MAC-move or flap logs.** They name the switch and the ports the loop passes through. If the flapping address belongs to a host, one of the two ports may be that host's own access port; the other is on the loop. 3. **Compare port counters.** Find the ports carrying the storm *in both directions* at high rate. An inter-switch link or an access port that is *receiving* a flood of broadcasts is a strong suspect, because a normal access port sends very few. 4. **Capture briefly.** A short capture showing one identical frame (same source, same content) repeating thousands of times confirms circulation rather than a chatty host. 5. **Check spanning tree on the suspect switches.** Is it running on that VLAN? Did a port that should be blocking start forwarding? Is a port receiving BPDUs that should not, or not receiving BPDUs that should? ## Stop it - **Cut a link that lies on the loop.** Shut the suspect inter-switch or recently patched port. Cutting any link in a cycle breaks that cycle, and the circulating copies drain out the access ports almost immediately. - **Do not shut the flapping host's access port.** The host is a victim whose frames are being looped. Shutting its port disconnects it and leaves the loop running. - If the culprit is unclear, shut **suspect redundant ports one at a time** and watch the broadcast counters drop. - **Broadcast rate limiting** on the ports can cap the damage while you work, but it never ends a loop: the copies that pass the limit keep circulating. ## Find out why spanning tree did not prevent it A loop survives only where spanning tree is absent or blind. Common root causes: - **Spanning tree off** on a switch or VLAN, or BPDU processing disabled on a port, so the redundant path was never seen. - **A device that does not pass BPDUs** in the loop; a small switch patched back through two wall ports, or a host or virtual machine bridging two interfaces, can behave this way. RFC 6325 names lost spanning-tree messages and added repeaters as causes of temporary loops. - **A one-way link**: BPDUs stop reaching a blocked port, its stored information expires, and it starts forwarding while the link still carries data the other way. - **An overloaded control plane**: once a storm pegs a CPU, the switch can miss BPDUs, so a blocked port expires its information and opens, a second loop that keeps the CPU pegged. Record which one it was, because each points to a different preventive control at the access edge, such as BPDU Guard on host-facing ports or Loop Guard where a one-way link could open a blocked port (both are implementation features, not names in the IEEE standard). ## Afterwards - Re-enable the cut link only once spanning tree is running across it and you have seen it block correctly. - Confirm the MAC tables have settled: the flapping host is learned on its own port again. - Add monitoring for MAC-move events and broadcast rate, so the next loop is caught in seconds rather than by users.
- Why can a switch's overloaded CPU during an Ethernet broadcast storm make the loop worse?BPDUs are processed by the switch's CPU. If the storm saturates it, BPDUs are missed, the stored spanning-tree information on a blocked port expires, and the port starts forwarding. That opens another loop, which feeds more broadcasts to the CPU. The storm sustains itself until a link is cut by hand.
- During a Layer 2 storm, why is a capture showing the same frame repeated more convincing than high broadcast counters alone?High broadcast counters can also come from a misbehaving host, a scanning tool or a burst of DHCP clients, all of them distinct frames from real senders. A loop replays identical frames, with the same source, content and checksum, thousands of times, which only circulation explains.
saying these in an interview costs you the question
- Shut the access port of the host whose MAC is flapping.
- Clearing the MAC address table will stop the storm.
- Broadcast rate limiting fixes a loop, so the cable can stay.
- A storm this size must be a routing loop or a DDoS attack.
- Spanning tree was running, so a Layer 2 loop is impossible.