skip to content

In a switched Ethernet LAN, a whole floor loses connectivity with line-rate broadcast on the uplinks, pegged switch CPUs and one MAC flapping between two ports; how do you confirm a Layer 2 loop and stop it?

level: seniorimportance: should knowfreq 22%

answer

  1. symptoms that arrive together
  2. the same frame, over and over
  3. management plane may be unreachable
  4. cut a link on the loop
  5. then ask why nothing blocked it

basics

~20 s

Line-rate broadcast plus a flapping MAC means a Layer 2 loop. Confirm it from flap logs and identical repeated frames, cut a link on the loop to stop it at once, then find why spanning tree did not block it.

solid answer

~50 s

Those three symptoms together are a Layer 2 loop. Confirm it: broadcast and multicast counters at line rate on many ports at once, MAC-move logs naming the ports the loop passes through, and a capture showing the **same** frame (an identical ARP request, say) repeating thousands of times. Expect in-band management to be slow or unreachable, so use a console or out-of-band path. To stop it, shut a port that lies **on** the loop: an inter-switch or recently patched port that both receives and sends the storm, never the flapping host's own access port, which is a victim. The storm drains almost at once. Then find the root cause: where spanning tree was off or defeated, such as a stray cable through a device that does not pass BPDUs, a host bridging two interfaces, a one-way link, or a CPU too busy to process BPDUs. Broadcast rate limiting caps the damage but never removes a loop.

go deeper

for a junior

Recognise the signature: line-rate broadcast, pegged switch CPUs and a MAC moving between ports all at once mean a switching loop. The fix is to cut a link in the loop.

for a middle

Explain how each symptom follows from flooding and source learning, and why the flapping host's own port is the wrong one to shut.

for a senior

Show the operational sequence: out-of-band access, flap logs, counters, a short capture, cutting a link that is on the loop, then the root cause, including the storm starving the CPU of BPDUs.

for a principal

Turn the incident into controls: edge-port protections, broadcast limits, MAC-move alerting, out-of-band management and smaller Layer 2 domains, each matched to a root cause actually seen.

## Recognise the signature A Layer 2 loop has a recognisable pattern, because its symptoms come from the same cause and arrive **together and suddenly**: | Observation | Why a loop produces it | |---|---| | Uplinks at line rate, mostly broadcast and multicast | Flooded frames circulate with no hop count to expire them, and every new broadcast adds more. | | Many access ports transmitting the same high broadcast rate | Every pass of the loop floods each copy to every host port in the VLAN. | | Logs of one MAC moving between two ports, many times a second | The switch records each frame's source against its arrival port, so looped copies of one host's frames keep overwriting the entry. | | Switch management CPUs pegged | Broadcasts addressed to the switch itself (ARP for its management address, for example) and constant MAC moves land on the control plane. | | Hosts sluggish or disconnected across the whole VLAN | Their network stacks process every broadcast copy, and unicast for flapping hosts is misdirected. | A routing loop looks different: it hits particular destinations, the traffic is unicast, and each packet dies when its `TTL` runs out. A failing uplink or a congested server does not make MAC entries flap. ## Confirm it 1. **Get a working management path.** In-band access through the looped VLAN may time out because the CPU is drowning. Use a console or out-of-band management network. 2. **Read the MAC-move or flap logs.** They name the switch and the ports the loop passes through. If the flapping address belongs to a host, one of the two ports may be that host's own access port; the other is on the loop. 3. **Compare port counters.** Find the ports carrying the storm *in both directions* at high rate. An inter-switch link or an access port that is *receiving* a flood of broadcasts is a strong suspect, because a normal access port sends very few. 4. **Capture briefly.** A short capture showing one identical frame (same source, same content) repeating thousands of times confirms circulation rather than a chatty host. 5. **Check spanning tree on the suspect switches.** Is it running on that VLAN? Did a port that should be blocking start forwarding? Is a port receiving BPDUs that should not, or not receiving BPDUs that should? ## Stop it - **Cut a link that lies on the loop.** Shut the suspect inter-switch or recently patched port. Cutting any link in a cycle breaks that cycle, and the circulating copies drain out the access ports almost immediately. - **Do not shut the flapping host's access port.** The host is a victim whose frames are being looped. Shutting its port disconnects it and leaves the loop running. - If the culprit is unclear, shut **suspect redundant ports one at a time** and watch the broadcast counters drop. - **Broadcast rate limiting** on the ports can cap the damage while you work, but it never ends a loop: the copies that pass the limit keep circulating. ## Find out why spanning tree did not prevent it A loop survives only where spanning tree is absent or blind. Common root causes: - **Spanning tree off** on a switch or VLAN, or BPDU processing disabled on a port, so the redundant path was never seen. - **A device that does not pass BPDUs** in the loop; a small switch patched back through two wall ports, or a host or virtual machine bridging two interfaces, can behave this way. RFC 6325 names lost spanning-tree messages and added repeaters as causes of temporary loops. - **A one-way link**: BPDUs stop reaching a blocked port, its stored information expires, and it starts forwarding while the link still carries data the other way. - **An overloaded control plane**: once a storm pegs a CPU, the switch can miss BPDUs, so a blocked port expires its information and opens, a second loop that keeps the CPU pegged. Record which one it was, because each points to a different preventive control at the access edge, such as BPDU Guard on host-facing ports or Loop Guard where a one-way link could open a blocked port (both are implementation features, not names in the IEEE standard). ## Afterwards - Re-enable the cut link only once spanning tree is running across it and you have seen it block correctly. - Confirm the MAC tables have settled: the flapping host is learned on its own port again. - Add monitoring for MAC-move events and broadcast rate, so the next loop is caught in seconds rather than by users.

  • Why can a switch's overloaded CPU during an Ethernet broadcast storm make the loop worse?
    BPDUs are processed by the switch's CPU. If the storm saturates it, BPDUs are missed, the stored spanning-tree information on a blocked port expires, and the port starts forwarding. That opens another loop, which feeds more broadcasts to the CPU. The storm sustains itself until a link is cut by hand.
  • During a Layer 2 storm, why is a capture showing the same frame repeated more convincing than high broadcast counters alone?
    High broadcast counters can also come from a misbehaving host, a scanning tool or a burst of DHCP clients, all of them distinct frames from real senders. A loop replays identical frames, with the same source, content and checksum, thousands of times, which only circulation explains.

saying these in an interview costs you the question

  • Shut the access port of the host whose MAC is flapping.
  • Clearing the MAC address table will stop the storm.
  • Broadcast rate limiting fixes a loop, so the cable can stay.
  • A storm this size must be a routing loop or a DDoS attack.
  • Spanning tree was running, so a Layer 2 loop is impossible.