A spanning-tree network shows its topology-change counter climbing all day and constant unicast flooding; how do you find the source and stop it?
answer
- the counter says when, not where
- follow the reports upstream
- overlapping 35-second windows
- host ports count as changes
basics
~20 sRecurring topology changes keep every bridge on 15-second MAC ageing, flooding frames to silent hosts. Trace the reports hop by hop to the bridge whose own port keeps changing, then fix it: a flapping link, or a host port that should be edge.
solid answer
~50 sEach change makes the 802.1D root set TC for 35 s at the IEEE defaults, and changes arriving closer together than that overlap, so bridges never leave **15 s ageing** and keep flooding frames to hosts silent for longer. The Bridge MIB (RFC 4188) shows *that* it happens: `dot1dStpTopChanges` keeps rising and `dot1dStpTimeSinceTopologyChange` stays small. To find *where*, walk toward the source: on each bridge, find the port the change arrived on (most implementations record it), move to the neighbour on that port, and repeat until you reach the bridge whose own port is transitioning. There, a rising per-port `dot1dStpPortForwardTransitions` count, or the optional `topologyChange` notification, names the port. Typical culprits are a flapping inter-switch link and host-facing ports going up and down. Fix the link; declare host ports **edge ports**, which RSTP does not count as changes.
go deeper
Recall that a topology change makes switches flood more, and that a port going up or down can cause one.
Explain why changes closer together than 35 s keep the tree on 15 s ageing and what the topology-change counters show.
Lead the hunt: walk the reports back to the originating bridge and port, separate a flapping link from noisy host ports, and fix the cause rather than the symptom.
Turn the incident into policy: edge ports on every host port, RSTP everywhere, alerting on the change rate, and smaller layer-2 domains so one bad port cannot flood the whole site.
## The symptom and its mechanism Users report sluggish, bursty performance; packet captures on access ports show unicast traffic for other hosts; switch monitoring shows the spanning-tree **topology change count** rising through the day. The mechanism is the topology-change machinery of **IEEE 802.1D** (an IEEE standard; the IETF's Bridge MIB, RFC 4188, restates its counters): 1. a bridge whose port changes state sends a TCN toward the root; 2. the root sets the **TC** flag in its Configuration BPDUs for **Max Age + Forward Delay**, 35 s at the IEEE defaults; 3. every bridge receiving TC ages its MAC entries after **Forward Delay**, 15 s, instead of the usual ageing time, 300 s as 802.1D-1998 recommends; 4. any host silent for more than 15 s loses its entry, and frames to it are **flooded** across the VLAN. One change is a short flood. The trouble is repetition. ## Why recurring changes mean constant flooding The arithmetic is simple. If topology changes arrive more often than every 35 s, each new TC period begins before the last one ends, so the tree **never leaves short ageing**. | Interval between changes | Share of time on 15 s ageing (approx.) | |---|---| | every 10 minutes | 35 / 600, about 6 % | | every 2 minutes | 35 / 120, about 29 % | | every 30 s | windows overlap: all of it | A server room where several hosts reboot, or one link that flaps every few seconds, is enough. The flooding also reaches every bridge in that spanning tree, not just the neighbourhood of the fault. ## Proving it is happening The Bridge MIB gives the evidence, polled from any bridge: - `dot1dStpTopChanges` — the number of topology changes the bridge has detected; a steady climb is the signature; - `dot1dStpTimeSinceTopologyChange` — time since the last one; if it never gets large, changes are continuous; - `dot1dStpPortForwardTransitions` — per port, the number of moves from learning to forwarding; a port with a climbing count is changing state repeatedly. These tell you **when** and, at the right bridge, **which port**. They do not tell a distant bridge where the change began: the TCN and the TC flag carry no bridge or port ID. ## Finding the source Work toward the bridge that originates the changes: 1. Start at any bridge, often the root, and find the port on which the most recent topology-change report arrived. Recording that port, and the sender, is an implementation feature most switches offer. 2. Go to the neighbour on that port and repeat. Under 802.1D the TCNs travelled rootward along this exact path, so walking it backwards leads to the origin. 3. Stop at the bridge where the change comes from one of its **own** ports. Its per-port forward-transition count, or RFC 4188's optional `topologyChange` notification, sent "when any of its configured ports transitions from the Learning state to the Forwarding state, or from the Forwarding state to the Blocking state", names the port. 4. Correlate with that port's link up and down history. ## Common culprits and fixes - **A flapping inter-switch link** — a failing optic, a bad cable, a duplex mismatch. Fix the physical layer; each flap is a genuine change. - **Host ports that go up and down** — rebooting servers, laptops docking, power-saving network cards. Under classic 802.1D each one is a topology change unless the port is an edge port. Declare host-facing ports **edge ports**: under RSTP an edge port does not count as a change when it reaches forwarding. How edge ports and their guards are configured is a separate subject. - **A neighbour bouncing between roles** — for instance an unstable link toward a bridge that keeps winning and losing a designated role. Stabilise the link or the priorities. What not to do: setting the normal ageing time permanently low only makes flooding permanent; lowering timers on a non-root bridge has no effect at all. ## Where RSTP helps Under RSTP (IEEE 802.1w, now in 802.1D-2004) a port going down or to discarding is not a topology change, only a non-edge port moving to forwarding is, and edge ports never count. Moving to RSTP and marking host ports as edge removes most of the noise; what is left points at real instability.
- Why can't a distant bridge tell from the TC flag where the change started?Under 802.1D the TCN is a 4-octet header with no bridge or port ID, and the root announces TC in its own Configuration BPDUs. Every bridge learns only that a change happened. Finding the origin means walking the path the TCNs took, bridge by bridge, using the arrival port that most implementations record.
- The topology-change counters are flat, yet unicast flooding continues. What else could explain it?Constant flooding has other causes: a MAC table that is full, asymmetric paths where return traffic never passes a switch so it never learns the source, or a very low normal ageing time. Check the table's size and the Bridge MIB's ageing time before blaming spanning tree.
saying these in an interview costs you the question
- Topology changes come only from failing links between switches.
- The TC flag tells every bridge which port caused the change.
- Lowering the normal MAC ageing time to 15 s for good is a harmless fix.
- Frequent topology changes are cosmetic as long as the root bridge stays the same.
- Under RSTP, any port going down floods a topology change through the network.