skip to content

After a spanning-tree reconvergence, why do some silent hosts stay unreachable for a while, and why does the switched network flood unicast traffic?

level: seniorimportance: should knowfreq 15%

answer

  1. the tree moved, the tables did not
  2. talkers heal themselves, the silent do not
  3. flush trades a blackhole for a flood
  4. edge ports never trigger the flush

basics

~20 s

Switches still map MAC addresses to the old path. Hosts that transmit are relearned at once; frames to silent hosts go the wrong way until entries are aged out or flushed, and after a flush unlearned destinations are flooded.

solid answer

~50 s

A switch forwards known unicast by its MAC table, and spanning tree reconvergence moves the path without telling the tables. Entries learned before the change still point towards the old path. A host that transmits fixes this at once, because every switch on the new path relearns its source address. A **silent host**, such as a printer that only answers or a server that waits for requests, stays unreachable until its stale entries are removed. Normal ageing would take up to 300 s, the default 802.1D-1998 recommends (RFC 4188). Spanning tree therefore signals a topology change after reconvergence, and 802.1D then ages entries after one Forward Delay while RSTP flushes them immediately. The price is flooding: once entries are flushed, every frame to a not-yet-relearned host goes out of every forwarding port in the VLAN.

go deeper

for a junior

Remember that after spanning tree reroutes, switches must relearn where hosts are. Hosts that talk are found at once; silent ones can be briefly unreachable.

for a middle

Explain the stale entry, the topology-change flush (Forward Delay ageing in 802.1D, an immediate flush in RSTP) and why flushing causes a burst of unknown-unicast flooding.

for a senior

Diagnose the pattern: rising topology-change counters, a host port without edge status, a silent host's brief blackhole, and a VLAN-wide flood that is not a loop.

for a principal

Treat flood exposure as a reason to keep layer-2 domains small. Large VLANs multiply the cost of every flush, and a routed design confines it.

## What the tables believe after the tree moves An Ethernet switch forwards a unicast frame by looking up its destination MAC address in the **forwarding table**, which the IETF's Bridge MIB, RFC 4188, calls the Forwarding Database. Entries are learned from source addresses and age out after a period of silence. RFC 4188 notes that 802.1D-1998 recommends a default ageing time of **300 seconds**. Spanning-tree reconvergence changes which ports forward, but it does not rewrite those entries. After a failover: - switches on the **old path** may still hold entries pointing towards a link that is now down or blocked; - switches that **saw no failure at all** still send frames for hosts behind the changed segment out of the port they learned them on; - the **new path** has no entries yet for the hosts it now carries. ## Two symptoms from one cause | Host behaviour | What happens after the tree moves | How it ends | |---|---|---| | Transmits regularly | Its first frame teaches every switch on the new path where it is | Heals itself within one frame | | Silent (answers only, or idle) | Frames towards it follow stale entries and are dropped at the dead or blocked port | Only when the stale entry is aged out or flushed | | Any host, after a flush | Its address is unknown, so frames to it are flooded on every forwarding port in the VLAN | When it next transmits and is relearned | So the same event can produce a **blackhole** for some hosts and a **flood** for the whole VLAN, often one after the other. ## Why spanning tree flushes, and how fast If switches waited for normal ageing, a silent host could stay unreachable for up to 300 s after the tree had already healed. Spanning tree therefore follows a reconvergence with a **topology change**. How the topology change travels and how long it stays set are the mechanism of the BPDU and its flags; what matters here is its effect on the tables: 1. Under **802.1D** (the IEEE standard), bridges temporarily age dynamic entries using **Forward Delay** instead of the normal ageing time. RFC 4188 says Forward Delay "is also used when a topology change has been detected and is underway, to age all dynamic entries". A stale entry for a silent host lives for up to about 15 s at the IEEE default, not 300 s. 2. Under **RSTP** (IEEE 802.1w, folded into 802.1D-2004), bridges **flush** affected entries as soon as the change is signalled, so stale entries disappear almost immediately. ## The price: flooding A flushed table treats every destination as unknown. Until each host is heard again, switches flood frames for it on every forwarding port in its VLAN: - a large VLAN with busy flows (backups, storage replication, video) briefly sends that traffic to every access port; - slow links, such as a wireless bridge or a remote site, can congest under traffic that was never meant for them; - hosts receive and discard frames addressed to other machines. The flood is temporary and is not a loop. It ends host by host as each one transmits. ## Edge ports keep flushes rare A topology change, and therefore a network-wide flush, should follow real changes in the tree. Under RSTP only a **non-edge** port moving to `forwarding` signals one; an **edge port** (a host-facing port; one vendor calls it PortFast) never does, whether it comes up or goes down. Under 802.1D, a port moving into or out of forwarding signals a change, and implementations commonly exempt their edge-port feature from this. When host ports are not marked as edge ports, every laptop being unplugged, docked or rebooted can trigger a flush across the VLAN. A building of desks arriving at nine o'clock then produces a steady stream of flushes and floods that looks like random slowness. ## Diagnosing it 1. **Check the counters.** RFC 4188 exposes how many topology changes a bridge has detected and the time since the last one. A counter that keeps climbing on a stable network means some non-edge port is flapping. 2. **Find the source.** Implementations usually record which port last triggered a change. Follow it to a host port that lacks edge status, or to a flapping link. 3. **Separate blackhole from flood.** A silent host unreachable for a few seconds after a failover is the stale-entry window. A VLAN-wide traffic burst at the same moment is the flush. 4. **Do not "fix" it by shortening ageing permanently.** An ageing time of a few seconds removes stale entries faster but floods constantly, every time a host pauses.

  • Why not flush every MAC table whenever any port changes state, to be safe?
    Every flush turns all known destinations into unknown ones, so the VLAN floods until each host is relearned. Host ports coming and going do not change the path between other hosts, which is why edge ports are exempt from signalling a topology change. Flushing on every port event would trade a rare, short blackhole for constant flooding.
  • A host behind the changed path still cannot be reached a minute after failover. What do you suspect?
    Spanning tree's own flush should have cleared stale entries within Forward Delay (802.1D) or at once (RSTP). A minute points elsewhere: either the topology change never reached the switch holding the stale entry, or the fault is not in layer-2 forwarding at all. Check the topology-change counter and the time since the last change on each switch along the path.

saying these in an interview costs you the question

  • Once every port is forwarding again, every host is reachable at once
  • Stale MAC entries always live for the full 300-second ageing time
  • Unicast flooding after a failover means a loop has formed
  • A host port going up or down never affects other hosts' traffic
  • Permanently shortening MAC ageing to seconds is a harmless fix