skip to content

On a congested router queue carrying many TCP flows, why does tail drop cause global synchronisation, and how do RED and weighted RED avoid it?

level: seniorimportance: should knowfreq 19%

answer

  1. drops only when the buffer is full
  2. many flows lose packets together
  3. smoothed average, two thresholds
  4. linear drop probability ramp
  5. per drop precedence profiles

basics

~20 s

Tail drop discards only when a queue is full, so many flows lose packets at once, slow down together and leave the link idle. RED drops randomly and early as the average queue grows; weighted RED sets thresholds per drop precedence.

solid answer

~50 s

With **tail drop** a shared queue discards only when full, so a burst of overflow hits many flows at the same moment. Each TCP sender reads its loss as congestion and slows down; they slow together, the link goes underused, they speed up together and overflow again. RFC 7567 calls this global synchronisation of flows. **RED** tracks a smoothed queue average and, between a minimum and a maximum threshold, drops arriving packets with a probability rising linearly to a configured maximum, so a few flows back off at a time and the average queue stays short. **Weighted RED** runs separate thresholds per drop precedence, so AF13 is shed before AF12 and AF12 before AF11, as RFC 2597 requires. ECN-capable traffic can be marked instead of dropped (RFC 3168). Keep it off the voice priority queue, and note that RFC 7567 no longer recommends RED as the default algorithm.

go deeper

for a junior

Recall that tail drop discards only when a queue is full, while RED starts dropping a few packets early as the queue grows.

for a middle

Explain global synchronisation, RED's smoothed average with two thresholds and a linear ramp, and how weighted RED orders drops by drop precedence.

for a senior

Show where you enable weighted RED, where you keep it off, how ECN marking changes the picture, and how to read drop counters per precedence.

for a principal

Weigh buffer sizing, AQM choice and tuning burden across a fleet, given RFC 7567's move away from RED as the default toward self-tuning algorithms.

## Tail drop and its drawbacks The simplest queue has a maximum length; once it is full, arriving packets are discarded. This is **tail drop**, since the packet at the tail is the one lost. RFC 7567 (BCP 197, which replaced the recommendations of RFC 2309) lists four drawbacks: 1. **Full queues**: tail drop signals congestion only once the buffer is full, so queues stay full for long periods and every packet waits behind them. 2. **Lock-out**: a single connection or a few flows can monopolise the queue space and starve others. 3. **Packet bursts**: a large burst delays other packets and disrupts the senders' pacing. 4. **Control-loop synchronisation**: flows sharing a bottleneck can fall into step, causing periodic disruption. ## Global synchronisation When a shared queue overflows, many flows lose packets in the same instant. Each congestion-controlled TCP sender treats its loss as congestion and cuts its sending rate (how much and how it recovers is a topic of TCP's own congestion control). Because they all cut at once, the link goes underused; because they all ramp up again together, the queue overflows together again. RFC 7567 describes bursts of loss producing a global synchronisation of flows throttling back, followed by a sustained period of lowered link utilisation. The result is the worst of both: a queue that swings between empty and full, high delay at the peaks and wasted capacity in the troughs. ## Random Early Detection **RED** drops a few packets early and at random, before the queue is full, so flows back off one at a time. RFC 2597 requires these properties of an Assured Forwarding dropping algorithm and gives RED as its example: 1. Track a **smoothed** (averaged) congestion level, so short bursts are queued rather than punished. 2. Below a **minimum threshold**, drop nothing. 3. Between the minimum and the **maximum threshold**, drop arriving packets with a probability rising linearly from zero to a configured maximum. 4. Above the maximum threshold, drop every arriving packet of that precedence. Because the drops are random they land on different flows at different moments, which breaks the synchronisation. Because they start early, the average queue stays short, which lowers delay. RFC 2597 also requires the discard rate of a flow's packets within one precedence to be proportional to that flow's share of the traffic, so the heaviest flows absorb most of the drops. ## Weighted RED and drop precedence **Weighted RED (WRED)**, the usual implementation name, runs several RED profiles inside one queue, chosen by drop precedence or by codepoint. RFC 2597 requires that the dropping parameters be independently configurable for each drop precedence and each AF class, and that a packet with a lower drop precedence is never forwarded with smaller probability than one with a higher drop precedence. An illustrative profile for a backup class (example values, not from any standard): | Codepoint | Minimum threshold | Maximum threshold | Maximum drop probability | |---|---|---|---| | AF11 | 30 packets | 40 packets | 10 % | | AF12 | 25 packets | 40 packets | 10 % | | AF13 | 20 packets | 40 packets | 10 % | Traffic a policer re-marked to AF12 or AF13 as excess starts being shed at a shallower average queue, protecting in-profile AF11. ## Marking instead of dropping RFC 3168 lets a router that would drop a packet to signal incipient congestion instead set the Congestion Experienced codepoint in the packet's ECN field, when the packet is ECN-capable. RFC 7567 frames AQM as dropping or ECN-marking packets before a queue becomes full. The sender slows down without a packet being lost. RFC 3168 forbids marking in place of a drop that is made for reasons other than congestion, such as an edge discarding a class it does not admit. ## Where not to use it, and what RFC 7567 changed - **The voice priority queue**: RFC 4594's summary table marks the Telephony class with no active queue management. Voice payloads do not slow down when packets are lost, so early drops only add loss; policing and admission control keep that queue short instead. - **Unresponsive traffic**: RFC 7567 notes that AQM reduces drops only where end-to-end congestion control dominates the traffic. - **As a default algorithm**: RFC 7567 still strongly recommends deploying AQM but no longer recommends RED, or any specific algorithm, as the default. RED's parameters proved hard to set well, so it was rarely enabled by default; RFC 7567 asks for algorithms that tune themselves for common deployments.

  • If tail drop can cause synchronisation, why not simply make the buffer much larger?
    A larger buffer delays the first drop but keeps the queue standing full for longer, adding delay to every packet, including interactive ones. RFC 7567 says queue limits should reflect the bursts a device must absorb, not a steady-state queue, and that normally small queues can give both higher throughput and lower delay. Early, random drops address the cause instead.
  • Why does RED drop packets randomly rather than from the largest flow?
    Picking a flow needs per-flow state that a queue manager does not keep. Random drops achieve a similar effect statistically: a flow holding more of the queue's packets is more likely to be hit, and RFC 2597 requires a flow's discard rate within one precedence to be proportional to its share. Randomness also spreads drops over time, which desynchronises the senders.

saying these in an interview costs you the question

  • RED drops packets only once the queue is completely full.
  • Weighted RED drops AF11 before AF13 because AF11 is the lower value.
  • Larger buffers are the fix for global synchronisation.
  • WRED should be enabled on the voice priority queue to protect calls.
  • RFC 7567 makes RED the mandatory default AQM algorithm.
  • RED reacts to every instantaneous spike in queue length.