skip to content

A TCP bulk transfer across per-packet load-balanced parallel links shows frequent fast retransmits although the path drops almost nothing: what in TCP's acknowledgment behaviour explains this?

level: seniorimportance: nice to knowfreq 18%

answer

  1. a duplicate ACK has several causes
  2. overtaken looks the same as lost
  3. displaced by three or more
  4. the receiver saw it twice

basics

~20 s

A TCP receiver sends an immediate duplicate ACK for each segment arriving above a gap, and a reordered segment leaves the same gap as a lost one. Overtaken by three or more segments, it triggers fast retransmit of data that was only late.

solid answer

~50 s

RFC 5681 has the receiver send an immediate duplicate ACK whenever a segment arrives above a gap, and it names three causes of duplicates: dropped segments, **reordering**, and replication of ACKs or data. The sender cannot tell them apart from the ACK alone. Per-packet load balancing sends consecutive segments over links with different delays, so a segment is often overtaken by several later ones; each overtaking segment produces a duplicate ACK. Once three arrive, fast retransmit resends a segment that was merely late, and the sender cuts its rate as if there had been loss. When the original and the retransmission both reach a receiver that implements **D-SACK** (RFC 2883), it reports the duplicate range in its first SACK block, telling the sender the retransmission was spurious. RFC 8985's RACK replaces duplicate counting with a time-based reordering window.

go deeper

for a junior

Recall that a duplicate ACK repeats the same acknowledgment number because data arrived after a gap, and that a gap can come from delay as well as loss.

for a middle

Explain RFC 5681's three causes of duplicate ACKs and why a segment overtaken by three others triggers fast retransmit even though nothing was lost.

for a senior

Diagnose the pattern from the protocol: many retransmissions, few drops, D-SACK reports of duplicated data, and a path spreading one flow's packets across links.

for a principal

Weigh fixing reordering in the network, by keeping flows on one path, against relying on reordering-tolerant loss detection at every endpoint.

## The symptom A long transfer runs over a path where a device spreads **individual packets** of one flow across several parallel links. Loss on the path is close to zero, yet the sender performs frequent fast retransmits and its throughput is far below what the links can carry. The cause lies in what a TCP **duplicate acknowledgment** can and cannot say. ## How duplicate ACKs are generated RFC 5681 gives the receiver's rules: - Out-of-order data segments SHOULD be acknowledged **immediately**, not delayed. - When a segment arrives **above a gap** in the sequence space, the receiver SHOULD send an immediate duplicate ACK, repeating the acknowledgment number it last sent. - When a segment fills all or part of a gap, it SHOULD send an immediate ACK. With SACK in use (RFC 2018), each of these duplicate ACKs SHOULD also carry a SACK option whose first block covers the segment that just arrived. ## Why the sender cannot tell reordering from loss RFC 5681 lists three network events that produce duplicate ACKs: 1. **Dropped segments**: every segment after the lost one triggers a duplicate until the loss is repaired. 2. **Reordering** of data segments, which RFC 5681 notes is not rare along some paths. 3. **Replication** of ACK or data segments by the network. At the moment the duplicates arrive they look identical. RFC 8985 states the consequence directly: the sender can only distinguish reordering from loss **in hindsight**, if the missing range is filled later without a retransmission. ## What goes wrong on a reordering path Suppose segment S travels the slower of two links and four later segments overtake it. | Arrival order | Receiver's ACK | Sender's view | |---|---|---| | S+1 | duplicate, SACK S+1 | 1st duplicate | | S+2 | duplicate, SACK S+1..S+2 | 2nd duplicate | | S+3 | duplicate, SACK S+1..S+3 | 3rd duplicate: fast retransmit of S | | S+4 | duplicate | recovery already under way | | S (original) | ACK jumps past S+4 | the loss was never real | | S (retransmission) | ACK carrying a D-SACK block for S | the retransmission was unnecessary | The fast retransmit rule of RFC 5681 treats **three duplicate ACKs** as a loss signal. RFC 8985 describes exactly this failure: when the reordering degree exceeds that threshold, duplicate counting causes a **spurious fast recovery** and an unnecessary reduction of the congestion window. The sender pays twice, once for resending data and once for slowing down; the size of that slowdown belongs to the congestion-control topic. ## How the protocol exposes the mistake - **D-SACK** (RFC 2883): when the retransmitted copy of S arrives after the original, the receiver has the same bytes twice. It reports that duplicate range in the **first SACK block** of its ACK. The sender can then infer that the retransmission was spurious. D-SACK needs no negotiation beyond SACK-Permitted. - **RACK** (the RACK-TLP loss detection of RFC 8985, Standards Track): instead of counting duplicates, the sender marks a segment lost when a segment sent **later** has been delivered and the original has still not been acknowledged after roughly a round-trip time plus a **reordering window**. That window adapts: it SHOULD grow when D-SACK shows a spurious retransmission, and it MUST be bounded, with SRTT as the recommended bound. How the sender repairs its rate after a spurious retransmission, and the timers involved, belong to the retransmission and congestion-control topics. ## Diagnosing it - Look at the duplicates the receiver reports: a high rate of **D-SACK** blocks means data is arriving twice, which points at reordering or premature retransmission rather than loss. - Compare the retransmission rate with the actual drop rate on the path. Many retransmissions and almost no drops is the reordering signature. - Check whether the path spreads packets of one flow across links. Keeping each flow on one link, for example by choosing the link from the flow's addresses and ports, removes the reordering at its source. ## Misconceptions - A duplicate ACK does **not** prove a loss; it proves only that data arrived above a gap, or that something was duplicated. - Reordering is **not** harmless just because every byte eventually arrives: the retransmission and the rate cut happen before the late segment lands.

  • How does a TCP sender learn that one of its retransmissions was unnecessary?
    Through D-SACK (RFC 2883). If both the original and the retransmitted copy reach the receiver, the second arrival duplicates bytes it already holds, and the receiver reports that range in the first SACK block of its ACK. Because the sender knows it sent those bytes twice, a D-SACK tells it both copies arrived, so the retransmission was spurious. Without SACK the sender has no such evidence.
  • Why not just raise the duplicate-ACK threshold on paths that reorder?
    A higher threshold delays recovery from real losses and needs more segments in flight to trigger at all. RFC 8985 notes that adapting the threshold, for instance to half the FlightSize, still fails for some reordering patterns, such as a short last segment of an application-limited flight arriving first. It therefore uses a time-based reordering window, grown when D-SACK reveals a spurious retransmission.

saying these in an interview costs you the question

  • A duplicate ACK always means a segment was lost in the network.
  • Reordering cannot cause retransmissions because every byte eventually arrives.
  • Fast retransmit fires on the first duplicate ACK.
  • D-SACK needs its own option negotiated separately in the SYN.
  • An ACK that only updates the advertised window counts as a duplicate ACK.