skip to content

Why does losing the last segment of a TCP response often stall it for a whole RTO, while a mid-stream loss recovers in about one round trip?

level: seniorimportance: should knowfreq 28%

answer

  1. who produces the duplicate ACKs?
  2. nothing behind the hole
  3. fewer than three followers
  4. probe before the timer
  5. resend the highest segment

basics

~20 s

Fast retransmit needs later segments to produce duplicate ACKs; a tail loss has too few behind it, so only the RTO, normally at least 1 s (RFC 6298), recovers it. RACK-TLP (RFC 8985) instead probes after about 2 x SRTT.

solid answer

~50 s

Fast retransmit is driven by arrivals: each segment reaching the receiver after a gap produces a duplicate ACK, and the sender acts on the third. If the lost segment is the last one, or one of the last three, fewer than three segments follow it, so the third duplicate never comes. The sender's only remaining signal is the retransmission timer, which RFC 6298 normally floors at 1 s, far longer than a typical RTT; the timeout also cuts `cwnd` to one segment. Request-response traffic makes this common, because every response ends in a tail. **RACK-TLP** (RFC 8985, Standards Track, requires SACK) targets it: after about `2*SRTT` with no ACK, the sender sends a **tail loss probe**, new data if it has any, otherwise the highest-sequenced segment. The ACK for the probe exposes the missing segments, so fast recovery repairs them instead of an RTO.

go deeper

for a junior

Recall that fast retransmit needs segments after the lost one, so losing the final segment leaves TCP waiting for its timeout.

for a middle

Explain why fewer than three followers means no fast retransmit, and why an RTO costs so much more than one round trip.

for a senior

Diagnose tail-latency stalls on short responses from the protocol, and explain how RACK-TLP's probe after about two SRTTs converts them into fast recovery.

for a principal

Weigh enabling SACK and RACK-TLP across a fleet against its costs, such as an extra probe segment and scoreboard state, for request-response workloads.

## Fast retransmit depends on what comes after the loss A TCP receiver acknowledges **cumulatively**: its ACK names the next byte it expects. When one segment goes missing and later ones arrive, every later arrival makes the receiver repeat the same ACK, a **duplicate ACK**. Under RFC 5681, the sender treats the **third** duplicate ACK as evidence of loss and retransmits at once: **fast retransmit**. The repair therefore costs roughly one round trip. The key word is *later*. Duplicate ACKs are manufactured by segments that follow the hole. No followers, no duplicates. ## The tail case, traced Consider a response sent as ten segments: 1. **Segment 5 lost.** Segments 6 to 10 arrive, producing five duplicate ACKs. The third triggers fast retransmit; the loss is repaired about one round trip later. 2. **Segment 8 lost.** Only 9 and 10 follow it, so just two duplicate ACKs arrive. The threshold of three is never met. 3. **Segment 10 lost.** Nothing follows it. The receiver has nothing new to acknowledge, so the sender hears silence. In cases 2 and 3 the sender falls back to the **retransmission timer**. RFC 6298 computes the RTO from measured round-trip times and says it SHOULD be at least **1 second**, often many times the actual RTT. After the timeout the sender also resets its congestion window to one segment under RFC 5681. RFC 5681 recommends Limited Transmit (RFC 3042), sending new data on the first two duplicate ACKs to generate more of them, but at the end of a response there is usually no new data to send. ## What the stall costs RFC 8985 §3.2 gives a worked example: a sender transmits 100 segments and only the last three are lost. Without a probe, the RTO expires, the sender retransmits the first missing segment and slow-starts from a window of one. The whole transfer takes **three RTTs plus one RTO**. Had the same losses happened three segments earlier, fast recovery would have repaired them within one round trip and the window would not have been reset. The pattern bites hardest on: - **short request-response exchanges**, where the entire response is a tail; - **small windows**, where fewer than four segments are ever in flight; - **application-limited senders** that pause after each burst of data. ## The tail loss probe **RACK-TLP** (RFC 8985, Standards Track) adds a timer that fires well before the RTO. It **requires SACK** (RFC 2018), because it reads the receiver's selective acknowledgments to find holes. The probe timeout (PTO) is chosen as follows: - with an RTT estimate, `PTO = 2 * SRTT`, since an ACK normally arrives within one SRTT and exactly one would risk spurious probes; - when only one segment is in flight, the sender MAY add an allowance for a delayed ACK; - with no RTT estimate yet, `PTO = 1 second`, matching RFC 6298's initial RTO; - if the RTO would fire sooner than the PTO, the probe is sent at the RTO time instead. When the PTO fires, the sender SHOULD send **previously unsent data** if it has any and the receive window allows; otherwise it retransmits the **highest-sequenced segment** sent so far. At most one probe may be outstanding, and after sending one the sender re-arms the RTO timer, so the timeout remains the last resort. ## Why the probe works The probe gives the receiver something to acknowledge. Its SACK shows the receiver holds the end of the flight, which exposes every earlier hole. In the 100-segment example, the probe retransmits segment 100; the SACK for it lets RACK detect that 98 and 99 were lost and trigger fast recovery. Total time: **four RTTs**, with the congestion window only partly reduced rather than reset. If the probe is lost too, RTO recovery happens as before. | Situation | What repairs it | Typical cost | |---|---|---| | Loss with at least three segments behind it | Fast retransmit on duplicate ACKs | About one RTT | | Tail loss, no probe | Retransmission timer | One RTO, normally at least 1 s, plus slow start from one segment | | Tail loss with RACK-TLP | Tail loss probe, then fast recovery | About 2 x SRTT before the probe, then about one RTT | ## Misconceptions - Fast retransmit cannot repair every single loss in one round trip; it needs followers. - The probe does not wait for the RTO; it fires after about two smoothed RTTs. - RACK-TLP defines no new TCP option; it is sender-side logic that relies on SACK.

  • Why does a tail loss probe retransmit the highest-sequenced segment rather than the first unacknowledged one?
    An acknowledgment for the last segment, whether for its original copy or the probe, shows the receiver holds the end of the flight, which exposes every earlier hole to RACK at once. RFC 8985 chooses it to sidestep the retransmission ambiguity: it does not matter which copy the SACK answers. Retransmitting the first missing segment would reveal less.
  • What happens under RACK-TLP if the tail loss probe itself is lost?
    After sending a probe, RFC 8985 has the sender re-arm the retransmission timer rather than the probe timer. If no acknowledgment comes back, the RTO fires and the sender performs ordinary RTO recovery, resetting the congestion window. The outcome matches having no probe, at the cost of one extra segment sent.

saying these in an interview costs you the question

  • Fast retransmit can repair any single lost segment within one round trip
  • A tail loss stalls because the receiver is waiting on its delayed-ACK timer
  • The tail loss probe is sent only after a full RTO has expired
  • RACK-TLP needs a new TCP option negotiated during the handshake
  • Losing the last segment is so rare that tail stalls do not matter