skip to content

When a TCP path's round-trip time suddenly jumps from 80 ms to about 2 seconds, why does the sender retransmit data that was never lost, and how can it tell afterwards?

level: seniorimportance: nice to knowfreq 15%

answer

  1. late is not lost
  2. a timeout sized for yesterday's path
  3. whose clock value comes back?
  4. the receiver reports a duplicate
  5. undoing a needless cut

basics

~20 s

An RTO built from 80 ms samples sits at the 1 s floor, so a 2 s delay fires it while the originals are merely late. An ACK echoing the original's TSval, or a D-SACK for the duplicate, exposes the spurious timeout.

solid answer

~50 s

RFC 6298 builds the RTO from recent samples: with a steady 80 ms RTT the formula gives roughly 100 ms, which the 1 s floor lifts to 1 s. A sudden 2 s delay, from a link-layer retry on a radio link or a path change, outlasts that, so the timer fires on data that is still in transit. The sender retransmits the earliest unacknowledged segment, doubles the RTO and, under RFC 5681, cuts `cwnd` to one segment. When the late ACKs arrive it may even resend further segments it wrongly presumed lost. Detection is after the fact: with the **timestamp option** (RFC 7323), the first new ACK echoes the **original** transmission's `TSval`, proving the original arrived; with **D-SACK** (RFC 2883), the receiver reports the duplicate copy. F-RTO (RFC 5682) infers it from ACKs for never-retransmitted data. Once detected, the sender may undo the needless congestion-window cut.

go deeper

for a junior

Recall that a timeout can fire on data that is only delayed, so TCP sometimes resends segments that were never lost.

for a middle

Explain why an RTO computed from a steady 80 ms path cannot cover a sudden 2 s delay, and what the sender does when the timer fires.

for a senior

Walk the full cost of a spurious timeout and show how timestamp echoes, D-SACK and F-RTO each let the sender detect it and undo the cut.

for a principal

Judge the minimum-RTO trade for a fleet with mixed paths: faster loss recovery on short paths against spurious timeouts on paths prone to delay spikes.

## What the estimator believed A TCP sender sets its **retransmission timeout (RTO)** from recent round-trip time (RTT) samples, using RFC 6298: `RTO = SRTT + max(G, 4*RTTVAR)`, where `SRTT` is the smoothed RTT and `RTTVAR` the smoothed variation. Take a path that has been steady for a while: - samples cluster around 80 ms, so `SRTT` is about 80 ms; - they barely vary, so `RTTVAR` is a few milliseconds and `4*RTTVAR` is about 20 ms; - the formula gives about 100 ms, and RFC 6298 rule 2.4 says it SHOULD be rounded up to **1 second**. The estimator is right about the path as it was. It has no way to anticipate a **delay spike**: a burst of link-layer retransmissions on a radio link, a handover, a route change, or a deep queue building in front of the flow. If the RTT jumps to about 2 seconds for a few seconds, every segment in flight is late but none is lost. ## The timeline of a spurious timeout 1. **t = 0**: the timer was last restarted by an ACK; several segments are outstanding. 2. **t = 1 s**: no ACK has arrived, so the timer expires. Following RFC 6298 §5 the sender **retransmits the earliest unacknowledged segment**, **doubles the RTO** to 2 s, and restarts the timer. 3. Under RFC 5681 the timeout also sets the congestion window to **one full-sized segment** and sets the slow-start threshold to about half the data that was in flight (the details belong to congestion control). 4. **t = about 2 s**: the delayed ACKs for the **original** segments finally arrive. The receiver later gets the retransmitted copy as a **duplicate**. 5. Without further help, the sender treats these ACKs as the start of slow start after a loss. RFC 8985 §3.5 describes how a sender that marked all outstanding segments lost on the timeout goes on to retransmit the next ones as well, although they were never lost either. ## What it costs - **Wasted capacity**: duplicate segments cross the path and are thrown away by the receiver. - **A needless rate collapse**: the congestion window restarts from one segment and the threshold is halved, so throughput takes many round trips to recover. - **A distorted estimate**: without timestamps, **Karn's algorithm** forbids sampling the retransmitted segment, so the ACK that best describes the spike is discarded; the backed-off RTO of 2 s is what prevents a second spurious timeout if the delay persists. ## How a sender can tell afterwards | Method | Signal the sender sees | Specification | |---|---|---| | Timestamp echo | The first ACK after the timeout echoes the **original** segment's `TSval` in `TSecr`, older than the retransmission's | Timestamp option, RFC 7323; the Eifel detection algorithm, RFC 3522, cited there | | Duplicate SACK | The receiver reports, in a D-SACK block, that it received a range twice | RFC 2883 | | Forward RTO-Recovery | After the timeout, returning ACKs acknowledge segments that were **never** retransmitted | F-RTO, RFC 5682, cited by RFC 8985 | **RACK** (RFC 8985) is related but different: it does **not** detect spurious timeouts. What it does is limit the damage, because on an RTO it marks only the first outstanding segment lost automatically, and others only once enough time has passed since they were sent. ## Responding once it is detected RFC 2883 frames the response as an option rather than a rule. Having learned that a retransmission timeout was unnecessary, a sender can: - **undo the window reduction**, restoring the slow-start threshold to the old congestion window and slow-starting back up to it (congestion control owns the mechanics); - **adjust how it sets the RTO** so the same delay pattern does not trigger another spurious timeout; - with timestamps, **take a valid RTT sample** of the spike, since the echoed `TSval` identifies the original transmission. RFC 6298 notes this is what lets a sender converge on a correct estimate when the RTT exceeds 1 second. ## Why the conservative floor exists This scenario is exactly why RFC 6298 keeps a **1 second minimum RTO** as a SHOULD: the research it cites found that a large minimum is needed to avoid spurious retransmissions. Stacks that choose a lower floor, about 200 ms in Linux for example, recover faster from real losses on short paths but hit spurious timeouts on smaller spikes, which makes the detection mechanisms above more valuable.

  • Without timestamps, why can't the TCP sender learn about the 2-second delay from the ACK of the retransmitted segment?
    That ACK is ambiguous: it may answer the original or the retransmission, so Karn's algorithm requires discarding it as an RTT sample. The estimator learns about the slower path only from later segments acknowledged without retransmission. Meanwhile the doubled RTO of 2 s is what keeps a second spurious timeout from firing if the delay continues.
  • Would lowering the minimum RTO below 1 second make spurious timeouts from delay spikes more or less likely?
    More likely. A lower floor lets the RTO track a short, steady path closely, so a smaller spike outlasts it. The gain is faster recovery from real losses. RFC 6298 chose 1 second as the conservative trade; implementations with a lower floor lean harder on timestamp or D-SACK detection to undo the damage.

A shop reorders stock the moment a delivery runs later than usual. When the delayed van turns up right behind the replacement, the shop has paid twice and cut its orders for nothing. Checking the dispatch date on the delivery note shows the first van came after all.

saying these in an interview costs you the question

  • A retransmission timeout always means a segment was lost
  • Spurious timeouts are harmless because the receiver simply drops the duplicate copy
  • When the delay spike ends, TCP automatically restores the congestion window it cut
  • Without timestamps, the ACK for the resent segment gives a valid RTT sample of the spike
  • D-SACK stops the spurious retransmission from being sent in the first place