Why can't TCP use a fixed retransmission timeout, and how does RFC 6298 compute the RTO from measured round-trip times?
answer
- paths differ by orders of magnitude
- an average plus a spread
- gains of one-eighth and one-quarter
- the variance term's multiplier
- a conservative floor
basics
~20 sRound-trip times differ hugely between paths and drift over time, so a fixed timeout is too slow or spuriously early. RFC 6298 sets RTO = SRTT + max(G, 4 x RTTVAR), with gains 1/8 and 1/4, rounded up to 1 s.
solid answer
~50 sA timeout that suits a 1 ms LAN would fire constantly on a 600 ms satellite path, and one sized for the satellite would idle a LAN for most of a second after every loss; even one path's RTT drifts as queues grow. So TCP measures. RFC 6298 keeps `SRTT`, a smoothed average of RTT samples (gain `alpha = 1/8`), and `RTTVAR`, a smoothed mean deviation (gain `beta = 1/4`). The first sample R sets `SRTT = R` and `RTTVAR = R/2`; later samples update `RTTVAR` first, then `SRTT`. The timeout is `RTO = SRTT + max(G, 4*RTTVAR)`, G being the clock granularity, and it SHOULD be rounded up to 1 second. Before any sample, the RTO SHOULD start at 1 s (RFC 2988 used 3 s). The variance term is the point: a jittery path gets a wide safety margin, a steady one a tight margin.
code
pseudocode · 12 lineson_rtt_sample(R):
if no_sample_yet:
SRTT = R
RTTVAR = R / 2
else:
RTTVAR = (1 - 1/4) * RTTVAR + (1/4) * abs(SRTT - R) # uses old SRTT
SRTT = (1 - 1/8) * SRTT + (1/8) * R
RTO = SRTT + max(G, 4 * RTTVAR)
if RTO < 1 second:
RTO = 1 second # SHOULD, RFC 6298 (2.4)
if MAX_RTO is set:
RTO = min(RTO, MAX_RTO) # MAX_RTO >= 60 s (2.5)go deeper
Recall that TCP measures round-trip times and sets its retransmission timeout from them, because a fixed timeout cannot fit both fast and slow paths.
Explain SRTT and RTTVAR, the RFC 6298 formula with K = 4, the 1 s initial value and floor, and walk through one update by hand.
Show how variance widens the timeout on jittery paths, why the floor exists, and how implementation floors below 1 s trade spurious timeouts for faster recovery.
Discuss where a fixed-gain estimator breaks down, such as per-ACK sampling with timestamps or datacentre RTTs far below the floor, and what a fleet-wide floor change would risk.
## Why a fixed timeout cannot work A TCP sender that has sent data and heard nothing back must eventually decide the data was lost and send it again. The waiting period is the **retransmission timeout (RTO)**. Choosing it is a balance: - **Too short** and the timer fires while the data or its acknowledgment is still in transit. The sender retransmits needlessly (a **spurious retransmission**), wastes capacity and, because TCP treats a timeout as a congestion signal, slashes its own sending rate for nothing. - **Too long** and a real loss leaves the connection idle far longer than necessary. The right value depends on the **round-trip time (RTT)** of the path, and RTTs span orders of magnitude: well under a millisecond inside a data centre, tens of milliseconds across a continent, hundreds of milliseconds over a geostationary satellite. Even on one path the RTT moves as queues fill and drain. No constant can serve all of these, so RFC 9293 §3.8.1 requires the RTO to be computed dynamically with the algorithm of **RFC 6298** (which obsoletes RFC 2988). ## The two state variables RFC 6298 keeps two numbers per connection and a clock granularity `G`: - **`SRTT`**, the smoothed round-trip time: an exponentially weighted moving average of RTT samples. - **`RTTVAR`**, the round-trip time variation: a smoothed average of how far each sample lands from `SRTT`. - **`G`**, the granularity of the clock used for timing; it guarantees the variance term is never zero. Averaging alone is not enough. The original RFC 793 procedure set the timeout to a fixed multiple (1.3 to 2.0) of a smoothed RTT, and RFC 1122 records that it failed because it assumed RTT variation would be small and constant. Jacobson's algorithm, which RFC 6298 standardises, adds the variation explicitly. ## The algorithm step by step 1. **Before any sample**, the RTO SHOULD be **1 second** (rule 2.1). RFC 2988 and RFC 1122 used 3 seconds; RFC 6298 still allows 3 s or any value above 1 s. 2. **First sample R**: `SRTT = R`, `RTTVAR = R/2`, `RTO = SRTT + max(G, 4*RTTVAR)`. 3. **Each later sample R'**: first `RTTVAR = (3/4)*RTTVAR + (1/4)*|SRTT - R'|`, then `SRTT = (7/8)*SRTT + (1/8)*R'`, then recompute the RTO. The order is a MUST: the deviation is measured against the old average. 4. **Floor**: if the RTO is below 1 second it SHOULD be rounded up to 1 second (rule 2.4). 5. **Ceiling**: an upper bound MAY be applied, provided it is at least 60 seconds (rule 2.5). The constants `alpha = 1/8` and `beta = 1/4` are SHOULDs; `K = 4` is fixed in the formula. ## A worked example Assume a clock fine enough that `G` is negligible, on a long-delay path: | Sample | RTTVAR | SRTT | RTO = SRTT + 4*RTTVAR | |---|---|---|---| | First, 600 ms | 300 ms | 600 ms | 1800 ms | | Second, 680 ms | 0.75*300 + 0.25*80 = 245 ms | 0.875*600 + 0.125*680 = 610 ms | 1590 ms | | Third, 610 ms | 0.75*245 + 0 = 183.75 ms | 610 ms | 1345 ms | The RTO starts wide because one sample says little about variation, then tightens as consistent samples arrive. On a short path the floor dominates: a first sample of 2 ms gives `2 + 4*1 = 6 ms`, which is rounded up to 1 second. ## Why the floor is so large RFC 6298 justifies the 1 second minimum as conservative: research it cites found a large minimum RTO is needed to avoid spurious retransmissions, and the RFC says future work might show a smaller minimum is acceptable. Implementations differ here: some widely deployed stacks use a floor well below 1 second (Linux, for example, uses about 200 ms). That is an **implementation choice**, not the RFC's rule. ## Where the samples come from The estimator is only as good as its inputs. RFC 6298 §3 requires at least one RTT measurement per round trip and forbids samples from retransmitted segments (**Karn's algorithm**), because their ACKs are ambiguous. With the TCP **timestamp option** (RFC 7323), nearly every ACK can yield a sample; RFC 7323 warns that feeding many samples per round trip into the same `alpha` and `beta` shortens the history the estimator remembers. ## Misconceptions - The RTO is not "twice the average RTT"; that was the flavour of the RFC 793 procedure that RFC 6298's predecessors replaced. - The 1 second floor is a SHOULD in RFC 6298, not a MUST, and 200 ms is an implementation value, not an RFC value. - A single slow sample moves `SRTT` by only one-eighth of the difference, though it moves `RTTVAR` faster.
- Why does RFC 6298 update RTTVAR before SRTT when a new RTT sample arrives?`RTTVAR` measures how far the new sample lands from the estimate that existed before it, `|SRTT - R'|`. Updating `SRTT` first would pull the average toward the sample and understate the deviation, making the timeout too tight exactly when the path became less predictable. RFC 6298 makes the order a MUST.
- Is the 1-second minimum RTO a hard requirement of the TCP specifications?No. RFC 6298 says the RTO SHOULD be rounded up to 1 second, as a conservative guard against spurious retransmissions, and acknowledges research may later justify a smaller floor. Some widely deployed stacks use a lower floor, about 200 ms in Linux for example; that is an implementation choice rather than the RFC's rule.
saying these in an interview costs you the question
- The RTO is simply twice the average round-trip time
- RFC 6298 requires a minimum RTO of 200 milliseconds
- Current TCP specifications require an initial RTO of 3 seconds
- SRTT just tracks the latest sample, so one slow packet resets it
- RTTVAR counts how many segments were recently retransmitted