How does a TCP sender detect that a segment was lost, and what are the two ways it decides to retransmit it?
answer
- silence versus repeated feedback
- a clock that adapts to the path
- the same cumulative ACK, again
- what if nothing follows the loss
basics
~20 sA TCP sender infers loss two ways: its retransmission timer (RTO) expires with data still unacknowledged, or a third duplicate ACK arrives and triggers fast retransmit. Fast retransmit usually fires first; the timer is the slower backstop.
solid answer
~50 sTCP never observes a loss directly; it infers one from acknowledgments. The first signal is the **retransmission timeout**: while data is outstanding a timer runs, and if it expires (after an RTO computed from measured round-trip times per RFC 6298, normally no less than 1 s) the sender retransmits the earliest unacknowledged segment and doubles the RTO. The second is **fast retransmit** (RFC 5681): segments arriving past a gap make the receiver repeat the same cumulative ACK, and on the **third duplicate ACK** the sender resends the missing segment without waiting for the timer. Fast retransmit typically repairs a loss in about one round trip; a timeout costs at least the RTO and, under RFC 5681, cuts `cwnd` to one segment. The timer stays as the backstop for losses followed by too few segments to produce three duplicates.
go deeper
Recall the two triggers: the retransmission timer expiring, and three duplicate ACKs causing fast retransmit. Be clear that TCP learns about loss only from acknowledgments.
Explain why the receiver produces duplicate ACKs after a gap, why the threshold is three, and what the timer resends and how its value doubles on each expiry.
Show when fast retransmit cannot fire, such as tail loss or a tiny window, and why a timeout is so much costlier: the wait plus the collapse of the congestion window.
Frame loss detection as a latency budget: timer-driven recovery dominates tail latency for short request-response traffic, which is what motivates SACK and RACK-TLP deployment.
## What "loss" means to a TCP sender TCP runs over IP, which may drop, delay, duplicate or reorder packets without telling anyone. A TCP **sender** therefore never learns directly that a segment was lost. All it sees is the stream of **acknowledgments** (ACKs) coming back from the receiver. An ACK in TCP is **cumulative**: its acknowledgment number says "I have every byte before this one". From that stream the sender has to infer loss, and RFC 9293 and its companion documents give it exactly two ways to do so. ## Signal 1: the retransmission timer The **retransmission timeout (RTO)** is TCP's safety net. RFC 9293 §3.8.1 requires the RTO to be computed with the algorithm of **RFC 6298**, which derives it from measured round-trip times (smoothed average plus four times the smoothed variation) and says it SHOULD be rounded up to at least **1 second**. RFC 6298 §5 recommends managing the timer like this: 1. When a segment carrying data is sent and the timer is not running, start it for RTO seconds. 2. When an ACK acknowledges new data, restart it. 3. When all outstanding data is acknowledged, stop it. 4. If it expires, **retransmit the earliest unacknowledged segment**. 5. **Double the RTO** ("back off the timer") and start the timer again. The timer works whatever the traffic pattern: it needs no further segments and no cooperation from the receiver beyond an eventual ACK. Its weakness is speed. The RTO is deliberately conservative, usually several round trips long, and RFC 5681 treats a timeout as a severe congestion signal: the congestion window (`cwnd`) drops to one full-sized segment and the sender slow-starts again (the congestion-control rules themselves belong to that topic). ## Signal 2: fast retransmit on duplicate ACKs When one segment is lost but later segments arrive, the receiver cannot advance its cumulative ACK past the hole. Each later arrival therefore produces an ACK carrying the **same** acknowledgment number: a **duplicate ACK**. RFC 5681 says the sender SHOULD use **fast retransmit**: after **three duplicate ACKs** (with no ACK in between that moves the left edge of the window), it retransmits what appears to be the missing segment **without waiting for the timer**. Why three? A single duplicate ACK only says a segment arrived out of order, and networks do reorder. Waiting for three makes loss the more likely explanation, at the price of a little delay. ## The two signals side by side | | Retransmission timeout | Fast retransmit | |---|---|---| | Trigger | RTO expires with data outstanding | Third duplicate ACK | | Specification | RFC 6298 (required by RFC 9293) | RFC 5681 (SHOULD) | | Typical delay to repair | At least the RTO, normally >= 1 s | About one round trip | | Needs later segments in flight | No | Yes, at least three | | What is resent | Earliest unacknowledged segment | The segment that appears missing | | Congestion response | `cwnd` to 1 segment, slow start | Milder: fast recovery | ## Why the timer is still needed Fast retransmit depends on **later segments** reaching the receiver. It cannot fire when: - the lost segment is among the last three sent, so fewer than three segments follow it; - the window is so small that fewer than three segments are in flight at all; - the retransmitted segment is itself lost; - every segment of a window is lost, or all the returning ACKs are. In each case the timer is the only thing that will ever restart the flow. That is why every TCP must implement it, while fast retransmit is a performance optimisation layered on top. Extensions such as **SACK** (selective acknowledgments, RFC 2018) and **RACK-TLP** (RFC 8985) make loss detection sharper and shrink the cases that fall through to the timer, but they do not remove it. ## Misconceptions to avoid - **"TCP receivers send a negative acknowledgment for a missing segment."** TCP has no NAK; duplicate ACKs and SACK blocks are positive statements from which the sender infers the gap. - **"The timeout is a fixed constant."** It is measured per connection and changes as round-trip samples arrive. - **"Fast retransmit replaces the timer."** It only works when enough segments follow the loss; the timer remains the backstop. - **"A timeout and a fast retransmit are equally costly."** A timeout costs a long wait and a collapse of the sending rate.
- Why does a TCP sender wait for three duplicate ACKs instead of retransmitting on the first one?A duplicate ACK only says a segment arrived past a gap. IP networks reorder packets, so one or two duplicates are often a late segment rather than a lost one. RFC 5681 uses three duplicates as the point where loss becomes the likelier explanation; retransmitting on every duplicate would resend data that was merely late and trigger needless congestion responses.
- When the retransmission timer expires, does the TCP sender resend all outstanding data at once?RFC 6298 rule 5.4 says to retransmit the earliest unacknowledged segment, then double the RTO and restart the timer. The congestion window is now one segment, so the sender cannot blast the rest; it proceeds as ACKs return. How eagerly the remaining segments are treated as lost is left to the loss-recovery algorithm in use.
saying these in an interview costs you the question
- TCP receivers send a NAK naming the segment that went missing
- Fast retransmit fires as soon as the first duplicate ACK arrives
- The retransmission timeout is a fixed constant, the same on every path
- A lost segment is only ever resent after its timer expires
- A timeout and a fast retransmit cost the sender about the same