A TCP receiver's application stops reading its socket; what happens as the advertised window reaches zero, and why must the sender keep probing?
answer
- slow consumer, full buffer
- window updates are pure ACKs
- one octet past the window
- first probe after one RTO
- exponential backoff, no deadline
basics
~20 sWhen a TCP receiver's buffer fills, it advertises a zero window and the sender stops sending new data. The sender then sends periodic zero-window probes, because the ACK that reopens the window carries no data and is not retransmitted if lost.
solid answer
~40 sUnread bytes fill the receiver's buffer, each ACK advertises a smaller window, and at zero the sender must stop sending new data; its own send buffer then fills and the writing application blocks. When the reader resumes, the receiver sends a window update, but that is a pure ACK and RFC 1122 notes pure ACKs are not reliably delivered. If it is lost, both sides wait forever. So RFC 9293 requires **zero-window probing**: the sender transmits at least one octet (or a retransmission) past the closed window, first after the zero window has lasted one retransmission timeout, then at exponentially increasing intervals. The receiver must ACK each probe with its current window. While probes are acknowledged, the sender MUST keep the connection open, however long it takes.
go deeper
Recall that a receiver whose application stops reading eventually advertises a window of zero, and the sender must pause new data.
Explain why probing exists: window updates are pure ACKs that nobody retransmits, so the sender probes, first after one RTO, then with exponential backoff.
Diagnose a stalled stream as a slow consumer, not a dead peer: probes are acknowledged, so TCP keeps it open and the application must impose its own deadline.
Weigh the spec's choice to allow indefinite zero windows against resource exhaustion from many stalled peers, and decide where the timeout policy belongs.
## The scenario: a consumer that stops reading A server streams data to a client over TCP. The client application gets stuck (a blocked thread, a paused process, a disk it cannot write to) and stops reading its socket. Nothing is wrong with the network or with TCP. What follows is flow control doing its job, step by step: 1. Arriving bytes are acknowledged and held in the client's **receive buffer**, because the application is not reading them. 2. Each acknowledgment advertises a smaller window, since free buffer space is shrinking. The sender's right edge stops advancing. 3. The buffer fills and the client advertises a **window of zero**. The sender may not transmit new data. 4. On the server, unsent data piles up in its **send buffer**. When that fills too, the server application's writes block (or, on a non-blocking socket, report that they would block). 5. The connection is idle on the wire but healthy. It can stay that way for a long time. ## Why the sender must probe The receiver will eventually read and reopen its window by sending a **window update**: an acknowledgment that carries a larger `Window` value but no data. RFC 1122 §4.2.2.17 points out the trap: **ACK segments that contain no data are not reliably transmitted**. Nothing retransmits a lost pure ACK. If the window update is lost, the receiver believes it has reopened the window and waits for data, while the sender still believes the window is zero and waits for an update. Without another mechanism the connection would hang forever. That mechanism is **zero-window probing**. RFC 9293 §3.8.6.1 makes it mandatory: - probing of zero windows **MUST** be supported (MUST-36); - the sender regularly transmits **at least one octet of new data** if it has some, or a retransmission, even though the send window is zero; - the first probe **SHOULD** be sent once the zero window has existed for the **retransmission timeout** period (SHLD-29); - the interval between successive probes **SHOULD** increase **exponentially** (SHLD-30). Implementations commonly call the timer that drives this the **persist timer**; the name is an implementation convention, while the behaviour above is the RFC's. ## What the receiver does with a probe A probe is a real segment, so it elicits a real answer. RFC 9293 requires that when a receiver with a zero window gets a segment, it must still send an acknowledgment showing its **next expected sequence number** and its **current window**. Two outcomes: | Receiver state | What happens to the probe byte | What the ACK says | |---|---|---| | Window still zero | not accepted (outside the window) | same acknowledgment number, window 0 | | Window has reopened | accepted as ordinary data | acknowledgment advanced by one, new window | Either way the sender learns the true window. If the original update was lost, the next probe recovers it, bounded by the probe interval. ## How long a zero window may last RFC 9293 allows a receiver to keep its window closed **indefinitely** (MAY-8), and as long as it keeps acknowledging the probes, the sender **MUST** keep the connection open (MUST-37). The RFC's illustration is a printer daemon that stops reading because the printer ran out of paper: a legitimate pause, not a failure. The RFC notes this is subject to the implementation's resource-management concerns. For a running system that means: - A connection stuck at zero window is **alive**, not dead. The peer answers every probe; only its application is stalled. - TCP will not end it for you. A limit on how long a write may block belongs in the **application**, as a write deadline or an idle timeout, not in hopes that TCP gives up. - If probes stop being acknowledged, the situation changes: the peer is unreachable, and the implementation's retransmission limits eventually abort the connection. How many attempts it makes is an implementation choice. ## Keeping it apart from neighbouring mechanisms - **Not congestion.** A zero window says nothing about the network; the congestion window can be large while the receive window is zero. - **Not keepalive.** Keepalive probes an *idle* connection with no data waiting, to detect a dead peer; zero-window probing happens when data *is* waiting and the receiver has said it cannot take it. - **Not a reset condition.** A zero window by itself is no reason to reset the connection. RFC 9293 also requires a receiver to process RST and URG on incoming segments even while its window is zero (MUST-66).
- What does a TCP receiver do with the data byte carried in a zero-window probe?If its window is still zero, the byte lies outside the window and is not accepted, but RFC 9293 requires an acknowledgment showing the next expected sequence number and the current window of zero. If the window has reopened, the byte is accepted as normal data and the ACK advertises the new window. Either answer tells the sender the truth.
- Can a zero-window stall hold TCP connection resources indefinitely, and what should bound it?Yes. RFC 9293 lets a receiver keep its window closed indefinitely and requires the sender to keep the connection open while probes are acknowledged; it notes this is subject to the implementation's resource-management concerns. TCP will not end a healthy but stalled connection, so the bound belongs in the application: a write deadline or idle timeout that closes connections whose peer stops consuming.
saying these in an interview costs you the question
- A zero window means the network is congested and the sender should back off.
- A zero window means the connection is broken and should be reset.
- The receiver's window update is retransmitted if lost, so probing is unnecessary.
- The sender waits silently until the receiver announces that its window has reopened.
- Zero-window probes and keepalive probes are the same mechanism with the same timing.