skip to content

A TCP bulk transfer over a 1 Gbit/s path with a 100 ms round trip never exceeds about 5 Mbit/s, with no packet loss; how do the receive window and window scaling explain this?

level: seniorimportance: should knowfreq 30%

answer

  1. window divided by round-trip time
  2. bandwidth-delay product
  3. 16-bit field, 65,535 bytes
  4. shift count fixed at SYN
  5. buffer must match the BDP

basics

~20 s

A TCP connection moves at most one receive window per round trip. An unscaled 65,535-byte window over 100 ms caps it near 5.2 Mbit/s; filling this 12.5 MB bandwidth-delay product needs window scaling (shift 8 or more) and a matching buffer.

solid answer

~40 s

Throughput is bounded by `window / RTT`. The `Window` field is 16 bits, so without scaling the receiver can advertise at most 65,535 bytes: 65,535 bytes per 0.1 s is about 5.2 Mbit/s, exactly the ceiling seen. The path's **bandwidth-delay product** is 1 Gbit/s x 0.1 s = 12.5 MB, so the window must be roughly 190 times larger. RFC 7323's Window Scale option supplies a shift count (at most 14, for about 1 GiB), but only if **both** SYNs carry it; it is then fixed for the connection's lifetime. Scaling only lets the receiver express a big window: it must also have a receive buffer near the BDP and a reader that keeps up. No loss plus in-flight data pinned at the advertised window says the receiver, not the path, is the limit.

go deeper

for a junior

Recall that the TCP window field is 16 bits, so an unscaled window tops out at 65,535 bytes.

for a middle

Compute a bandwidth-delay product and the window / RTT ceiling, and explain that the window must cover the BDP to keep a path full.

for a senior

Diagnose a loss-free but slow long-haul transfer as window-limited: check both SYNs for Window Scale, then the advertised windows, then the receive buffer.

for a principal

Weigh window size against memory: every byte advertised is buffer promised per connection, so aim for the BDP of the paths you serve, not the maximum.

## Why throughput is capped by the window A TCP sender may have at most one window of unacknowledged data in flight, and it learns that data arrived only one **round-trip time** (RTT) later. So a single connection's throughput is bounded by: `throughput <= window / RTT` To keep a path full, the window must be at least the path's **bandwidth-delay product** (BDP): the number of bytes that fit "in the pipe" during one round trip. `BDP = bandwidth x RTT` RFC 7323 calls a path with a large BDP a **long, fat network** (LFN) and names the first problem basic TCP has on it: the header's `Window` field is **16 bits**, so the largest window it can express is 65,535 bytes, just under 64 KiB. ## Working the numbers For the scenario in the question, a 1 Gbit/s path with a 100 ms RTT: 1. **BDP** = 1,000,000,000 bit/s x 0.1 s = 100,000,000 bits = **12,500,000 bytes** (12.5 MB). 2. **Ceiling without scaling** = 65,535 bytes / 0.1 s = 655,350 bytes/s, about **5.2 Mbit/s**, roughly half a percent of the link. 3. **Shift needed**: 12,500,000 / 65,535 is about 191, so the scale factor must be at least 256, a **shift of 8** (65,535 x 2^8 = 16,776,960 bytes). A shift of 7 allows only 8,388,480 bytes, which is short. A loss-free path, steady acknowledgments and a sender whose in-flight data sits exactly at the advertised window are the signature of a **receive-window-limited** connection. ## What window scaling changes RFC 7323 extends the window to 30 bits without changing the header. The **Window Scale** option carries a shift count; each side keeps the true window as a 32-bit value and puts only its shifted form in the 16-bit field: - the receiver sends `SEG.WND = RCV.WND >> Rcv.Wind.Shift`; - the sender computes `SND.WND = SEG.WND << Snd.Wind.Shift`. The rules that matter for flow control: - **Offered only in SYN segments** and used only if **both** sides send it; otherwise both shifts are zero. The scale is therefore **fixed for each direction for the life of the connection**. How the option is laid out and negotiated is a header subject; its effect on `rwnd` is this one. - The window field of a **SYN or SYN-ACK is never scaled**, so the first window a peer sees is at most 65,535 bytes. - The shift is **limited to 14**, allowing windows up to 2^30 bytes, about **1 GiB**. A received value above 14 is logged and treated as 14. - The shift sets the **granularity**: with shift S the receiver can only advertise multiples of 2^S bytes. RFC 7323 §2.4 notes this rounding can force a receiver to retract its window slightly, which senders must tolerate. - Scaling applies only to the advertised window. The congestion window is a local variable and is not quantised by it. ## Scaling is permission, not capacity A scale factor only lets the receiver **express** a large window. The window it advertises is still its free buffer space. RFC 7323 notes that the maximum receive buffer determines the scale factor, and that buffer is typically set by default and can be overridden by the program before the connection opens. So filling a long fat pipe needs three things together: | Requirement | Why | Fails when | |---|---|---| | Scale negotiated | lets rwnd exceed 65,535 bytes | either SYN lacks the option | | Receive buffer >= BDP | the window advertises real free space | the buffer stays small | | Reader keeps up | freed space reopens the window | the application reads slowly | The sender's congestion window must also grow to the BDP, and loss on the path limits that separately; a clean path with a capped window points at flow control instead. ## Diagnosing and fixing the scenario Reason from the protocol: 1. Compare achieved throughput with `window / RTT`. About 5 Mbit/s at 100 ms is exactly 64 KiB per round trip: the window, not the link, is the ceiling. 2. Check the handshake: did **both** SYNs carry Window Scale? If either lacks it, the connection is stuck at 16 bits for its lifetime and no later change helps. 3. If scaling is on, check the advertised windows during the transfer. Windows that stay small point at a small receive buffer or a slow reader. 4. Fix by allowing a receive buffer of at least the BDP before the connection opens, which lets the stack pick a large enough shift. Many stacks grow receive buffers automatically, an implementation feature rather than a protocol rule. Oversizing is not free: every byte of advertised window is buffer the receiver has promised, so a window far above the BDP costs memory per connection without adding throughput.

  • Why does RFC 7323 cap the TCP window scale shift at 14?
    TCP judges whether data is old or new by whether its sequence number lies within 2^31 bytes of the window's left edge. Sender and receiver windows can be out of phase by up to a window, so twice the maximum window must stay below 2^31, meaning the window must be below 2^30. A 16-bit field shifted by 14 stays under that bound; a larger received value is treated as 14.
  • If window scaling is negotiated with a large shift, is a TCP connection guaranteed to fill a long fat pipe?
    No. The shift only lets the receiver advertise a large window; the advertised value is still its free buffer space, so a small receive buffer or a slow reader keeps it small. The sender's congestion window must also grow to the bandwidth-delay product, and loss on the path limits that separately.
  • Why can a TCP endpoint not turn on window scaling after a slow connection is already established?
    RFC 7323 allows the Window Scale option only in SYN segments and ignores it anywhere else, so the shift for each direction is fixed when the connection opens. A connection that started without scaling stays limited to 65,535 bytes; only a new connection, set up with a large enough receive buffer, can use a bigger window.

saying these in an interview costs you the question

  • Window scaling widens the header's window field beyond 16 bits.
  • The window scale can be renegotiated mid-connection when throughput drops.
  • The window advertised in a SYN segment is already scaled.
  • A faster link would fix it, because throughput is limited by bandwidth alone.
  • Negotiating a large shift fills the pipe even if the receive buffer stays small.