Over a site-to-site tunnel, TCP handshakes and small requests succeed but large downloads stall; why does clamping the MSS in passing SYNs fix it?
answer
- endpoints only know their own link
- encapsulation shrinks the usable MTU
- only full-sized segments disappear
- missing ICMP feedback
- rewrite both SYN and SYN-ACK
basics
~20 sEndpoints derive MSS from their own 1500-byte links, but tunnel encapsulation shrinks the path MTU, so full-sized segments are dropped and, with ICMP feedback filtered, never shrink. Rewriting passing SYNs' MSS to fit the tunnel stops oversized segments being built.
solid answer
~50 sEach endpoint derives its `MSS` from its own interface, 1460 on a 1500-byte IPv4 link, and cannot see that a tunnel in the middle adds encapsulation headers. The handshake, ACKs and short requests are small, so they pass; the first full-sized data segments are too big for the tunnel. Stacks doing Path MTU Discovery send with Don't Fragment set, so the tunnel router drops those packets and relies on an ICMP message to make the sender shrink; if a firewall filters that message, the sender keeps retransmitting the same size and the transfer stalls. MSS clamping makes the tunnel router lower the MSS option in every SYN and SYN-ACK it forwards to at most the tunnel MTU minus 40 bytes for IPv4, so both senders build segments that fit from the start. It is a middlebox practice, not part of RFC 9293, and it only covers TCP flows whose SYNs cross that router.
go deeper
Recall that MSS is the TCP payload limit each side announces in its SYN, and that tunnels add headers that make the usable packet size smaller.
Explain why the handshake succeeds while bulk data fails, and compute a clamp value from a tunnel MTU for IPv4 and for IPv6.
Diagnose a path MTU black hole from its symptom pattern, clamp both directions correctly, and state what clamping cannot cover, such as UDP, QUIC and asymmetric routing.
Weigh clamping as standard edge hygiene against fixing ICMP handling across the estate, and decide where each belongs in a network design.
## Why the symptom has this shape Take a site-to-site tunnel whose usable MTU is 1400 bytes, joining two sites whose hosts sit on 1500-byte Ethernet links. The failure unfolds in a fixed order: 1. The client's `SYN` announces `MSS` 1460, derived from its own 1500-byte interface. The server's `SYN-ACK` does the same. 2. The handshake segments are tens of bytes long, so they cross the tunnel easily. 3. The client's request is a few hundred bytes, so it also fits. 4. The server answers with full-sized segments: 1460 bytes of data plus 40 bytes of headers is a 1500-byte IP packet, 100 bytes too big for the tunnel. 5. A sender doing Path MTU Discovery sets **Don't Fragment** on IPv4 (and IPv6 routers never fragment at all), so the tunnel router drops the packet and should return an ICMP error telling the sender to use a smaller size. 6. If a firewall filters that ICMP, the sender never learns. It retransmits the same 1500-byte packet, which is dropped again, and the download stalls while small traffic keeps working. This pattern, handshake fine and small exchanges fine but bulk transfer dead, is the signature of a **path MTU black hole**. ## What each endpoint can know RFC 9293 defines the MSS a host advertises as "the effective MTU minus the fixed IP and TCP headers" of its own interface. Neither endpoint is directly connected to the tunnel, so neither can include its overhead. The MSS option was designed to describe the receiver, not the path; discovering the path is Path MTU Discovery's job, and that is exactly the part that failed. ## How MSS clamping works The router at the tunnel edge sees every `SYN` and `SYN-ACK` for flows that cross it, and it rewrites them in flight: - It reads the `MSS` option (kind 2) in each segment with `SYN` set. - If the value exceeds the limit, it **lowers** it to the limit. It never raises a smaller value. - Because RFC 9293 includes all options in the TCP checksum, it must **recompute the checksum** after the edit. - It applies this in **both directions**, because each announced MSS limits only what its announcer receives. The endpoints then build segments that fit from the start, and Path MTU Discovery never has to fire for those flows. ## Choosing the clamp value The limit is the smallest MTU the router knows on the path, minus the **fixed** headers: | Tunnel MTU | Inner IP version | Fixed headers | Clamp to | |---|---|---|---| | 1400 | IPv4 | 20 + 20 | 1360 | | 1400 | IPv6 | 40 + 20 | 1340 | Do not subtract TCP options such as timestamps. RFC 6691 and RFC 9293 both say the advertised MSS ignores options and the **sender** trims each segment's data by the option bytes it carries. A clamp slightly too low only costs a little efficiency; a clamp too high leaves the black hole in place. A PPPoE access link is another classic home of the same problem, because its encapsulation also shrinks the MTU below Ethernet's 1500. ## Limits of the technique | Limitation | Why | |---|---| | TCP only | UDP and QUIC carry no MSS option to rewrite | | Needs to see the SYNs | With asymmetric routing, a SYN that bypasses the router is not clamped | | Fixed at setup | A path change later in the connection is not reflected | | Hides the real fault | Filtered ICMP still breaks other protocols and other paths | ## Where it sits among the fixes - **Restore the feedback.** Let the ICMP messages Path MTU Discovery depends on through; how that message and the discovery process work is a separate topic. - **Probe instead of relying on ICMP.** RFC 9293 strongly recommends Packetization Layer Path MTU Discovery, which searches for a working size without ICMP. - **Clamp at the narrow link.** Cheap, immediate and local, which is why it is widely applied on tunnels, but it remains a workaround for TCP.
- Why must MSS clamping rewrite both the SYN and the SYN-ACK rather than just the client's SYN?Each MSS option limits only what its announcer receives. Clamping the client's SYN shrinks segments the server sends to the client, but the server's SYN-ACK still announces 1460, so the client keeps sending full-sized segments toward the server. An upload in that direction would still stall. Clamping both leaves every sender building segments that fit the tunnel.
- Why is MSS clamping usually called a workaround rather than the fix for a path MTU black hole?The root fault is that the sender never receives the ICMP feedback Path MTU Discovery relies on. Clamping repairs only TCP flows whose SYNs cross the clamping router, and only for the MTU it knows about. UDP and QUIC traffic, asymmetric paths and later path changes are still exposed, so the durable fixes are restoring ICMP delivery or using packetization-layer probing.
- Should the clamp value leave room for the TCP timestamps option?No. RFC 9293 and RFC 6691 define the advertised MSS as the MTU minus only the fixed IP and TCP headers. A sender that adds 12 bytes of timestamps must itself shrink each segment's data by those 12 bytes, so an MTU of 1400 is correctly clamped to 1360 for IPv4.
saying these in an interview costs you the question
- MSS clamping lowers the interface MTU on the endpoints
- Clamping only the client's SYN fixes both directions
- MSS clamping also protects UDP and QUIC flows through the tunnel
- The clamp value should be the tunnel MTU itself
- The clamp must also subtract the timestamps option bytes