skip to content

When an IPsec VPN client falls back to carrying IKE and ESP over TCP on a UDP-blocking network, what does the tunnel lose, and why does RFC 9329 make it a last resort?

level: seniorimportance: should knowfreq 20%

answer

  1. configured, not negotiated
  2. one connection carries every SA
  3. loss becomes delay
  4. voice retransmitted, not dropped
  5. retransmit over UDP before switching

basics

~20 s

Over TCP, every Child SA shares one reliable connection: loss turns into delay for all flows, real-time traffic is retransmitted, and per-SA QoS and DF copying are lost. RFC 9329 therefore says to prefer direct or UDP-encapsulated ESP.

solid answer

~50 s

**RFC 9329** (which obsoletes RFC 8229) frames IKE and ESP messages on one TCP connection per IKE SA - a stream prefix, then a 16-bit Length before each message. It must work on TCP 4500; port 443 and an optional TLS layer are configured extras for networks that pass only web traffic. Support is configured on both peers, not negotiated. The costs come from TCP's reliability: a lost outer segment stalls every inner flow, inner TCP sees delay and may retransmit spuriously, and voice is retransmitted late instead of dropped. All Child SAs share one connection, so they get one DSCP marking, and the DF bit cannot be copied outward. Hence the RFC's rules: prefer direct or UDP-encapsulated ESP, try UDP first and retransmit at least once before falling back, and move back to UDP when MOBIKE changes networks.

go deeper

for a junior

Know that IPsec can ride TCP when UDP is blocked, and that this is a fallback with a performance price, not the default.

for a middle

Describe RFC 9329's framing, its mandatory TCP 4500 port, the configured rather than negotiated support, and why loss turns into delay.

for a senior

Name the concrete losses - shared head-of-line blocking, late voice, one QoS class, DF handling - and the fallback rules: UDP first, one retransmission, back to UDP after a move.

for a principal

Decide when the reach is worth the cost, how quickly clients fall back, and how you would measure how many sessions are stuck on TCP.

## What TCP encapsulation is When a network drops UDP, ESP-in-UDP on port 4500 cannot help. **RFC 9329**, *TCP Encapsulation of IKE and IPsec Packets* (Standards Track, obsoleting RFC 8229), carries both IKE and ESP inside a TCP connection instead: - The **TCP Originator** opens the connection and sends a stream prefix, the ASCII bytes `IKETCP`, once at the start. - Every IKE message or ESP packet is then preceded by a **16-bit Length** field that marks message boundaries in the stream. AH is not supported. - An originator should open **one TCP connection per IKE SA**, and every Child SA's ESP traffic shares it. - Every implementation **must support TCP port 4500**. Other ports - 443 is the obvious choice on restrictive networks - are optional configured alternatives. - **Appendix A** adds optional TLS around the stream, so the tunnel can pass middleboxes that expect TLS and can use HTTP `CONNECT` through a web proxy. Because IKE already secures everything, both sides *should* allow a NULL cipher for that TLS layer. - Support is **configured on both peers, not negotiated** in IKE, precisely because the networks that need it may block the UDP that a negotiation would ride on. ## What the tunnel loses ESP is designed to ride an unreliable network; TCP is not one. RFC 9329's own performance section lists the costs: | Property | ESP direct or in UDP | ESP over TCP | |---|---|---| | Outer loss | Seen by inner flows as loss | Hidden and turned into delay | | Head-of-line blocking | None between packets | One lost segment stalls every Child SA | | Real-time traffic | Late packets dropped | Retransmitted and delivered late | | Per-SA QoS | Each Child SA can carry its own DSCP | One connection, effectively one class | | DF bit | Can be copied from inner to outer header | Controlled by the TCP stack | | Framing overhead per packet | IPv4 20 + UDP 8 = 28 bytes | IPv4 20 + TCP 20 + Length 2 = 42 bytes | Some of these need unpacking: - **TCP-in-TCP.** Inner TCP connections see outer loss only as delay. They may fire their own retransmission timeouts, retransmit spuriously, and settle at a lower sending rate than the path allows. Stacked timers can also feed back into more loss - the effect usually called TCP meltdown - which RFC 9329 says can be limited by carefully managing the two layers' timeouts. - **Real-time traffic.** Voice prefers a lost packet to a late one; over TCP it gets the late one. - **Shared fate.** Every flow through the gateway shares the stall, so RFC 9329 recommends limiting how many inner connections share a TCP-encapsulated tunnel when latency-sensitive traffic is present. - **Overhead.** The 42-byte figure assumes a TCP header without options, against 28 bytes for UDP encapsulation - 14 more bytes per packet before any TLS record overhead. - **Visibility.** The `IKETCP` prefix is easy to recognise, and RFC 9329 notes that operators who block IPsec on purpose can filter on it or on TCP 4500. ## Why it is a last resort The RFC turns these costs into rules about *when* to use it: 1. Implementations **should prefer** direct ESP or UDP encapsulation, and **must favour** them whenever possible. 2. An initiator with no prior knowledge of the network **should try IKE over UDP first**, and the initial message should be retransmitted at least once before falling back, unless the network is known to block UDP. 3. Switching from UDP to TCP starts a **new `IKE_SA_INIT`** with a new SPI and recalculated NAT detection payloads. 4. With **MOBIKE (RFC 4555)**, an initiator that changes networks sends `UPDATE_SA_ADDRESSES` over UDP first and returns to ESP over IP or UDP 4500 if that works, even if the previous network forced TCP. ## Telling the two failures apart TCP encapsulation is the answer to *UDP filtering*, not to *mapping expiry*. A tunnel that dies after idling on a network that passes UDP needs inside-originated keepalives, not TCP. A tunnel that never completes `IKE_SA_INIT` over UDP while web traffic works is the case TCP encapsulation exists for - and it should stay on TCP only as long as the network forces it.

  • Why must a client that moved off a UDP-blocking network try UDP again instead of keeping its TCP connection?
    Because every cost of TCP carriage stays with the session as long as it rides TCP. RFC 9329's MOBIKE rules send UPDATE_SA_ADDRESSES over UDP first after a network change; if UDP answers, ESP goes back to IP or UDP 4500 even though the previous network forced TCP.
  • Why does RFC 9329 let the TLS layer around TCP encapsulation use a NULL cipher?
    The tunnel's security comes from IKE and ESP; TLS is there only to get past middleboxes. Encrypting twice costs CPU and bytes for no gain, so both sides should allow a NULL cipher, though the RFC notes TLS 1.3 offers only AEAD suites and had no recommended NULL suite.

saying these in an interview costs you the question

  • RFC 9329 requires IKE and ESP over TCP to use port 443.
  • TCP encapsulation is negotiated inside IKE, so any client can try it.
  • Voice calls improve over TCP because lost packets get retransmitted.
  • Each Child SA gets its own TCP connection, so one loss stalls one flow.
  • Wrapping the stream in TLS adds to the tunnel's security.