skip to content

An IPsec ESP tunnel between two gateways on 1,500-byte links receives full-size 1,500-byte IPv4 packets with DF set; what happens, and what inner MTU fits?

level: seniorimportance: must knowfreq 35%

answer

  1. the tunnel makes every packet bigger
  2. DF set or clear decides
  3. who sends the ICMP, to whom
  4. fragment before or after encrypting

basics

~20 s

With AES-GCM the IPsec tunnel-mode packet would be 1,556 bytes, too big, so the gateway SHOULD drop it and send an ICMP PMTU message to the source; the largest inner packet that fits is 1,446 bytes.

solid answer

~50 s

With AES-GCM tunnel mode the overhead is 56 bytes for this size, so a 1,500-byte inner packet becomes 1,556 bytes. RFC 4301 §8.2.1 says what to do when a packet will not fit the SA's path MTU: an IPv4 packet with `DF` set SHOULD be discarded and an ICMP PMTU message sent back; with `DF` clear the gateway SHOULD fragment, before or after encryption, and send no ICMP; an IPv6 packet is discarded with a Packet Too Big. The largest inner packet that fits 1,500 is **1,446** bytes (1,438 with AES-CBC and a 16-byte HMAC). If that ICMP is filtered, hosts never learn it and full-size packets vanish while small ones pass. The operator's fix is to give the tunnel an MTU of the computed inner size so hosts learn it; MSS clamping is the general TCP workaround.

go deeper

for a junior

Remember that a tunnel makes every packet bigger, so a full-size packet no longer fits a link of the same MTU.

for a middle

Compute the inner MTU for a named transform and explain what DF set, DF clear and IPv6 each make the gateway do.

for a senior

Diagnose the small-works, large-stalls pattern on a tunnel, say who receives each ICMP message, and pick between pre- and post-encryption fragmentation.

for a principal

Set estate-wide tunnel MTU policy from the transforms you allow, weighing ICMP filtering, reassembly load on gateways and per-transform recomputation.

## Why the tunnel creates the problem IPsec tunnel mode puts a whole new IP header plus the ESP fields around every packet. RFC 4301 §8 says it plainly: applying AH or ESP increases the packet's size and may push it past the path MTU of the SA. With AES-GCM (8-byte IV, 16-byte ICV), a 1,500-byte IPv4 packet becomes: - 20 (outer IPv4) + 8 (ESP header) + 8 (IV) + 1,500 + 2 (trailer) + 2 (padding to a 4-byte boundary) + 16 (ICV) = **1,556 bytes**. That is 56 bytes too many for a 1,500-byte link, so the gateway has to decide what to do with every full-size packet. ## What RFC 4301 tells the gateway to do The security association database keeps a **PMTU** value per SA, which the gateway compares against the size each packet will have after encapsulation. When a packet would exceed it, RFC 4301 §8.2.1 sets three cases: | Original packet | Gateway SHOULD | ICMP to the source | |---|---|---| | IPv4, `DF` set | discard it | send an ICMP PMTU message (Destination Unreachable, fragmentation needed and DF set) | | IPv4, `DF` clear | fragment it, before or after encryption per its configuration, and forward | do not send one | | IPv6 | discard it | send a Packet Too Big | For the scenario in the question, that is the first row: the packet is dropped, and the source is told the next-hop MTU. The MTU the source needs is the path MTU minus the gateway's own overhead; RFC 4301 describes that accounting when it updates an SA's PMTU, "taking into account the size of the AH or ESP header... any crypto synchronization data, and the overhead imposed by an additional IP header, in the case of a tunnel mode SA". ## Computing the inner MTU For AES-GCM in tunnel mode on a 1,500-byte path, the fixed part is 20 + 8 + 8 + 16 = 52 bytes, leaving 1,448 bytes for the inner packet, the padding and the 2 trailer bytes. 1,448 is a multiple of 4, so a 1,446-byte inner packet needs no padding at all and fills the packet exactly: **inner MTU 1,446**. A TCP connection over IPv4 without TCP options then carries at most 1,446 - 40 = 1,406 bytes of data per segment. Other transforms give other answers, which is why the number must be computed rather than remembered: - AES-CBC with `AUTH_HMAC_SHA2_256_128`: fixed part 20 + 8 + 16 + 16 = 60, leaving 1,440, a multiple of 16 → **1,438**. - An IPv6 outer header (40 bytes) with AES-GCM → **1,426**. ## Fragmenting: before or after encryption When `DF` is clear, the gateway has two choices, and RFC 4301 leaves the choice to its configuration: 1. **After encryption.** The 1,556-byte ESP packet is split into outer IP fragments. RFC 4303 §3.4.1 requires the receiving gateway to reassemble them **before** ESP processing and to discard anything offered to ESP that still looks like a fragment. That reassembly costs the far gateway memory and CPU, and RFC 4303 §3.3.4 notes that accepting fragments for reassembly creates denial-of-service exposure. 2. **Before encryption.** The gateway fragments the inner packet first, then encrypts each fragment as its own ESP packet. The far gateway decrypts whole ESP packets and forwards inner fragments, which the destination host reassembles. Each fragment pays the full ESP overhead, and the SA must be one that may carry fragments. Only tunnel mode can do the second: RFC 4301 lets transport mode protect whole datagrams only, because a single header cannot hold fragmentation state for both the plaintext and the ciphertext packet. ## The outer DF bit and the path between gateways RFC 4301 §8.1 requires that the outer header's `DF` bit be configurable per SA as **set, clear or copy from inner header** (when both headers are IPv4). If routers between the gateways meet an outer packet too big for them, their ICMP goes to the **encapsulating gateway** — the outer source — not to the host. The gateway may act on that unauthenticated message, map it to the SA, lower the SA's PMTU, and tell the host on the next oversized packet. RFC 4301 §8.2.2 also requires the SA's PMTU to be **aged**, so a value learned once is periodically retested. ## How it looks in production - Pings and small requests work; large transfers stall. Full-size packets with `DF` set are dropped, and if the ICMP is filtered anywhere, nothing tells the sender to shrink. - The robust fix is to compute the inner MTU for the transform actually negotiated and give the tunnel that MTU, so hosts and routers learn it directly. Clamping the TCP MSS on the tunnel is the common extra safeguard for TCP, and it is a general mechanism, not an IPsec one.

  • Why might an IPsec gateway be configured to fragment before encryption rather than after?
    Fragments made after encryption must be reassembled at the receiving gateway before ESP can authenticate them (RFC 4303 §3.4.1), which costs buffers and CPU and gives an attacker a reassembly target. Fragmenting the inner packet first means every ESP packet arrives whole; the destination host reassembles. The price is a full ESP overhead per fragment, and it works only in tunnel mode.
  • Why can't IPsec transport mode use the same fragment-before-encrypting approach?
    RFC 4301 defines transport mode SAs as never carrying fragments. With one IP header there is nowhere to keep separate fragmentation state for the plaintext and the ciphertext, so the receiver could not tell pre-IPsec fragments from fragments made after encryption. Tunnel mode's inner and outer headers each keep their own.
  • A router between the two IPsec gateways sends an ICMP fragmentation-needed message; who receives it, and what happens next?
    It goes to the outer source, the encapsulating gateway, since that router saw only the outer header. If configured to act on such unauthenticated messages, the gateway maps it to the SA and lowers the SA's PMTU after subtracting its ESP and outer-header overhead. It then sends its own ICMP to the host when the next oversized packet arrives (RFC 4301 §8.2.1).

saying these in an interview costs you the question

  • The gateway silently fragments every oversized packet, so DF never matters
  • The tunnel's inner MTU is 1,500 minus the 20-byte outer header
  • An IPsec gateway fragments an oversized IPv6 packet like a DF-clear IPv4 one
  • Fragmenting after encryption is free because the far host reassembles
  • ICMP from routers between the gateways goes straight back to the original host