In a VPN tunnel behind a NAT, which peer must send the keepalives, and how should their interval relate to the translator's UDP mapping timer?
answer
- direction of the refresh
- outbound refresh is the required one
- shorter than the shortest timer
- 20 seconds, only when idle
- not a liveness signal
basics
~20 sThe peer behind the NAT must send them, because RFC 4787 only requires outbound packets to refresh a mapping. The interval must undercut the shortest UDP mapping timer on the path; RFC 3948 defaults to 20 seconds of idleness.
solid answer
~50 sThe peer **behind** the translator sends them. RFC 4787 requires a NAT's mapping to be refreshed by outbound packets and only *allows* inbound refresh, so the gateway's packets cannot be relied on - and once the mapping is gone they are dropped anyway. IKEv2 says the same: the side that finds itself behind a NAT should start sending RFC 3948 keepalives. The interval must be shorter than the shortest UDP mapping timer on the path. RFC 4787 sets a two-minute floor and recommends five minutes, but real translators vary, so RFC 3948 sends a one-octet keepalive after **20 seconds** with nothing else sent. WireGuard's persistent keepalive is optional; the paper gives no interval and the usual 25 seconds is an implementation's. Keepalives hold the mapping only: RFC 3948 forbids using them to judge whether the peer is alive.
go deeper
Remember that a NAT forgets idle UDP flows, so the peer behind it sends small keepalives, 20 seconds by default for IPsec.
Explain why only outbound packets reliably refresh a mapping, how RFC 3948 decides when to send a keepalive, and why that is separate from checking whether the peer is alive.
Show that you choose the interval from the shortest timer your users actually meet, not from the RFC floor, and that you know which side of the tunnel has to send.
Weigh battery and gateway packet rate against reachability, and decide which clients need inbound reachability enough to justify keepalives at all.
## What a keepalive is for A **NAT mapping** is the translator's memory that internal address and port A:a are currently represented by external address and port X:x. For UDP there is no connection setup or teardown to watch, so the translator runs a **mapping timer**: if no packet uses the mapping for long enough, it is deleted. A VPN tunnel that goes quiet - a laptop on a coffee break, a phone in a pocket - loses its mapping, and the gateway can no longer reach it. A **NAT keepalive** is a small packet sent while the tunnel is otherwise idle, only to keep that timer from firing. ## Which peer sends it The direction matters more than the interval: - **RFC 4787, REQ-6**: a NAT *must* refresh a mapping on **outbound** packets (inside to outside). Refresh on inbound packets is only *permitted*, and the RFC notes that letting outside packets keep a mapping alive is itself a security risk. **RFC 7857** later added that inbound packets which match a mapping but are refused by policy *should not* refresh it. - So a keepalive sent by the gateway on the public side may or may not refresh anything, and once the mapping has expired the translator has nowhere to deliver it. - **IKEv2 (RFC 7296)** builds this in: when NAT detection shows a peer that it is the one behind a translator, that peer *should* start sending the keepalives defined in RFC 3948. - **WireGuard's** persistent keepalive is configured on the peer that needs its mapping held open - in remote access, the client behind the NAT. ## How long the interval can be The interval has to be shorter than the **shortest mapping timer on the path**, and the sender cannot see that timer. | Source | What it says | Whose value | |---|---|---| | RFC 4787 REQ-5 | UDP mapping timer must not expire in under **2 minutes**; **5 minutes or more** recommended | A requirement on NATs (BCP 127) | | RFC 4787 REQ-5a | Shorter timers allowed for specific well-known destination ports (0-1023) | A permitted exception | | RFC 3948 section 4 | Send a keepalive if nothing else went to the peer for **M = 20 s** (default) | The IPsec specification | | WireGuard whitepaper | Persistent keepalive is optional; no interval given | The paper | | Implementation documentation | **25 s** is the common WireGuard setting | An implementation, not the protocol | Why does RFC 3948 use 20 seconds when RFC 4787 promises two minutes? RFC 3948 was published in 2005 and RFC 4787 in 2007, and RFC 4787 itself records great variation among deployed translators. A floor is a requirement on NATs, not a guarantee that every NAT on a hotel or mobile path meets it, so tunnel specifications pick a conservative default. RFC 3948 ties the keepalive to need: a peer *should* send one when a NAT was detected and no other packet went to the peer in the last `M` seconds. A busy tunnel's own traffic keeps the mapping alive, so the keepalive only fills the silences. ## Keepalive is not liveness These packets are easy to confuse with peer-health checks, and the specifications keep them apart: - **RFC 3948**: reception of NAT-keepalive packets **must not** be used to decide whether a connection is live. IKEv2 checks liveness with an empty `INFORMATIONAL` exchange (sometimes called dead peer detection). - **WireGuard's passive keepalive** is a different mechanism from the persistent one. After **Keepalive-Timeout (10 s)**, a peer that has received data but has nothing to send replies with an empty authenticated message. It only answers traffic, so a fully idle tunnel stays silent and its mapping can still expire. Only the persistent keepalive holds an idle mapping. - Over TCP, **RFC 9329** says ESP NAT-keepalives *should not* be sent, because TCP mappings last far longer than UDP ones (RFC 5382 sets at least 2 hours 4 minutes for an established TCP connection). ## What it costs Keepalives are small but constant: 1. An RFC 3948 keepalive is a 1-octet payload + 8-byte UDP header + 20-byte IPv4 header = **29 bytes**. At one per 20 seconds that is 3,600 / 20 = **180 packets per hour**, about 5,220 bytes. 2. A WireGuard keepalive is a transport message with a zero-length packet: 16 bytes of header plus a 16-byte authentication tag = 32 bytes of UDP payload, or **60 bytes** with UDP and IPv4. At 25 seconds: 3,600 / 25 = 144 per hour, about 8,640 bytes. 3. At the gateway, 10,000 idle clients at 20 seconds add 10,000 / 20 = **500 packets per second**. The bytes are trivial; the real cost is waking a mobile device's radio and keeping it awake. Too short wastes battery, too long loses the mapping, and the right value is the shortest timer your users' networks actually use.
- Can a VPN client decide that its IPsec gateway is dead because NAT keepalives stopped arriving?No. RFC 3948 forbids using received NAT-keepalives to judge liveness; they exist only to hold the translator's mapping, and only the peer behind the NAT is required to send them. IKEv2 checks liveness with an empty INFORMATIONAL exchange, which is authenticated and expects a reply.
- Does an IPsec tunnel carried over TCP under RFC 9329 still need NAT keepalives?RFC 9329 says ESP NAT-keepalives should not be sent over TCP, and a receiver must silently drop any that arrive. TCP mappings are kept far longer than UDP ones; RFC 5382 sets at least 2 hours 4 minutes for an established connection.
saying these in an interview costs you the question
- The gateway's keepalives hold the client's NAT mapping open.
- Any interval under five minutes is safe, because RFC 4787 guarantees it.
- WireGuard's protocol mandates a 25-second keepalive.
- Receiving NAT keepalives proves the IPsec peer is still alive.
- WireGuard's passive keepalive keeps an idle tunnel's mapping open.