After a laptop's VPN tunnel through carrier-grade NAT sits idle, the gateway cannot reach the laptop until it sends traffic; why, and how do UDP and TCP tunnels recover?
answer
- the translator forgot, not the peers
- only the inside can open a mapping
- new source port, same keys
- authenticated packet updates the endpoint
- TCP Originator reconnects
basics
~20 sThe translator's idle mapping expired, so the gateway's packets have nowhere to go until the laptop creates a new one. UDP tunnels recover on the laptop's next authenticated packet; TCP tunnels need the laptop to open a new connection.
solid answer
~50 sThe tunnel's keys are fine; the **translator** forgot the mapping because nothing crossed it for longer than its timer. Only an outbound packet can create a new one, so anything the gateway sends - data, a handshake, a liveness check - is dropped until the laptop speaks. Its next packet often arrives from a new external port, sometimes a new address. A **WireGuard** hub accepts it once it authenticates and updates the peer's endpoint, as the whitepaper's roaming design describes. An **IKEv2** gateway that is not behind a NAT should do the same for an integrity-validated packet (RFC 7296), and MOBIKE gives an explicit alternative. A **TCP-carried** tunnel is harder: the connection itself is dead, so the laptop must open a new one; RFC 9329 makes the TCP Originator reconnect and the responder accept existing SAs from the new address. The fix for the gateway's blind spot is a keepalive from the laptop.
go deeper
Recall that the NAT, not the VPN, forgets an idle tunnel, and that only the client behind the NAT can recreate the mapping.
Explain how WireGuard and IKEv2 follow a peer whose external port changed, and why the keys and security associations survive.
Diagnose the one-way blackout: gateway-initiated traffic fails while client traffic heals it. Contrast UDP recovery with RFC 9329's TCP reconnect and say which side owns each step.
Decide which users need gateway-initiated reachability, and price inside-originated keepalives and reconnect storms against that need.
## What actually broke Two things hold a remote-access tunnel together: the **cryptographic state** in both peers (keys, sequence counters, the security association) and the **translation state** in every NAT on the path. An idle period damages only the second. Behind **carrier-grade NAT** - a translator run by the access provider, shared by many subscribers - the laptop has no control over the timer and often no idea the translator exists. When nothing crosses the mapping for longer than its UDP timer, the translator deletes it. From then on: - The gateway still believes the laptop is at the old external address and port. - Anything the gateway sends there - tunnelled data, a WireGuard handshake initiation, an IKEv2 liveness check - reaches a translator with no matching mapping and is typically discarded. - Nothing from the outside can repair this. **RFC 4787** requires mappings to be refreshed by outbound packets, inbound refresh is optional, and an unsolicited inbound packet does not create a mapping. So the gateway is blind until the laptop sends something. ## Recovery in UDP tunnels The laptop's next outbound packet creates a **new mapping**, often with a different external port and, behind a carrier pool, possibly a different external address. What the gateway does with a packet from an unfamiliar source decides whether recovery is instant: - **WireGuard.** The whitepaper's roaming design uses the outer source address of every *correctly authenticated* packet as the peer's current endpoint. The hub replies to the new port with no handshake and no configuration change. If the idle spell outlasted **Reject-After-Time (180 s)**, the old session is unusable and a fresh 1-RTT handshake runs first, started by whichever side has data to send. If the hub tries first, its initiations go to the dead mapping, are retried every **Rekey-Timeout (5 s)**, and give up after **Rekey-Attempt-Time (90 s)**. - **IKEv2 with ESP-in-UDP.** **RFC 7296** says a host that is not behind a NAT *should* send all packets to the address and port of a newly received, integrity-validated packet and store them as the SA's new address. Because an old replayed packet could move the SA back, this is only safe with replay protection. The peer *behind* the NAT *should not* follow changes this way, since that would let a single forged packet break the connection. **MOBIKE (RFC 4555)** replaces the implicit update with an explicit `UPDATE_SA_ADDRESSES` exchange. Either way, keys and SAs survive; only the address the gateway sends to changes. ## Recovery in TCP tunnels TCP mappings last much longer: **RFC 5382** says an established connection's idle timer must not be shorter than **2 hours 4 minutes**. When one does expire, or a translator reboots, the TCP connection itself is dead, and both TCP stacks find out the slow way, through retransmission timeouts or a reset. 1. **RFC 9329** (IKE and ESP over TCP) makes the **TCP Originator** - in remote access, the laptop - responsible for re-establishing the connection. 2. It opens a new TCP connection, sends the `IKETCP` stream prefix again, and continues the existing IKE session; with TLS on top, a new TLS handshake is needed, and RFC 9329 recommends session resumption. 3. The **TCP Responder** *must* accept packets for existing SAs from the new source address and port. 4. Unless a `DELETE` was sent, both sides keep their SAs for their normal lifetime, so the IKE SA is not renegotiated. A TLS-based VPN follows the same shape - the client reconnects, then resumes or re-authenticates - but how it does so is each implementation's design, since no RFC defines that family. ## Side by side | Tunnel | What the idle spell kills | Who recovers it | Cost of recovery | |---|---|---|---| | WireGuard | The UDP mapping | Laptop's next authenticated packet | None, or one handshake after 180 s | | IKEv2, ESP-in-UDP | The UDP mapping | Laptop's next validated packet, or MOBIKE | None for the SAs | | IKE and ESP over TCP | The TCP connection | Laptop as TCP Originator | TCP (and TLS) handshake | | TLS-based VPN on TCP | The TCP connection | Client reconnect | Implementation's reconnect | ## What fixes the gateway's blind spot Recovery always waits for the laptop. That is fine for a user who opens a browser, and not fine for anything the gateway or the corporate side must initiate: a push, an incoming call, a management connection to the laptop. - Send **keepalives from the laptop**, at an interval shorter than the shortest UDP timer on its networks; RFC 3948 defaults to 20 seconds, WireGuard's persistent keepalive is optional. - Do not expect gateway-side settings to help: a longer gateway liveness timer only waits longer, and gateway-originated keepalives cannot recreate a mapping. - Check liveness with the protocol's own mechanism - IKEv2's `INFORMATIONAL` exchange, WireGuard's reply expectation - not with NAT keepalives, which RFC 3948 forbids using for that.
- Does a WireGuard hub need configuration to follow a client whose external port changed?No. The whitepaper's roaming design takes each correctly authenticated packet's outer source address as the peer's current endpoint, so the hub replies to the new port automatically. An on-path attacker can rewrite that unauthenticated outer address, but the paper notes it gains only denial of service, which it already had by dropping packets.
- If the translator reboots rather than timing out, what changes for a gateway with thousands of clients?Every client's mapping disappears at once. UDP clients recover on their next packet, so the gateway sees a burst of endpoint updates. TCP-carried clients all lose their connections and reconnect together, so the gateway takes a wave of TCP and TLS handshakes; RFC 9329 recommends TLS session resumption for this kind of re-establishment.
saying these in an interview costs you the question
- The tunnel's keys must be renegotiated whenever the client's NAT port changes.
- The gateway can reopen the client's mapping by sending keepalives towards it.
- A TCP-carried tunnel survives a NAT reboot because TCP retransmits.
- Raising the gateway's liveness timeout fixes unreachable idle clients.
- Carrier-grade NAT holds UDP mappings for five minutes, so idle tunnels cannot expire.