skip to content

For a remote-access VPN serving staff on hotel Wi-Fi, carrier-grade NAT and locked-down guest networks, how do you choose the transports, the fallback order and the keepalive policy?

level: principalimportance: nice to knowfreq 10%

answer

  1. measure where users connect from
  2. UDP first, TCP for reach
  3. how long before falling back
  4. keepalives only where inbound matters
  5. log which rung each session used

basics

~20 s

Offer UDP first and a TCP or TLS rung on 443 for networks that drop UDP, fall back after a short UDP retry and re-probe UDP on every move, and send inside-originated keepalives only for clients that must be reachable.

solid answer

~50 s

Start from data: which networks your users meet, how many drop UDP, and how short their UDP mapping timers are. Then build a ladder. **UDP first** - ESP-in-UDP or WireGuard - because it keeps loss behaving like loss. A **TCP rung** for UDP-blocking networks: RFC 9329 TCP encapsulation on 4500 or a configured 443, with TLS if middleboxes demand it, or a TLS-based VPN; a WireGuard-only design has no such rung. Fall back after at least one UDP retransmission, remember per network what worked, and **re-probe UDP after each move** so sessions do not stay on TCP. Send keepalives from clients only where inbound reachability matters, at an interval below the shortest timer you measured; suppress them over TCP. Finally, log which rung every session uses - that number tells you what the design costs.

go deeper

for a junior

Recall the ladder: UDP when possible, a TCP or TLS fallback for networks that drop UDP, and keepalives from the client side.

for a middle

Explain why UDP is tried first, what RFC 9329 says about the fallback, and why only the client's keepalives reliably hold a mapping.

for a senior

Show how you would detect sessions stuck on TCP, tune the fallback wait, and pick a keepalive interval from measured timers rather than the RFC floor.

for a principal

Own the trade between reach, performance, battery and gateway cost, and the choice between one TCP-capable protocol and a mixed estate. Back each choice with data on where users connect.

## Start from where users connect There is no single correct transport for remote access; there is a population of networks, and the design should be fitted to it. Before choosing, find out: - What share of sessions start on networks that **drop UDP** entirely (hotel, guest, some corporate visitor networks). - How short the **UDP mapping timers** are on the paths users actually take, especially mobile carrier-grade NAT. RFC 4787 requires at least two minutes and recommends five, but records great variation in practice. - Which clients need to be **reachable from the inside** - pushed policy, incoming calls, management sessions to the device - and which only ever start traffic themselves. - What your gateways can afford: per-session TCP and TLS state, handshake rate, keepalive packet rate. ## Build the ladder | Rung | Example transports | Gains | Costs | |---|---|---|---| | 1. UDP | ESP-in-UDP on 4500 (RFC 3948), WireGuard, a TLS-based VPN's UDP data channel | Loss stays loss; best latency and throughput | Fails where UDP is dropped; mappings need keepalives | | 2. TCP | IKE and ESP over TCP 4500 (RFC 9329) | Passes networks that allow TCP but not UDP | TCP-in-TCP, shared head-of-line blocking, one QoS class | | 3. TCP or TLS on 443 | RFC 9329 on a configured 443 with optional TLS and HTTP `CONNECT`, or a TLS-based VPN | Passes networks that allow little but web traffic | Rung 2's costs plus TLS overhead and handshakes | The protocol choice constrains the ladder. **WireGuard** gives an excellent rung 1 and nothing else, because its whitepaper defines only UDP. **IKEv2** can cover all three, but only if both client and gateway are configured for RFC 9329, since TCP support is not negotiated. **TLS-based VPNs** often cover all three in one design, defined by each implementation rather than an RFC. A mixed estate - a UDP-only protocol plus a separate TCP-capable fallback - doubles what you operate. ## Choose the fallback order and timing 1. **Try UDP first.** RFC 9329 says an initiator with no prior knowledge should, and that implementations must favour direct or UDP-encapsulated ESP whenever possible. 2. **Fall back after a short, bounded wait.** RFC 9329 recommends at least one UDP retransmission before switching. Too long and users stare at a spinner on every hotel network; too short and clients land on TCP when UDP was merely slow. 3. **Remember what worked.** RFC 9329 lets implementations shorten the timeout from historical data, so a network that blocked UDP yesterday can go straight to TCP. 4. **Re-probe UDP after every move.** With MOBIKE, the initiator sends `UPDATE_SA_ADDRESSES` over UDP first after a network change and returns to UDP if it answers. Without this, a session that touched one bad network stays on TCP for the rest of the day. ## Decide the keepalive policy - Keepalives must come **from the client**, because NATs are required to refresh mappings on outbound packets only. - They matter only for clients that must be **reachable while idle**. A client that only initiates traffic recovers on its next packet: WireGuard and IKEv2 gateways both follow a validated packet's new source address. - Pick the interval from the **shortest timer you measured**, not from RFC 4787's floor. RFC 3948's default is 20 seconds; WireGuard's persistent keepalive is optional, with 25 seconds a common implementation setting. - Account for the cost at both ends: 12,000 idle clients at 20 seconds add 12,000 / 20 = 600 packets per second at the gateway, and every keepalive can wake a phone's radio. - **Suppress keepalives on the TCP rungs**, as RFC 9329 directs, since TCP mappings last far longer. ## Operate it - **Log the rung** every session lands on and the time it took to get there. A rising TCP share is a performance problem you can see before users report it. - **Size the TCP rungs for reconnect storms**: when a translator reboots, every TCP-carried client reconnects at once, and TLS session resumption, which RFC 9329 recommends, softens the handshake load. - Remember that hiding on 443 is not invisibility: RFC 9329 notes that operators blocking IPsec on purpose can filter its stream prefix, so a guest network's policy may still win.

  • How would you set the timeout before a client abandons UDP for TCP?
    Long enough for at least one UDP retransmission, as RFC 9329 recommends, so a slow network is not mistaken for a blocking one; short enough that users on UDP-blocking networks are not left waiting. Then let history shorten it: a client that saw a network block UDP before can go straight to the TCP rung there.
  • When would you leave keepalives switched off for a group of clients?
    When those clients never need to be reached while idle. Their own next packet recreates the mapping, and both WireGuard and IKEv2 gateways follow the new source address, so the only cost of expiry is a slower first packet. Turning keepalives off then saves battery and gateway packet rate.

saying these in an interview costs you the question

  • Put everyone on TCP 443 so connection tickets go away.
  • One keepalive interval fits every platform and every network.
  • Once a session falls back to TCP, it may as well stay there.
  • A WireGuard-only design covers networks that drop UDP.
  • Tunnelling on 443 means no network operator can identify or block it.