An office's NAT edge router fails over to a standby with identical static and pool configuration but no session table; why do established connections stall while new ones succeed?
answer
- state lives only in the active box
- replies with no binding to follow
- configuration is not history
- silence, not a reset
basics
~20 sTranslation state lives only in the device that created it. The standby has the configuration but not the bindings and sessions, so replies to existing sessions match nothing and are dropped, while each new session's first packet builds fresh state.
solid answer
~50 sA NAT's mappings are **state**, not configuration. Each session's first packet created an entry on the old active router; replies are un-translated only by finding it. The standby inherits the rules but none of that history. For **dynamic-pool source NAT**, a reply arriving at `203.0.113.19` has no binding saying which inside host owns it, so it is dropped; the inside host's next outbound segment may even be bound to a different pool address. **Static** mappings fare better, because their binding follows from configuration, but whether a mid-session packet is accepted without a session entry is an implementation choice, and many stateful devices drop it. Nobody tells the endpoints: RFC 5382 lets a NAT reset or silently abandon a connection, and after a sudden failure only silence is possible, so applications wait out their own timeouts. New sessions start with a first packet and succeed at once.
go deeper
Remember that a NAT keeps a table of active sessions and needs it to send replies back to the right inside host; a fresh device starts with that table empty.
Explain which part of the table is configuration and which is history, and why a reply to a dynamic-pool binding has nothing to match on a device that never saw the session start.
Diagnose the stall: no resets, retransmissions into a void, rebinding to a different pool address, static flows depending on mid-session acceptance, and the fixes of state sync and application reconnects.
Weigh state replication, static identities and application resilience against each other, and treat every stateful translator as a shared-fate point when designing edge redundancy.
## Why a NAT failure is worse than a router failure A plain IPv4 router forwards each packet on its destination address and keeps no memory of conversations, so when traffic shifts to a parallel router, existing connections continue. A NAT is different. RFC 2993 puts it directly: if a NAT fails, "even when it recovers the communication that was passing through it will still fail (because the NAT no longer translates it using the same mappings)", which it calls more severe than the failure of a router. RFC 3022 section 2.1 adds that when a stub domain has more than one exit point, "it is of great importance that each NAT has the same translation table". The translation table holds two kinds of information: - **Configuration**: static one-to-one mappings and the list of pool addresses. Both devices have this. - **History**: which inside host was bound to which pool address, and which sessions exist, with their tuples and timers. Only the device that saw each session's first packet has this. A standby with no state synchronisation starts with configuration and an empty history. ## What happens to each kind of session Take the office edge with inside network `10.20.0.0/16`, a static mapping of mail relay `10.20.5.25` to `203.0.113.16`, and a dynamic pool `203.0.113.17` to `203.0.113.23`. | Session through the old active router | Packet reaching the standby | Outcome | |---|---|---| | Workstation `10.20.4.31` bound to `203.0.113.19` | inbound reply to `203.0.113.19` | no binding names an inside host; dropped | | Same workstation | its next outbound TCP segment, not a SYN | dropped by devices that require a session, or bound afresh, possibly to another pool address the server does not know | | Mail relay on its static mapping | inbound or outbound mid-session packet | the address rewrite is reproducible from configuration; survives only if the device accepts a packet with no session entry | | Any brand-new session | a TCP SYN or a first UDP packet | new binding and session created; works | The asymmetry in the question follows. RFC 2663 section 2.5 says a NAT recognises the start of a TCP session by the SYN flag set with ACK clear, and treats the first UDP packet with an unseen tuple as a new session. New connections present exactly such a packet; existing ones never will again. When the standby does rebind an existing workstation's outbound packet to a different pool address, the remote server receives a segment from an address and port belonging to no connection it knows. It discards the segment or answers with a reset, and either way the original connection does not recover. ## Why the endpoints hang instead of failing fast RFC 5382 section 5 says that a NAT abandoning a live TCP connection MAY send TCP RST packets to both endpoints or MAY abandon it silently, and notes that notification is impossible when state is lost to a power failure. A crashed active router sends nothing, and the standby does not know the sessions existed. So: 1. Each side keeps retransmitting unacknowledged data with exponential backoff. 2. Idle connections notice nothing until they next send, or until a keepalive fires. 3. The application sees a stall that ends only at its own or TCP's timeout. Inbound ICMP errors quoting packets of the lost sessions do not help either: RFC 5508 says a NAT with no active mapping for the quoted packet SHOULD silently drop the error. ## Designing around it - **Synchronise state between the pair.** Many devices offer session-table replication to the standby as an implementation feature; it is what makes failover invisible, at the cost of replication traffic and a consistency window. - **Prefer static mappings for long-lived critical flows.** A binding derived from configuration can at least be reproduced on the standby. - **Make applications recover.** Keepalives, idle timeouts and reconnect logic turn an indefinite stall into a quick reconnection. - **Keep both directions on one translator.** State is useful only where the packets arrive; replies routed through a different translator meet the same empty table. The general lesson, from RFC 2993's end-to-end argument, is that any device holding per-session state becomes a point where the sessions' fates are shared.
- Why does a static one-to-one mapping survive a stateless failover better than a dynamic pool binding?A static binding is a pure function of configuration: inside `10.20.5.25` always becomes `203.0.113.16`, so the standby rewrites the same way. A dynamic binding records history, which pool address happened to be free when the host first sent. The standby cannot reconstruct it and may pick another address. The static flow still needs the device to accept a mid-session packet without a session entry, which is an implementation choice.
- Why do the endpoints not receive a TCP reset when the active translator dies?RFC 5382 lets a NAT that abandons a live connection send RSTs or stay silent, and notes that notifying endpoints is impossible when state is lost to a power failure. A crashed device sends nothing and the standby does not know the session existed, so the endpoints only retransmit and wait for their own timeouts.
saying these in an interview costs you the question
- Copying the NAT configuration to the standby is enough for seamless failover.
- Established connections survive because both routers use the same public address pool.
- The standby rebuilds dynamic bindings automatically from the first reply it receives.
- A NAT failover is no worse than a router failover, since routing reconverges.
- The translator always sends TCP resets when it loses a session.