Why is a stateful IPv4 NAT's session table a denial-of-service target, and what do the NAT RFCs require when it cannot create a new mapping?
answer
- every new flow costs state
- minimum timers keep entries alive
- one subscriber versus everyone
- drop the new, keep the old
- soft error versus prohibited
basics
~20 sEvery new flow costs translator memory and lingers for the RFC minimum timers, so one scanning or flooding host can fill the table for everyone. When full, a NAT drops the new packet with an ICMP error and keeps existing mappings.
solid answer
~50 sA NAT keeps a mapping, and usually per-destination session and filter state, for every flow, and the behavioural RFCs set minimum idle lifetimes: 2 minutes for UDP (RFC 4787 REQ-5), 2 hours 4 minutes for established TCP and 4 minutes for partially open or closing TCP (RFC 5382 REQ-5). Table size is roughly new-flow rate times lifetime, so an infected or scanning inside host, or a SYN flood against a forwarded port, can exhaust it for every user; unsolicited inbound packets that match no mapping create no mapping. RFC 6888 asks a carrier-grade NAT for per-subscriber limits on ports and state (REQ-4, REQ-5), and REQ-11 says that when it cannot create a mapping it MUST drop the packet, SHOULD send ICMP Destination Unreachable code 1, and MUST NOT delete existing mappings to make room. RFC 5508 REQ-8 names code 13 for NATs in general.
code
pseudocode · 10 lineson outbound packet p with no matching mapping:
sub = subscriber_of(p.source_address)
if sub.mapping_count >= sub.mapping_limit or table.is_full():
drop(p)
send_icmp(to = p.source_address, type = 3, code = 1)
notify_management()
return # existing mappings are never evicted
m = create_mapping(p.source_address, p.source_port)
sub.mapping_count = sub.mapping_count + 1
translate_and_forward(p, m)go deeper
Recall that a NAT stores an entry for every connection passing through it, and that this storage can run out.
Explain how new-flow rate and the minimum idle timers set the table's size, and why scanning hosts and floods against forwarded ports fill it.
Show the production response: per-subscriber limits, rate limits and transitory timers, the drop-new-keep-old rule, and telling state exhaustion apart from port exhaustion.
Plan capacity for a shared translator: sizing against worst-case subscribers, fairness limits, and how timer policy trades protection against breaking idle sessions.
## What the table holds A stateful NAT keeps an entry for every **mapping** (inside address and port to public address and port) and, alongside it, per-destination **session** and **filter** state; RFC 6888 REQ-5 names sessions and filters as the state an operator should be able to limit. Every entry costs memory and lookup work, and the table is finite. Unlike a router's forwarding table, which grows with the topology, this one grows with **traffic**. ## Why it fills, and who can fill it Entries are created by **new flows** and survive until they have been idle for at least the minimum the behavioural RFCs set: | Flow | Minimum idle lifetime | Source | |---|---|---| | UDP mapping | 2 minutes; 5 minutes or more recommended (shorter allowed for specific well-known destination ports) | RFC 4787 REQ-5 | | TCP, established | 2 hours 4 minutes | RFC 5382 REQ-5 | | TCP, partially open or closing | 4 minutes | RFC 5382 REQ-5 | The table's occupancy is therefore roughly **new-flow rate x how long each entry lingers**. Typical ways it fills: - An inside host that is infected or scanning opens thousands of connections a second to random destinations; every SYN that gets no answer sits in transitory state. - An application that sends many short UDP exchanges to many destinations. - An inbound SYN flood against a **forwarded** service: each attempt matching a forwarding rule creates session state. - By contrast, unsolicited inbound packets that match no mapping create **no mapping** - they are not forwarded. At most, a NAT following RFC 5382 REQ-4 holds an unsolicited SYN for at least 6 seconds before answering it, and it may drop such SYNs silently under attack - so an outside attacker usually needs a forwarded port, or a foothold inside, to fill the table. In a shared translator, one subscriber's flood exhausts the table for everybody, which is why RFC 6888 treats per-subscriber limits as a fairness and denial-of-service question. ## What the RFCs require when it is full RFC 6888 REQ-11 sets the rule for a carrier-grade NAT that cannot create a dynamic mapping because of resource constraints or quotas: 1. It MUST drop the packet that would have created the mapping. 2. It SHOULD send an ICMP Destination Unreachable with **code 1 (Host Unreachable)** to the sender - preferred over code 13 because RFC 1122 lists it as a soft error, so the sender's TCP does not abort the attempt outright. 3. It SHOULD notify a management system, if one is configured. 4. It MUST NOT delete existing mappings to make room for the new one. For NATs in general, RFC 5508 REQ-8 says to drop the packet and send an ICMP Destination Unreachable with **code 13 (Communication Administratively Prohibited)**. The shared principle: failing a **new** connection is better than killing an **established** one, because applications cope with setup failures far better than with a working session vanishing. ## Containing it - **Per-subscriber limits** - RFC 6888 REQ-4 on the number of external ports and REQ-5 on state memory - plus a rate limit on how fast one subscriber may create mappings. - **Shorter transitory timers** where security matters: RFC 7857 lets a NAT configure the partially-open timeout below 4 minutes, while still recommending 4 minutes as the default. - **Active liveness checks** while under attack, which RFC 5382's security considerations allow instead of waiting out passive timers. - Do **not** answer it by cutting the established-TCP timeout far below the RFC minimum on its own: healthy but idle connections then vanish without any endpoint closing them. ## Watching for it The early signal is not a full table but its slope: occupancy that climbs with no matching rise in users, many entries in the transitory state, or one inside address owning a disproportionate share of sessions. Those point at a scanning host or a flood against a forwarded service long before new connections start failing, and they are far cheaper to act on than an outage. ## State versus ports Keep this separate from **port exhaustion**: an address-and-port translator can also run out of external ports on its public addresses while it still has memory to spare. Both are capacity limits of the same box, but the remedies differ - more public addresses or wider port allocation for ports, limits and timer policy for state. In an interview, naming which resource ran out is half the diagnosis.
- Why does RFC 6888 prefer ICMP code 1 over code 13 when a carrier-grade NAT cannot create a mapping?RFC 1122 lists Host Unreachable as a soft error, so a TCP sender treats it as advisory and keeps retrying rather than aborting the connection attempt at once. Since the shortage may be momentary, the CGN wants the attempt to survive; RFC 5508's general NAT rule uses code 13, Communication Administratively Prohibited, instead.
- Why not simply evict the oldest mappings when the table is full?Evicting kills sessions that are working, and applications handle a failed connection setup far better than an established connection silently breaking. RFC 6888 REQ-11 therefore says a CGN MUST NOT delete existing mappings to make room; the new flow is the one that fails.
saying these in an interview costs you the question
- Unsolicited inbound packets fill a NAT's mapping table directly.
- A full NAT should evict the oldest mappings to admit new flows.
- Cutting the established-TCP timeout to a few minutes is a safe fix.
- Table exhaustion and port exhaustion are the same problem with the same fix.
- One subscriber's traffic cannot affect other users of a shared NAT.