Instances behind an AWS NAT gateway start failing to connect to one busy third-party API while every other destination works fine, and the gateway's ErrorPortAllocation metric is non-zero. What is happening, and what are your options?
answer
- only one destination is failing
- the tuple leaves one field free
- capacity counts per destination, per address
- entries linger after the request ends
- reuse connections before adding addresses
basics
~20 sThe NAT gateway has run out of source ports for that one destination. Its simultaneous-connection capacity is counted per unique destination and per associated IP address, so a single hot endpoint exhausts its port range while every other destination is unaffected. Add addresses, spread the load, or open far fewer connections.
solid answer
~60 sA NAT gateway multiplexes everything behind it onto its own address by rewriting source ports, and its capacity is counted **per unique destination** — repeated connections to the same destination address and port share one port pool. As of this writing AWS documents roughly 55,000 simultaneous connections to a single destination per associated IP address. A service that opens a fresh connection per request to one busy API burns through that pool, the gateway starts failing allocations, and that surfaces as the `ErrorPortAllocation` metric in the `AWS/NATGateway` CloudWatch namespace plus connection failures to that destination only. Three lines of attack: fix the client to reuse connections with keep-alive or a connection pool, so the same request volume needs far fewer flows; associate additional Elastic IPs with the NAT gateway, since each address brings its own port range toward that destination; or split callers across more NAT gateways. Note too that idle flows are held for a while before release, so churn costs more capacity than the concurrent request count suggests.
code
bash · 8 linesaws cloudwatch get-metric-statistics \
--namespace AWS/NATGateway \
--metric-name ErrorPortAllocation \
--dimensions Name=NatGatewayId,Value=nat-0123456789abcdef0 \
--start-time 2026-08-21T00:00:00Z \
--end-time 2026-08-21T06:00:00Z \
--period 300 \
--statistics Sumgo deeper
Know that a NAT gateway rewrites source ports to multiplex many instances onto one address, and that this pool is finite. Recognising ErrorPortAllocation as a NAT gateway metric is already good at this level.
Explain why the ceiling is per destination — the flow tuple leaves only the source port free when the gateway address and destination are fixed — and name the CloudWatch metrics that confirm it.
Diagnose from the shape of the symptom rather than a limit table, and rank the remedies correctly: connection reuse on the client first, extra addresses and gateway spread as capacity, private routing as the structural fix. Mention that idle hold time inflates port consumption far above apparent concurrency.
Own the standard that stops this recurring: connection pooling as a platform default in shared HTTP clients, egress capacity metrics that are alarmed on before they fail, and a policy on which third-party dependencies are allowed to sit on the shared NAT path at all.
## Why the failure is destination-specific The give-away in the symptom is that only *one* destination is failing. That rules out most network causes — a broken route, a missing gateway, or a saturated link would degrade everything — and points straight at how address translation allocates ports. When a NAT gateway forwards an outbound flow, it replaces the instance's private source address with its own and picks a source port so the reply can be matched back to the right originator. A flow is identified by the tuple of protocol, source address, source port, destination address, and destination port. Because the gateway's address and the destination are fixed for a given API, the only field left to vary is the **source port**, and a port range is finite. Two flows to *different* destinations may reuse the same source port harmlessly; two flows to the *same* destination cannot. So the capacity ceiling is per unique destination endpoint, not per gateway. Talking to a thousand different hosts is cheap; hammering one host is what runs out. ## The number, and how to use it AWS documents a NAT gateway as supporting up to about **55,000 simultaneous connections to each unique destination**, per IP address associated with the gateway. Quote it with the caveat that limits move, and then use it the way an engineer should — as an order of magnitude that tells you whether your connection *rate* is plausible, not as a number to recite. The more useful framing is the arithmetic behind the exhaustion. Ports are not released the instant a request finishes; a NAT gateway holds a translation entry for a while after a flow goes idle, and it drops idle TCP flows after a documented **350 seconds**. So the ports in use at any moment are roughly the connection establishment rate multiplied by how long each entry is held — not the number of requests in flight. A service issuing hundreds of new connections per second to one endpoint, each entry held for even a minute, is well into five figures of simultaneous entries while appearing to have only a handful of concurrent requests. That gap between apparent concurrency and actual port consumption is the whole trap. ## Confirming it Three signals, in order of directness: - **`ErrorPortAllocation`** in the `AWS/NATGateway` namespace — the count of times the gateway could not allocate a source port. Any non-zero value is the diagnosis, not a hint. - **`ActiveConnectionCount`** — the concurrent connection count through the gateway, useful for seeing the climb that precedes the failures. - **`IdleTimeoutCount`** — flows dropped for idleness, which tells you entries are being held rather than closed cleanly. ``` aws cloudwatch get-metric-statistics \ --namespace AWS/NATGateway \ --metric-name ErrorPortAllocation \ --dimensions Name=NatGatewayId,Value=nat-0123456789abcdef0 \ --start-time 2026-08-21T00:00:00Z --end-time 2026-08-21T06:00:00Z \ --period 300 --statistics Sum ``` On the application side the failure looks like connection timeouts or resets to one host, often intermittent at first and correlated with peak traffic — which is why teams chase the third party's status page for hours before looking at the gateway. ## The fixes, best first **Reduce the number of flows.** This is the real fix, and it is on the client. An HTTP client that reuses connections — keep-alive with a properly sized connection pool, or HTTP/2 multiplexing where the endpoint supports it — turns tens of thousands of short flows into a handful of long ones. A service that constructs a new client object per request, which is a very common defect, defeats pooling entirely and is usually the root cause. Fixing this also removes the handshake latency you were paying on every call. **Give the gateway more source ports.** You can associate additional Elastic IP addresses with a public NAT gateway; each address carries its own port range toward the same destination, multiplying capacity. It is a genuine fix and cheap to apply, but it is a multiplier on a design that is still burning ports, so treat it as a way to buy time or as headroom for a legitimately high-volume path. **Spread the load.** With per-AZ NAT gateways in place, callers in different zones already use different gateways with different addresses. If a single zone's callers are the problem, splitting them across additional subnets and gateways divides the port pressure. **Take the destination off the NAT path.** Where the destination is reachable by a private route rather than over the internet, the flows never touch the NAT gateway and the ceiling stops applying. That mechanism belongs to another discussion, but it is worth naming as the structural answer when the destination allows it. ## What a strong answer sounds like Start from the symptom's shape — one destination failing, others healthy — and derive the cause rather than reciting a limit. Name the metric that proves it. Then order the remedies with client-side connection reuse first, because adding addresses to absorb a bug the application will keep committing is treating the symptom. Mentioning that idle-entry hold time makes port consumption much higher than the concurrent request count is the detail that shows you have actually debugged this.
- Why does the failure affect only one destination rather than the whole gateway?Because the port pool is consumed per unique destination endpoint. Flows to different destinations differ in the destination fields of the tuple and can reuse the same source port, so they never compete. Only repeated connections to the same address and port contend for the same finite source-port range.
- Why does adding a second Elastic IP to the NAT gateway help?Each associated address gives the gateway a fresh source-port range toward the same destination, so total simultaneous connections to that endpoint scale with the number of addresses. It is a real capacity increase, but it multiplies a design that is still spending ports wastefully, so it should follow rather than replace fixing connection reuse.
- How does the gateway's idle handling make the problem worse than the concurrency numbers suggest?A translation entry survives after the exchange finishes and an idle TCP flow is only dropped after 350 seconds. Ports in use therefore track the new-connection rate multiplied by the hold time, not the number of requests in flight — so a service that looks like it has ten concurrent calls can be holding tens of thousands of entries.
saying these in an interview costs you the question
- Blames the third-party API because only that destination fails
- Thinks the limit is total connections through the gateway
- Reaches for a bigger NAT gateway, as if it had a size
- Adds Elastic IPs and never fixes the client's connection churn
- Assumes ports are freed the moment a request completes