skip to content

A Linux service that opens many short-lived outbound TCP connections to a single backend address starts failing with "cannot assign requested address". What resource has run out, what governs its size, and how would you fix it?

level: seniorimportance: should knowfreq 42%

answer

  1. uniqueness is over four values
  2. the client picks a local port
  3. the range is a sysctl
  4. sixty seconds per closed connection
  5. pool before you tune

basics

~20 s

It has run out of ephemeral source ports for that destination. Linux picks them from net.ipv4.ip_local_port_range, and each closed connection holds its port in TIME_WAIT for 60 seconds, so a high connect rate exhausts the range. Reusing connections through a keep-alive pool is the durable fix.

solid answer

~50 s

The error is `EADDRNOTAVAIL`, and it comes from `connect()`, not from the remote side — the client cannot find a free local port. A TCP connection is identified by the four-tuple of source address, source port, destination address and destination port; with the destination fixed and one source address, the only degree of freedom is the source port, drawn from `net.ipv4.ip_local_port_range` (commonly 32768–60999, about 28,000 values). Because the client is the side closing these connections, each one holds its port in TIME_WAIT for 60 seconds, so the sustainable rate to one destination is roughly the range size divided by 60 — a few hundred connections per second. Widening the range or enabling `net.ipv4.tcp_tw_reuse` buys headroom, and spreading across more source or destination addresses multiplies the space, but the real fix is to stop opening a connection per request: use a keep-alive connection pool.

code

bash · 3 lines
bash
sysctl net.ipv4.ip_local_port_range
sysctl -w net.ipv4.ip_local_port_range="10240 65535"
sysctl -w net.ipv4.tcp_tw_reuse=1

go deeper

for a junior

Know that outgoing connections consume a local port from a limited range, and that the error is raised by your own host rather than by the server you are calling.

for a middle

Explain the four-tuple uniqueness rule, name net.ipv4.ip_local_port_range, and connect the 60-second TIME_WAIT on the closing client to the sustainable connect rate.

for a senior

Diagnose it from the errno EADDRNOTAVAIL, distinguish it from EADDRINUSE and EMFILE, and rank the remedies: connection pooling first, then range width, tcp_tw_reuse and extra source or destination addresses as headroom.

for a principal

Frame it as an architecture signal. Decide where connection reuse belongs across the fleet, set alarms on connect rate per destination rather than on aggregate connections, and refuse tuning proposals that only move the cliff further out.

## The four-tuple is the resource TCP identifies a connection by four values: source IP, source port, destination IP, destination port. Two live connections may not share all four. When an application calls `connect()` without first calling `bind()`, the kernel silently chooses a source port on its behalf — an *ephemeral* port — and that choice must keep the four-tuple unique. The available pool is a sysctl: ``` sysctl net.ipv4.ip_local_port_range net.ipv4.ip_local_port_range = 32768 60999 ``` That is roughly 28,000 ports. Note carefully what the constraint is: the port must be unique *for that destination*, so 28,000 is not a global connection limit for the host. A machine can hold far more than 28,000 connections in total if they go to different destinations. Exhaustion is a per-destination phenomenon, which is exactly why it shows up when a service hammers one backend — one database, one internal API, one proxy — and never shows up in aggregate connection counts. ## Why TIME_WAIT makes it much worse In the request-per-connection pattern, the client is typically the side that closes, which makes it the active closer — and the active closer holds TIME_WAIT for 60 seconds on Linux. During that minute, that source port is unavailable for a new connection to the same destination. The arithmetic follows immediately: about 28,000 ports divided by 60 seconds gives roughly 470 new connections per second to a single destination before exhaustion. Cross that rate and `connect()` starts returning `EADDRNOTAVAIL`, which C reports as "Cannot assign requested address" and higher-level runtimes surface with wording of their own. The failure looks like a network or backend problem in application logs, which is what makes it a good interview scenario: the backend is perfectly healthy, and nothing on the network is wrong. The client has run out of a local kernel resource. ## The remedies, in order of how much they buy **Reuse connections.** This is the real answer and should be said first. An HTTP keep-alive pool, a database connection pool, a gRPC channel — anything that carries many requests over one connection — collapses the connect rate by orders of magnitude and removes the pressure entirely rather than deferring it. Every other measure is headroom. **Widen the range.** Lowering the bottom of `net.ipv4.ip_local_port_range` toward 1024 gains a few thousand ports. Do not overlap ports your host's own services bind, and remember it is a linear improvement against a problem that grows with traffic. **`net.ipv4.tcp_tw_reuse`.** Setting this allows the kernel to reuse a socket sitting in TIME_WAIT for a new *outgoing* connection when TCP timestamps show the old connection's segments cannot be confused with the new one. It targets precisely this client-side case. Its removed cousin `net.ipv4.tcp_tw_recycle` is not an option: it was deleted in Linux 4.12 for breaking clients behind NAT, and recommending it is a red flag. **Add source or destination addresses.** Because uniqueness is over the whole four-tuple, a second source IP doubles the space, and a backend reachable on several addresses or ports multiplies it the same way. If the application binds a specific source address before connecting, set `IP_BIND_ADDRESS_NO_PORT` on the socket — it tells the kernel not to pick a port at `bind()` time, when the destination is still unknown and the port must be globally unique, but to defer the choice to `connect()`, when the destination is known and the port only needs to be unique for that four-tuple. Without it, an explicit source bind can exhaust the range far earlier than necessary. **Change who closes.** If the server can be the one to close, TIME_WAIT accumulates there instead, where it does not consume a scarce client-side resource. This is often not under your control, but it is worth naming. ## Distinguishing the neighbouring failure Do not confuse this with `EADDRINUSE` ("Address already in use"), which is a `bind()`-time failure on a listening socket. `EADDRNOTAVAIL` at `connect()` time on an outbound-heavy client means ephemeral ports; the same client can also hit file-descriptor limits, which produce `EMFILE` ("Too many open files") instead. Three different resources, three different errnos — naming the right one is most of the diagnosis. ## The design lesson Ephemeral port exhaustion is almost always a symptom of a connection-per-request architecture that outgrew its traffic. Sysctl tuning raises the cliff; it does not remove it. The engineer who says "pool the connections, then tune for headroom, and alarm on the connect rate to each single destination" is answering at the level the question is really asked.

  • The host has only about 28,000 ephemeral ports. Does that cap its total outbound connections at 28,000?
    No — the four-tuple must be unique, not the port. The same source port may be reused for a different destination address or port, so the limit is per destination endpoint. That is why a host can hold hundreds of thousands of connections overall yet fail while talking to one busy backend, and why adding backend addresses multiplies the available space.
  • How do you tell ephemeral port exhaustion apart from hitting the process's file-descriptor limit?
    By the errno. Port exhaustion fails in `connect()` with `EADDRNOTAVAIL`, "Cannot assign requested address". A descriptor limit fails earlier, in `socket()` or `accept()`, with `EMFILE`, "Too many open files". They are different resources with different fixes — sysctls and pooling for one, `ulimit -n` or the service's LimitNOFILE for the other.
  • Why does binding an explicit source address before connect() make exhaustion happen sooner, and what fixes that?
    At `bind()` time the destination is unknown, so the kernel must pick a port unique across all destinations rather than merely for the eventual four-tuple. That wastes the range quickly. Setting `IP_BIND_ADDRESS_NO_PORT` on the socket defers the port choice to `connect()`, where the destination is known and the far weaker uniqueness requirement applies.

saying these in an interview costs you the question

  • The backend is refusing connections
  • The host can only ever have 28,000 connections open
  • Enable tcp_tw_recycle to free the ports
  • Just raise ulimit -n and it goes away
  • It is a NAT or firewall problem upstream

context