You stop a busy TCP server on Linux and start it again immediately, and it refuses to start with "Address already in use" even though no process is holding the port. Why does bind() fail, and what makes the restart succeed?
answer
- no process holds it, the kernel does
- the active closer waits
- twice the maximum segment lifetime
- one setsockopt call, before bind
- not the same option as REUSEPORT
basics
~20 sConnections the old server closed are still in TIME_WAIT, so the kernel still has sockets on that local port and refuses a plain bind. Setting SO_REUSEADDR on the listening socket before bind lets the server take the port back despite them.
solid answer
~50 sThe port is not held by a process — it is held by kernel state. Whichever side closes a TCP connection first sits in TIME_WAIT for twice the maximum segment lifetime, 60 seconds on Linux, so a server that just dropped thousands of connections leaves thousands of TIME_WAIT sockets on its local address and port. By default `bind()` refuses a local address that any of those sockets still occupy, and you get `EADDRINUSE`. The fix is `SO_REUSEADDR`, set with `setsockopt` *before* `bind`: it tells the kernel to allow the bind even though TIME_WAIT sockets exist on that address, while still refusing if another socket is actively listening there. Practically every server framework sets it by default. TIME_WAIT itself is not a bug to be tuned away — it exists so the final ACK can be retransmitted and so stray segments from the old connection cannot be mistaken for the new one.
code
python · 7 linesimport socket
s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
s.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
s.bind(("0.0.0.0", 8080))
s.listen(128)
print("listening on", s.getsockname())go deeper
Recall that the port can be occupied by leftover kernel state rather than by a running process, and that servers set SO_REUSEADDR on the listening socket to restart cleanly.
Explain the mechanics: the endpoint that closes first holds TIME_WAIT for 60 seconds on Linux, those sockets keep the local address occupied, and SO_REUSEADDR must be set before bind to override it.
Show you know why TIME_WAIT exists — retransmitting the final ACK and expiring stray segments — and push back on the folklore fixes, naming tcp_tw_recycle as removed in Linux 4.12 and tcp_fin_timeout as governing a different state.
Treat a large TIME_WAIT population as a design signal rather than a sysctl problem: decide which side should close, adopt keep-alive and connection pooling, and weigh the correctness guarantee TIME_WAIT provides before anyone proposes shortening it.
## Why a port stays busy with no process holding it A TCP connection is identified by four values: source address, source port, destination address, destination port. A listening socket is a much looser thing — a local address and port waiting for any peer. Both live in the same kernel table, and `bind()` consults that whole table, not the process list. That is why `lsof`-style reasoning ("nothing owns the port") and the kernel's answer ("the address is in use") can disagree. ## TIME_WAIT: what it is and who gets it TCP shuts down with a four-way exchange of FIN and ACK. The endpoint that sends the **first** FIN — the *active closer* — ends in the TIME_WAIT state after the exchange. The passive closer goes straight to CLOSED. This is the fact candidates most often get backwards: TIME_WAIT is not a server state or a client state, it is the *active closer's* state. A server that closes connections after each response accumulates TIME_WAIT; a server that lets clients close does not. TIME_WAIT lasts 2×MSL (twice the maximum segment lifetime). On Linux that duration is a compile-time constant of **60 seconds** and is not exposed as a sysctl. A frequent mistake is to lower `net.ipv4.tcp_fin_timeout` believing it shortens TIME_WAIT; it does not — it governs how long a socket may sit in FIN_WAIT_2 waiting for the peer's FIN. The state exists for two real reasons: 1. **The last ACK may be lost.** If the active closer's final ACK vanishes, the peer retransmits its FIN. Only a socket that still exists can answer it; a fully closed socket would answer with a RST, and the peer would log a spurious error. 2. **Stray segments must not contaminate the next connection.** Duplicated or delayed packets from the old connection can still be in flight. Holding the four-tuple for 2×MSL guarantees they expire before the same four-tuple can be reused. ## Why bind() refuses When your server exits, every connection it actively closed leaves a TIME_WAIT socket whose *local* address and port are the listening address and port. The new process asks to bind that same address and port. Without any option set, the kernel sees existing sockets on that local address and returns `EADDRINUSE`. A quiet server has none and restarts fine; a busy one that just dropped ten thousand connections fails for the next minute. This is why the bug reliably appears in production and never in local testing. ## SO_REUSEADDR is the intended answer ```python import socket s = socket.socket(socket.AF_INET, socket.SOCK_STREAM) s.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1) s.bind(("0.0.0.0", 8080)) s.listen(128) ``` Two details matter. First, it must be set **before** `bind()`; setting it afterwards changes nothing. Second, understand precisely what it grants on Linux: permission to bind a local address that has sockets in TIME_WAIT (and to bind a specific address when a wildcard bind exists, and vice versa). It does **not** let two processes hold live listening sockets on the same address and port — that is a different option, `SO_REUSEPORT`, with different semantics and a different purpose. Most server libraries set `SO_REUSEADDR` for you. Python's `socketserver` exposes `allow_reuse_address`; nginx, Apache and the Go and Java network stacks handle it internally. If you are writing a raw socket server, it is on you. ## The bad advice to recognise Interviewers often probe here because the internet is full of harmful tuning folklore: - **`net.ipv4.tcp_tw_recycle`** — aggressive TIME_WAIT recycling that broke connections from clients behind NAT, because it relied on per-source-IP timestamp monotonicity. It was **removed from Linux in 4.12** and does not exist on any current kernel. Recommending it is a red flag. - **`net.ipv4.tcp_fin_timeout`** — does not shorten TIME_WAIT, as above. - **`net.ipv4.tcp_tw_reuse`** — this one is real and useful, but it applies to *outgoing* connections: it lets the kernel reuse a TIME_WAIT socket for a new outbound connection when TCP timestamps make it safe. It does not help a listening socket bind. ## Framing it well in an interview The strong answer separates three things: the state (TIME_WAIT, held by the active closer for 60 seconds, protecting correctness), the symptom (`bind()` returning `EADDRINUSE` because kernel sockets, not processes, occupy the address), and the fix (`SO_REUSEADDR` before `bind`, which is a correctness-preserving permission and not a workaround). Then add the judgement: large TIME_WAIT counts on a server are usually a *design* signal — the server is closing connections it could keep alive, or clients are not reusing connections — rather than a number to suppress with sysctls.
- Does SO_REUSEADDR let two processes listen on the same port at the same time?No. On Linux it permits a bind despite sockets in TIME_WAIT, and permits mixing a wildcard bind with a specific-address bind — but a second live listener on the identical address and port is still rejected. Sharing a port between live listeners requires `SO_REUSEPORT`, which every socket in the group must set before bind.
- Your server shows tens of thousands of TIME_WAIT sockets. Is that a problem?Usually not on its own — each entry is small and expires in 60 seconds, and their presence just means your server is the side closing connections. It matters when it signals a design issue (no keep-alive, one connection per request) or when the same pattern appears on the *client* side, where each TIME_WAIT ties up an ephemeral source port and can exhaust the range.
- Why is lowering net.ipv4.tcp_fin_timeout not a fix for this failure?Because it controls a different state. `tcp_fin_timeout` bounds how long a socket waits in FIN_WAIT_2 for the peer's FIN after a half-close. TIME_WAIT's duration on Linux is a fixed 60-second constant with no sysctl at all, so changing `tcp_fin_timeout` leaves the sockets that block your bind exactly where they were.
saying these in an interview costs you the question
- The old process must still be running and holding the port
- TIME_WAIT always happens on the server side
- Lowering tcp_fin_timeout shortens TIME_WAIT
- Enable tcp_tw_recycle to clear it up
- SO_REUSEADDR lets two servers listen on one port