A client opening TCP connections to a Linux host gets an instant "connection refused" on one port, while a connect to another port on the same host hangs for around two minutes before failing. What is the kernel doing differently in each case?
answer
- an answer versus silence
- who sends the reset
- retransmit with exponential backoff
- DROP hangs, REJECT fails fast
- the listener's bound address matters
basics
~20 sConnection refused means the SYN reached the host and the kernel answered with a TCP RST because no socket was listening on that port. A hang means no answer came back at all, usually a dropped packet, so the client keeps retransmitting the SYN until it gives up.
solid answer
~50 sBoth are failures of the TCP handshake, but one gets an answer and the other gets silence. If a SYN arrives for a port where nothing is listening and no firewall rule intercepts it, the kernel immediately replies with a RST; the client's `connect()` fails right away with `ECONNREFUSED`. That is a definitive negative answer from a reachable host. If the SYN is silently discarded — a firewall rule that DROPs, a black-holed route, a host that is simply gone — nothing comes back, so the client's kernel retransmits the SYN with exponential backoff. On Linux that is bounded by `net.ipv4.tcp_syn_retries`, which defaults to 6 and works out to roughly 127 seconds before `connect()` fails with `ETIMEDOUT`. So refused tells you the host is up and reachable and the port is closed; a hang tells you your packets are disappearing and you do not yet know where.
go deeper
Know the one-line difference: refused means something answered and said no, a hang means nothing answered at all. Say plainly that a refusal proves the host is reachable.
Explain the mechanics — the kernel emits a RST when no listening socket matches, and an unanswered SYN is retransmitted with exponential backoff until the client's retry budget runs out.
Turn the distinction into a triage path: refusals send you to the process and its bind address, timeouts send you to firewalls, routing and host liveness. Mention that a bind to loopback refuses remote clients while local ones succeed.
Own the policy angle: rejecting fails clients fast and keeps debugging cheap, dropping denies scanners any confirmation the host exists. Decide where each belongs in your network, and make client connect timeouts explicit rather than inheriting the kernel's two-minute ceiling.
## What connect() is really waiting for A TCP client that calls `connect()` sends one packet — a SYN — and then waits. Everything the caller eventually sees is a consequence of what comes back, or of nothing coming back. The two failure modes candidates conflate, "connection refused" and "connection timed out", are exactly the answered case and the unanswered case. ## The refused case: an RST from the kernel When a SYN arrives on a Linux host, the kernel looks for a listening socket matching the destination address and port. This lookup is done by the kernel, not by any application: there is no process to consult, because by definition nothing has that port open. If no listening socket matches, the kernel replies with a TCP segment carrying the RST flag. The client's kernel receives the RST, aborts the pending connection attempt, and `connect()` returns `-1` with `errno` set to `ECONNREFUSED`, which the C library renders as "Connection refused". This happens in one round trip — on a LAN, sub-millisecond. The speed itself is diagnostic: an instant refusal proves the network path works in both directions and that the remote host's IP stack is alive and answering. A subtle and very common variant is a listener bound to the wrong address. A process that binds `127.0.0.1:8080` has a socket for the loopback address only. A SYN arriving on the machine's LAN address for port 8080 matches no socket, so the kernel refuses it exactly as if nothing were running. The service is up; it is just not listening where the client is knocking. Binding `0.0.0.0` (or `::`) is what makes a listener match every local address. ## The hanging case: silence and retransmission If the SYN is discarded rather than answered, the client learns nothing. TCP treats an unanswered SYN as a lost packet and retransmits, doubling the interval each time: roughly 1s, 2s, 4s, 8s, 16s, 32s, then a final wait. `net.ipv4.tcp_syn_retries` sets how many retransmissions are attempted; the default of 6 gives about 127 seconds in total before the kernel gives up and `connect()` fails with `ETIMEDOUT`. ``` sysctl net.ipv4.tcp_syn_retries net.ipv4.tcp_syn_retries = 6 ``` The usual causes of silence are a packet filter configured to DROP rather than reject, a security group or cloud firewall that discards non-permitted traffic by policy, a route that leads nowhere, or a host that is powered off. Note that a *drop of the reply* produces the same symptom: the server may have answered with a SYN-ACK that never made it home, for example because of asymmetric routing or a stateful middlebox. From the client's side these are indistinguishable, which is why the hang is a weaker signal than the refusal. ## Reject versus drop is a deliberate choice Firewalls can behave either way, and that is a policy decision. A rule that rejects sends something back — on Linux, netfilter's REJECT target defaults to an ICMP port-unreachable message, and can be told to send a TCP RST instead — so the client fails fast. A rule that drops sends nothing, so the client hangs for the full retry budget. Rejecting is friendlier to legitimate clients and to your own debugging; dropping is preferred on internet-facing edges because it gives a scanner no confirmation that the host exists at all. When you see one of your own services timing out where you expected a refusal, an intermediate DROP rule is the first hypothesis. ## Why this matters in practice The distinction turns "it does not connect" into a directed investigation: - **Refused** — the host and its IP stack are reachable. Stop looking at the network. Look at whether the process is running, which address and port it bound, and whether you are talking to the right host or the right port. - **Timed out** — packets are being lost in one direction or the other. Look at firewalls, security groups, routing and whether the host is up at all. The application is very likely blameless. A third outcome worth recognising is a connection that *is* established but then produces no data. That is not a handshake failure at all: the three-way handshake completed, so the port is open and something accepted you. That points at the application or at queueing behind the listener, not at reachability. ## The timeout is a client-side budget One last point candidates miss: the ~127-second hang is the *kernel's* default budget, not a law of TCP. Applications routinely impose a much shorter connect timeout of their own by using a non-blocking socket and giving up early, which is why the same unreachable host may fail in 2 seconds from one client and 127 seconds from another. The kernel's ceiling is the worst case, not the observed case.
- The process is running and listening, and local connections to it work, but remote clients get connection refused. What do you check first?The address it bound to. A socket bound to `127.0.0.1` matches only loopback traffic, so a SYN arriving on the host's LAN address finds no matching socket and the kernel refuses it. Confirm the bind address and change it to `0.0.0.0` or the specific external address. Container and VM port mappings produce the same symptom for the same reason.
- How does the failure differ if a firewall rejects with an ICMP port-unreachable message instead of dropping the packet?The client fails promptly instead of hanging, because it receives something rather than nothing. Linux's netfilter REJECT target defaults to ICMP port-unreachable, which the client's kernel treats as a fatal error for the pending connect. Configuring the reject to send a TCP reset instead makes the failure indistinguishable from having no listener at all.
- If the handshake succeeds but the client then waits a long time for the first byte, what has that ruled out?Everything about reachability. A completed three-way handshake proves packets flow both ways and that a socket accepted the connection. The delay is now above the handshake — the application not reading or responding, or the connection sitting in the kernel's accept queue because the server is not calling accept fast enough.
saying these in an interview costs you the question
- Connection refused means the host is down
- A hang always means the server is slow
- Refused and timed out are the same failure, worded differently
- The application decides to refuse the connection
- Retrying a refused connection immediately will make it work