skip to content

questions

6

A client opening TCP connections to a Linux host gets an instant "connection refused" on one port, while a connect to another port on the same host hangs for around two minutes before failing. What is the kernel doing differently in each case?

level: juniorimportance: must knowfreq 70%

answer

  1. an answer versus silence
  2. who sends the reset
  3. retransmit with exponential backoff
  4. DROP hangs, REJECT fails fast
  5. the listener's bound address matters

basics

~20 s

Connection refused means the SYN reached the host and the kernel answered with a TCP RST because no socket was listening on that port. A hang means no answer came back at all, usually a dropped packet, so the client keeps retransmitting the SYN until it gives up.

solid answer

~50 s

Both are failures of the TCP handshake, but one gets an answer and the other gets silence. If a SYN arrives for a port where nothing is listening and no firewall rule intercepts it, the kernel immediately replies with a RST; the client's `connect()` fails right away with `ECONNREFUSED`. That is a definitive negative answer from a reachable host. If the SYN is silently discarded — a firewall rule that DROPs, a black-holed route, a host that is simply gone — nothing comes back, so the client's kernel retransmits the SYN with exponential backoff. On Linux that is bounded by `net.ipv4.tcp_syn_retries`, which defaults to 6 and works out to roughly 127 seconds before `connect()` fails with `ETIMEDOUT`. So refused tells you the host is up and reachable and the port is closed; a hang tells you your packets are disappearing and you do not yet know where.

go deeper

for a junior

Know the one-line difference: refused means something answered and said no, a hang means nothing answered at all. Say plainly that a refusal proves the host is reachable.

for a middle

Explain the mechanics — the kernel emits a RST when no listening socket matches, and an unanswered SYN is retransmitted with exponential backoff until the client's retry budget runs out.

for a senior

Turn the distinction into a triage path: refusals send you to the process and its bind address, timeouts send you to firewalls, routing and host liveness. Mention that a bind to loopback refuses remote clients while local ones succeed.

for a principal

Own the policy angle: rejecting fails clients fast and keeps debugging cheap, dropping denies scanners any confirmation the host exists. Decide where each belongs in your network, and make client connect timeouts explicit rather than inheriting the kernel's two-minute ceiling.

## What connect() is really waiting for A TCP client that calls `connect()` sends one packet — a SYN — and then waits. Everything the caller eventually sees is a consequence of what comes back, or of nothing coming back. The two failure modes candidates conflate, "connection refused" and "connection timed out", are exactly the answered case and the unanswered case. ## The refused case: an RST from the kernel When a SYN arrives on a Linux host, the kernel looks for a listening socket matching the destination address and port. This lookup is done by the kernel, not by any application: there is no process to consult, because by definition nothing has that port open. If no listening socket matches, the kernel replies with a TCP segment carrying the RST flag. The client's kernel receives the RST, aborts the pending connection attempt, and `connect()` returns `-1` with `errno` set to `ECONNREFUSED`, which the C library renders as "Connection refused". This happens in one round trip — on a LAN, sub-millisecond. The speed itself is diagnostic: an instant refusal proves the network path works in both directions and that the remote host's IP stack is alive and answering. A subtle and very common variant is a listener bound to the wrong address. A process that binds `127.0.0.1:8080` has a socket for the loopback address only. A SYN arriving on the machine's LAN address for port 8080 matches no socket, so the kernel refuses it exactly as if nothing were running. The service is up; it is just not listening where the client is knocking. Binding `0.0.0.0` (or `::`) is what makes a listener match every local address. ## The hanging case: silence and retransmission If the SYN is discarded rather than answered, the client learns nothing. TCP treats an unanswered SYN as a lost packet and retransmits, doubling the interval each time: roughly 1s, 2s, 4s, 8s, 16s, 32s, then a final wait. `net.ipv4.tcp_syn_retries` sets how many retransmissions are attempted; the default of 6 gives about 127 seconds in total before the kernel gives up and `connect()` fails with `ETIMEDOUT`. ``` sysctl net.ipv4.tcp_syn_retries net.ipv4.tcp_syn_retries = 6 ``` The usual causes of silence are a packet filter configured to DROP rather than reject, a security group or cloud firewall that discards non-permitted traffic by policy, a route that leads nowhere, or a host that is powered off. Note that a *drop of the reply* produces the same symptom: the server may have answered with a SYN-ACK that never made it home, for example because of asymmetric routing or a stateful middlebox. From the client's side these are indistinguishable, which is why the hang is a weaker signal than the refusal. ## Reject versus drop is a deliberate choice Firewalls can behave either way, and that is a policy decision. A rule that rejects sends something back — on Linux, netfilter's REJECT target defaults to an ICMP port-unreachable message, and can be told to send a TCP RST instead — so the client fails fast. A rule that drops sends nothing, so the client hangs for the full retry budget. Rejecting is friendlier to legitimate clients and to your own debugging; dropping is preferred on internet-facing edges because it gives a scanner no confirmation that the host exists at all. When you see one of your own services timing out where you expected a refusal, an intermediate DROP rule is the first hypothesis. ## Why this matters in practice The distinction turns "it does not connect" into a directed investigation: - **Refused** — the host and its IP stack are reachable. Stop looking at the network. Look at whether the process is running, which address and port it bound, and whether you are talking to the right host or the right port. - **Timed out** — packets are being lost in one direction or the other. Look at firewalls, security groups, routing and whether the host is up at all. The application is very likely blameless. A third outcome worth recognising is a connection that *is* established but then produces no data. That is not a handshake failure at all: the three-way handshake completed, so the port is open and something accepted you. That points at the application or at queueing behind the listener, not at reachability. ## The timeout is a client-side budget One last point candidates miss: the ~127-second hang is the *kernel's* default budget, not a law of TCP. Applications routinely impose a much shorter connect timeout of their own by using a non-blocking socket and giving up early, which is why the same unreachable host may fail in 2 seconds from one client and 127 seconds from another. The kernel's ceiling is the worst case, not the observed case.

  • The process is running and listening, and local connections to it work, but remote clients get connection refused. What do you check first?
    The address it bound to. A socket bound to `127.0.0.1` matches only loopback traffic, so a SYN arriving on the host's LAN address finds no matching socket and the kernel refuses it. Confirm the bind address and change it to `0.0.0.0` or the specific external address. Container and VM port mappings produce the same symptom for the same reason.
  • How does the failure differ if a firewall rejects with an ICMP port-unreachable message instead of dropping the packet?
    The client fails promptly instead of hanging, because it receives something rather than nothing. Linux's netfilter REJECT target defaults to ICMP port-unreachable, which the client's kernel treats as a fatal error for the pending connect. Configuring the reject to send a TCP reset instead makes the failure indistinguishable from having no listener at all.
  • If the handshake succeeds but the client then waits a long time for the first byte, what has that ruled out?
    Everything about reachability. A completed three-way handshake proves packets flow both ways and that a socket accepted the connection. The delay is now above the handshake — the application not reading or responding, or the connection sitting in the kernel's accept queue because the server is not calling accept fast enough.

saying these in an interview costs you the question

  • Connection refused means the host is down
  • A hang always means the server is slow
  • Refused and timed out are the same failure, worded differently
  • The application decides to refuse the connection
  • Retrying a refused connection immediately will make it work

context

open as a page

You stop a busy TCP server on Linux and start it again immediately, and it refuses to start with "Address already in use" even though no process is holding the port. Why does bind() fail, and what makes the restart succeed?

level: middleimportance: must knowfreq 72%

basics

~20 s

Connections the old server closed are still in TIME_WAIT, so the kernel still has sockets on that local port and refuses a plain bind. Setting SO_REUSEADDR on the listening socket before bind lets the server take the port back despite them.

open as a page

On a Linux host where a web server and an application server run side by side, what do you gain and what do you give up by connecting them over a Unix domain socket instead of a TCP socket on 127.0.0.1?

level: middleimportance: should knowfreq 45%

basics

~20 s

A Unix domain socket skips the TCP/IP stack entirely: no handshake, no ephemeral ports, no TIME_WAIT, lower overhead, and access is controlled by filesystem permissions on the socket file. The cost is that it only works on the same host and leaves a stale file behind after an unclean exit.

open as a page

A Linux service that opens many short-lived outbound TCP connections to a single backend address starts failing with "cannot assign requested address". What resource has run out, what governs its size, and how would you fix it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

It has run out of ephemeral source ports for that destination. Linux picks them from net.ipv4.ip_local_port_range, and each closed connection holds its port in TIME_WAIT for 60 seconds, so a high connect rate exhausts the range. Reusing connections through a keep-alive pool is the durable fix.

open as a page

During a burst of new connections, clients of a Linux TCP server wait several seconds before their first request gets any response, while the server's own application logs show nothing wrong. Explain the two kernel queues behind listen() and what happens when the accept queue fills.

level: seniorimportance: should knowfreq 50%

basics

~20 s

A listening socket has two kernel queues: a SYN queue for half-open handshakes and an accept queue of completed connections waiting for the application to call accept(). When the accept queue overflows, Linux drops the client's final handshake packet by default, so the server retransmits its SYN-ACK and the client stalls for seconds.

open as a page

Linux's SO_REUSEPORT lets several processes hold listening sockets on the same TCP port. How would you decide whether to scale a service that way rather than having one process accept connections and hand them to workers?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

SO_REUSEPORT gives every worker its own accept queue and lets the kernel spread new connections across them by hashing the connection's four-tuple, removing the single-acceptor bottleneck. The tradeoffs are hashed rather than balanced load, and connections lost from a worker's queue when it exits.

open as a page