skip to content

questions

5

An HTTP client library exposes several separate timeout settings. What is the difference between a connect timeout, a read (socket) timeout, and an idle keep-alive timeout, and which failure does each one catch?

level: middleimportance: must knowfreq 62%

answer

  1. connect = handshake; read = gap between bytes; idle = pooled connection lifetime
  2. Read timeout is NOT a total duration cap
  3. Add a total/response timeout for a real deadline
  4. Client idle keep-alive < server keep-alive
  5. Also: write timeout, pool acquire timeout; defaults often infinite

basics

~20 s

Connect timeout bounds establishing the TCP (and often TLS) connection. Read timeout bounds waiting for the next bytes of a response - it is a per-read gap, not the total duration. Idle keep-alive timeout bounds how long an unused pooled connection stays open before being closed.

solid answer

~60 s

They guard different phases: - **Connect timeout** - from the first SYN until the TCP handshake completes. Catches an unreachable host, a dropped SYN (firewall blackhole), or a full backlog. Should be short: hundreds of milliseconds inside a datacentre. Some libraries time the TLS handshake separately. - **Read / socket timeout** - the maximum time to wait for **more data** on an established connection. Crucially it is a **gap between reads**, not a cap on the whole response: a server that dribbles a byte every second can hold the connection forever without ever tripping a 5-second read timeout. That is why many clients also offer a **total request timeout** covering headers plus body. - **Idle keep-alive timeout** - how long a pooled, unused connection may stay open before the client closes it. It manages resources and, when set below the server's own idle timeout, avoids handing out connections the server is about to close. Two more matter: **pool acquire timeout** (waiting for a free connection) and **write timeout** (a stalled request body). Set all of them explicitly; several libraries default to infinite.

go deeper

for a junior

Name the three phases - establishing the connection, waiting for the response, and how long an unused connection is kept - and say that each needs its own limit.

for a middle

Explain that the read timeout is a per-read gap rather than a total, add write and pool-acquire timeouts, and give sensible starting values.

for a senior

Justify values from measured latency, keep the client idle timeout under the server's, and separate error classes in telemetry so each timeout is independently diagnosable.

for a principal

Frame timeouts as a deadline budget across the call chain, tied to pool sizing and retry policy, with per-endpoint rather than global values.

## Why one timeout is never enough "The call took too long" is not one failure. A request passes through distinct phases, each with a different failure mode and a different sensible bound, so mature clients expose several timeouts. Getting them wrong produces either hung threads (too generous) or spurious failures under normal latency variance (too tight). ## Connect timeout Covers name resolution in some libraries and always the TCP handshake: SYN, SYN-ACK, ACK. Symptoms it catches: - host down or unreachable - usually an immediate `ECONNREFUSED` rather than a timeout; - a firewall silently dropping SYNs - here the OS retries with exponential backoff for a very long time (over two minutes on Linux defaults), so without an explicit connect timeout the client hangs; - a server whose accept backlog is full. Within a datacentre a connection should complete in single-digit milliseconds, so 100-500 ms is a reasonable bound; over the public internet, 1-3 s. Because it applies before any application work, it is safe to keep tight - and a connect timeout is the strongest evidence a request never reached the application at all, which matters for retry decisions. ## TLS handshake timeout Some libraries fold this into connect, some expose it separately. It bounds the handshake round trips, which can stall on a slow server or an unresponsive OCSP/CRL fetch. ## Read / socket timeout The most misunderstood setting. It is the maximum time the socket may be idle while the client is waiting for data - typically implemented as `SO_TIMEOUT` or an equivalent per-read deadline. It resets whenever a byte arrives. Consequences: - A 5-second read timeout does **not** guarantee a request completes in 5 seconds. A slow trickle keeps resetting it, and a large legitimate download also keeps resetting it, which is exactly why streaming works. - It cannot distinguish "server thinking" from "network path broken" - both look like silence. Because of that, clients increasingly expose a **response / total timeout** covering the entire exchange. Use the read timeout to detect stalls and the total timeout to enforce a deadline; for streaming or long-polling endpoints, keep the read timeout and raise or disable the total one. ## Write timeout Bounds how long sending the request body may stall, which happens when the server stops reading (a slow consumer or a full receive window). Without it, uploading to a wedged server hangs. ## Idle keep-alive timeout A property of the pool, not of a request: how long a connection with no traffic may sit idle before the client closes it. Servers have their own version (nginx `keepalive_timeout`, many managed load balancers 60 s), and they will close first if theirs is shorter. Keeping the client's idle timeout comfortably **below** the server's is the standard way to reduce the stale-connection race, where the client hands out a connection the server has already begun closing. Set it too low and you throw away warm connections and pay handshakes; too high and you accumulate idle sockets and hit that race. ## Pool acquire timeout When every pooled connection is leased, a new request waits for one. That wait is bounded by its own timeout and must be reported as its own error class - confusing it with a connect timeout sends people to debug the network when the problem is client-side concurrency. ## Putting it together A sane baseline for an internal JSON API: connect 200-500 ms, read 1-3 s, total request timeout aligned with the caller's deadline, write timeout similar to read, pool acquire a few hundred milliseconds, client idle keep-alive comfortably under the server's. Every value should be justified by measured latency (p99.9 plus margin), not copied from a tutorial - and every one of them should be set explicitly, because several widely used clients default to waiting forever.

  • Why can a request exceed a 5-second read timeout by minutes without ever triggering it?
    Because the read timeout measures the gap between successive reads, not total elapsed time. Each byte received resets the clock, so a server that emits data slowly - or a very large legitimate response - can run for minutes without a single gap exceeding 5 seconds. Enforcing an actual deadline requires a separate total or response timeout covering the whole exchange.
  • How should timeouts differ for a streaming or long-polling endpoint?
    Keep a read timeout to detect a genuinely dead connection, but size it to the maximum expected quiet period - for long polling that means longer than the poll interval - and do not apply a short total timeout, since the whole point is a long-lived response. Application-level heartbeats or keep-alive chunks let you keep the read timeout tight while still allowing an indefinite stream.

Connect timeout is how long you wait for someone to pick up the phone; read timeout is how long you tolerate silence once they are on the line; a total timeout is how long the whole call may last; idle keep-alive is how long you keep the line open when nobody is talking.

saying these in an interview costs you the question

  • Treating the read timeout as a cap on total request duration
  • Setting a single timeout value and assuming it covers connect, read and pool wait
  • Leaving library defaults in place without checking whether they are infinite
  • Setting the client's idle keep-alive timeout longer than the server's
  • Reporting a pool acquire timeout as a network or connect failure

context

open as a page

An HTTP call to a payment service times out after the client has finished sending the request body. Is it safe to retry, and what determines the answer?

level: seniorimportance: must knowfreq 50%

basics

~20 s

A timeout means unknown, not failed - the server may have processed the request and lost the response. Retry only if the failure proves the request never ran (connect failure, refused connection, reset with no bytes received) or the operation is genuinely idempotent server-side.

open as a page

A service intermittently fails with 'connection reset by peer' or an empty response on roughly one request in a thousand, always just after a quiet period, while the target service is healthy. What race is happening and how do you address it?

level: seniorimportance: must knowfreq 45%

basics

~20 s

The server's keep-alive idle timeout closed a pooled connection at the same moment the client leased it and wrote a request. The client never saw the FIN, so it sends into a dying socket. Keep the client's idle timeout below the server's, validate connections before use, and retry safe requests once.

open as a page

What does HTTP status 408 Request Timeout mean, who sends it and when, and how does it differ from 504 Gateway Timeout?

level: juniorimportance: should knowfreq 30%

basics

~20 s

408 is the origin server saying it gave up waiting for the client to send a complete request in time. 504 is a gateway or proxy saying it gave up waiting for an upstream server's response. 408 blames the inbound direction, 504 the outbound one.

open as a page

How would you choose timeout values for a service that sits in the middle of a call chain - an inbound caller, your service, and three downstream HTTP dependencies?

level: principalimportance: should knowfreq 40%

basics

~20 s

Start from the caller's deadline, not from downstream latency. Subtract your own processing, then give each downstream a timeout that fits the remaining budget including retries. Every inner timeout must be shorter than the outer one, and connect timeouts stay small.

open as a page