skip to content

What does TCP keepalive do on an idle connection, and is it enabled by default?

level: juniorimportance: should knowfreq 32%

answer

  1. an optional feature, not core TCP
  2. off unless the application asks
  3. two-hour idle floor
  4. probe below the window elicits an ACK
  5. idle, interval, count

basics

~20 s

TCP keepalive probes a connection idle for a set interval: an ACK means the peer still has it; an RST or repeated silence means it is gone. It is optional, off by default, and the idle default is at least two hours.

solid answer

~50 s

TCP keepalive is an **optional** mechanism (RFC 1122 §4.2.3.6, RFC 9293 §3.8.4) that probes a connection that has been **idle**, with no data or ACKs received and nothing outstanding, for a configured interval. The probe is a segment with sequence number `SND.NXT-1`, usually with no data, which the peer answers with an ACK. An RST means the peer no longer has the connection; no answer after several probes means the peer or path is gone, and the connection is aborted. The RFC requires it to be **off by default**, switchable per connection (`SO_KEEPALIVE` is the portable switch), and the idle interval to default to **no less than two hours**. The gap between probes and the probe count are implementation parameters. One unanswered probe MUST NOT be treated as death, because bare ACKs are not delivered reliably.

go deeper

for a junior

Know three facts: keepalive probes an idle TCP connection, it is off unless the application enables it, and the default idle time is at least two hours.

for a middle

Explain the probe (SND.NXT-1, answered by ACK or RST), the RFC's MUSTs, and which of idle, interval and count the RFC fixes versus leaves to the implementation.

for a senior

Reason about detection time as idle plus interval times count, and about what a probe does not prove: it tests the peer's TCP stack, not the application.

for a principal

Decide per protocol whether liveness belongs in TCP keepalive or in the application, weighing detection time, proxies in the path and the traffic a large fleet of probes generates.

## Why TCP needs a probe at all A TCP connection has no built-in heartbeat. If neither side sends anything, no segments flow, and a peer that has crashed, rebooted or lost its network path is **indistinguishable from a peer that is simply quiet**. RFC 9293 §3.5.1 calls a connection **half-open** when one side has closed or lost it without the other knowing. Half-open connections are discovered only when someone sends: the surviving side then gets an RST (if the peer rebooted and has no record) or no answer at all (if the peer or path is gone). An idle server holding a connection for a crashed client may never find out, and keeps the resources forever. Keepalive exists for exactly that case. RFC 1122 says it should only be invoked in server applications that might otherwise hang indefinitely and consume resources if a client crashes or aborts during a network failure. ## What the RFCs require Keepalive is deliberately **not** part of the core protocol. RFC 1122 §4.2.3.6 lists the objections: it can break perfectly good connections during transient failures, it consumes bandwidth on connections nobody is using, and it costs money on paths that charge per packet. So the requirements are hedged: - implementers **MAY** include keepalives (RFC 9293 MAY-5); - if they do, the application **MUST** be able to turn them on or off **per connection** (MUST-24), and they **MUST default to off** (MUST-25); - probes **MUST** only be sent when no sent data is outstanding and nothing has been received for an interval (MUST-26); - that interval **MUST** be configurable (MUST-27) and **MUST default to no less than two hours** (MUST-28); - a missing response to **any single probe MUST NOT** be treated as a dead connection (MUST-29), because ACK segments carrying no data are not reliably transmitted. ## What a probe looks like on the wire A keepalive probe is a segment with `SEG.SEQ = SND.NXT - 1`: one byte **before** the next sequence number. The RFC says it SHOULD carry no data, and MAY carry one garbage octet for old implementations that ignore empty probes. On a quiet connection this sequence number lies just outside the receiver's window, so the peer's TCP responds the way it responds to any unacceptable segment: it sends an **ACK** of its current state. The possible outcomes: | Peer state | Response | Result | |---|---|---| | Alive, connection intact | ACK | idle timer restarts | | Rebooted, no record of the connection | RST | connection aborted, application sees a reset error | | Crashed, powered off or path broken | nothing | more probes; after the configured count, connection aborted with a timeout error | ## The three knobs: idle, interval, count Implementations expose three parameters, but only the first comes from the RFC: 1. **Idle time**: how long the connection must be idle before the first probe. The RFC mandates a configurable value defaulting to at least two hours; many implementations ship with exactly that floor. 2. **Probe interval**: the gap between unanswered probes. Implementation-defined. 3. **Probe count**: how many unanswered probes before giving up. Implementation-defined, and more than one because of MUST-29. Detection time is therefore roughly idle + interval x count. With the two-hour default, a dead peer is noticed only after more than two hours, which is why applications that rely on keepalive shorten the idle time per connection. ## What keepalive is not - It is **not HTTP keep-alive**. HTTP's keep-alive is about reusing a connection for more requests; TCP keepalive is a liveness probe inside one connection. They share a name and nothing else. - It does **not** run while data is outstanding. Detecting a dead peer during a transfer is the retransmission mechanism's job, which aborts after its own retry limit. - It checks the **peer's TCP stack**, not the peer application: a hung process whose operating system is healthy still answers probes. - It is **off** unless the application enables it with `SO_KEEPALIVE`.

  • Why must a TCP keepalive implementation not declare a connection dead after one unanswered probe?
    Segments that carry no data, including the probe's ACK response, are not reliably transmitted: TCP does not retransmit a lost bare ACK. One lost packet would otherwise kill a healthy connection. RFC 9293 (MUST-29) therefore requires several unanswered probes before the connection is considered dead.
  • Why does a TCP keepalive probe use sequence number SND.NXT-1?
    On an idle connection the peer expects `SND.NXT` next, so `SND.NXT-1` is just outside its receive window. TCP answers any unacceptable segment with an ACK stating its current state, so the probe elicits a response without delivering new data to the peer's application. A peer that has lost the connection answers with RST instead.

saying these in an interview costs you the question

  • TCP keepalive is enabled on every connection by default.
  • The TCP specification requires a keepalive probe every 75 seconds.
  • TCP keepalive is the same feature as HTTP keep-alive.
  • One unanswered keepalive probe means the TCP peer is dead.
  • A successful TCP keepalive proves the peer application is healthy.