skip to content

Why does a long-lived TCP connection through a stateful firewall hang and fail on the first request after an idle period, and should you fix it with TCP keepalive or an application heartbeat?

level: seniorimportance: should knowfreq 30%

answer

  1. the box forgets, the ends do not
  2. silent drop, then retransmit backoff
  3. two-hour default is too long
  4. probes answered by the kernel, per hop
  5. heartbeat proves the application

basics

~20 s

The firewall's idle timer expired and it deleted the connection's state without telling either endpoint, so later segments are dropped or reset. Send traffic more often than that timer: TCP keepalive with a short idle time, or an application heartbeat.

solid answer

~50 s

Stateful middleboxes keep per-connection state with their own **idle timer**, often far shorter than TCP's two-hour keepalive default. When it expires the state is dropped silently, so both endpoints still think the connection is ESTABLISHED. The next request is then dropped or answered with an RST; if dropped, the sender **retransmits with exponential backoff** until its retry limit, which RFC 9293 says SHOULD correspond to at least 100 seconds, hence the hang. The fix is to send something more often than the middlebox timer. **TCP keepalive** needs no protocol change, but the idle time must be shortened per connection, its probes are answered by the peer's TCP stack (or by any TCP proxy in between) and prove nothing about the application. An **application heartbeat** crosses proxies end to end and proves the peer is responsive, at the cost of protocol support. Either way, the client must still handle a broken connection by reconnecting.

go deeper

for a junior

Remember that firewalls and NATs forget idle connections without telling the endpoints, and that periodic traffic keeps their state alive.

for a middle

Explain why the failure is a long hang: the dropped segment is retransmitted with backoff until the retry limit, while the idle peer keeps a half-open connection.

for a senior

Choose between keepalive and heartbeat by what each proves and where it terminates, set intervals below the shortest idle timer on the path, and insist on reconnect handling.

for a principal

Set a fleet-wide liveness policy: heartbeat intervals derived from the shortest middlebox timeout, jitter to avoid synchronised bursts, and client libraries that reconnect by default.

## Why the connection breaks while nobody is looking TCP itself has **no idle timeout**: an ESTABLISHED connection with no traffic can in principle last forever, and neither endpoint sends anything to prove it is still there. Stateful middleboxes in the path, such as firewalls and NAT devices, do not share that view. They keep an entry per connection and **expire entries that stay idle**, with timers chosen by the operator or vendor, often far shorter than the two hours TCP keepalive defaults to. RFC 8085 §3.5, while giving UDP guidance, notes that many middleboxes need keep-alive traffic for TCP connections too, just at a lower frequency than for UDP. When the entry expires, **nobody is told**. Both endpoints still hold the connection in ESTABLISHED. The next segment either end sends meets a middlebox with no matching state, and depending on the device it is: - **silently dropped**, or - answered with an **RST**, which aborts the connection immediately and surfaces as a reset error. ## Why it hangs instead of failing fast In the silent-drop case the client's request segment is lost, so its TCP does what it does for any loss: it **retransmits** with exponentially backed-off timeouts. RFC 9293 §3.8.3 says a connection is closed when retransmissions of one segment reach a threshold R2, and that R2 SHOULD correspond to **at least 100 seconds**; the application may set it per connection. So the request hangs for minutes, then fails with a timeout. The other side is worse off: if the server never sends, it never learns anything, and it keeps a **half-open** connection (RFC 9293 §3.5.1) and its resources indefinitely. ## Fix 1: TCP keepalive Keepalive (RFC 1122 §4.2.3.6, RFC 9293 §3.8.4) probes a connection that has been idle for a configured interval. A probe and its ACK are real segments, so they **refresh the middlebox's idle timer**, and a dead path is detected by unanswered probes. - It is **off by default** and must be enabled per connection (`SO_KEEPALIVE`). - The RFC requires the idle interval to default to **at least two hours**, so it must be **shortened below the middlebox timeout** or it never fires in time. The probe spacing and count are implementation parameters. - It needs **no change to the application protocol**. - The probe is answered by the **peer's TCP stack**, so a hung application still looks alive. - It is **per TCP connection**: a proxy that terminates TCP answers the probe itself, so it proves only the first hop. ## Fix 2: an application heartbeat A heartbeat is a message the application protocol defines: a ping frame, a no-op request, a periodic status message. - It travels **end to end** through TCP-terminating proxies, refreshing state on every leg. - A reply proves the **peer application** is running and responsive, and the sender can enforce its own deadline. - It needs **protocol support** on both sides and costs a few bytes per interval. ## Choosing | Question | TCP keepalive | Application heartbeat | |---|---|---| | Needs protocol change? | no | yes | | Default on? | no; idle default of 2 h or more | whatever the protocol defines | | Crosses a TCP-terminating proxy? | no, per hop | yes, end to end | | Detects a hung application? | no | yes | | Refreshes middlebox state? | yes, if frequent enough | yes, if frequent enough | A practical order: 1. If the protocol has, or can add, a heartbeat, use it with an interval comfortably below the shortest idle timeout on the path. 2. If it cannot, enable TCP keepalive with a short idle time on the long-lived connections. 3. In both cases, handle failure: detect a broken connection, **reconnect** and retry safe operations. RFC 8085 makes the same point for UDP: keep-alives are not a substitute for recovering from broken sessions. The interval is a trade-off: frequent enough to beat the shortest middlebox timer, infrequent enough not to waste bandwidth, battery or server work across a large fleet.

  • Why does TCP keepalive with its default settings usually fail to keep a connection alive through a stateful firewall?
    The RFC requires the keepalive idle interval to default to no less than two hours, and keepalive itself is off unless the application enables it. Middlebox idle timers are often far shorter, so the state is gone long before the first probe. The idle time has to be set per connection below the shortest timer on the path.
  • If a TCP-terminating proxy sits between client and server, what does a successful TCP keepalive on the client's connection prove?
    Only that the proxy's TCP stack still has the client's connection. The proxy answers the probe itself; the proxy-to-server leg is a separate TCP connection with its own state and timers. End-to-end liveness needs an application heartbeat that the proxy forwards.
  • Why did the request hang for minutes instead of failing at once after the middlebox dropped the state?
    The middlebox silently discarded the segment, so the sender treated it as loss and retransmitted with exponential backoff. RFC 9293 says the retransmission limit R2 SHOULD correspond to at least 100 seconds before the connection is closed, so the failure surfaces only after that.

saying these in an interview costs you the question

  • The firewall sends a FIN to both endpoints when its idle timer expires.
  • TCP connections have a built-in idle timeout that closed the connection.
  • Enabling SO_KEEPALIVE with default settings is enough to beat firewall timeouts.
  • A successful TCP keepalive proves the remote application is healthy.
  • With keepalive or heartbeats in place, the client never needs reconnect logic.