skip to content

Nagle, Delayed ACK & Keepalive

Nagle holds small writes until earlier data is acknowledged, delayed ACK holds the acknowledgement, and keepalive probes an idle peer. Interviewers ask about the stall the first two cause together.

on this pageshow

questions

5

In TCP, what does Nagle's algorithm do with small application writes, and what changes when an application sets TCP_NODELAY on the socket?

level: middleimportance: must knowfreq 48%

answer

  1. one small segment in flight
  2. tinygrams: 1 data byte, 40 header bytes
  3. wait for the ACK or a full MSS
  4. per-connection off switch the RFC requires

basics

~20 s

Nagle's algorithm makes a TCP sender hold new small data while earlier data is still unacknowledged, sending it when the ACK arrives or a full-sized segment accumulates. TCP_NODELAY disables it, so small writes go out immediately.

solid answer

~40 s

Nagle's algorithm (RFC 896, now RFC 9293 §3.7.4) fights the **small-packet problem**: one byte of data carried in a 41-byte IPv4 packet. If a TCP sender has unacknowledged data outstanding (`SND.NXT > SND.UNA`), it buffers further small user data, regardless of the PSH bit, until that data is acknowledged or it has a full-sized segment (`Eff.snd.MSS` bytes) to send. When nothing is outstanding, a small write goes out at once. So at most one small segment is in flight per round trip, and small writes are coalesced. RFC 9293 says a TCP SHOULD implement it and MUST let an application disable it per connection. `TCP_NODELAY` is the portable socket option that does that: small segments are sent immediately, trading more packets and header overhead for lower latency.

code

pseudocode · 16 lines
pseudocode
on application_write(data):
    queue.append(data)
    try_send()

try_send():
    while queue not empty:
        if len(queue) >= EFF_SND_MSS:
            send_segment(take(queue, EFF_SND_MSS))   # full segment: never held by Nagle
        elif SND_NXT == SND_UNA or nodelay:
            send_segment(take_all(queue))            # nothing outstanding, or Nagle disabled
        else:
            return                                   # small data waits for an ACK

on ack_received():
    advance SND_UNA
    try_send()

go deeper

for a junior

Remember the one-liner: while earlier data is unacknowledged, a TCP sender holds new small writes; TCP_NODELAY switches that off for one connection.

for a middle

Explain the exact condition from RFC 9293: unacknowledged data outstanding, buffer until ACK or a full MSS, regardless of PSH, and why that means at most one small segment per round trip.

for a senior

Show judgement about when TCP_NODELAY helps: it suits whole-message request/response traffic, but an application that dribbles tiny writes needs its writes batched, not just Nagle disabled.

for a principal

Frame the trade-off as latency versus packet and header efficiency, and argue for message-oriented write paths in shared libraries rather than per-service socket-option folklore.

## The problem Nagle solved John Nagle described the **small-packet problem** in RFC 896 (1984). A user typing into a remote terminal produced one TCP segment per keystroke: **1 byte of data plus 40 bytes of IPv4 and TCP header**, a 41-byte packet and a 4000% overhead. On a lightly loaded network that is merely wasteful; on a congested one those tiny packets (often called *tinygrams*) add to the congestion, cause loss and retransmission, and can push throughput so low that connections are aborted. Earlier systems fixed this with a timer: hold small data for a fixed 200-500 ms hoping more arrives. RFC 896 points out the flaw: no single time limit suits both a fast LAN and a congested path with a multi-second round trip. Nagle's answer is **adaptive and timer-free**: let the acknowledgements pace the sender. ## The rule, as the specification states it RFC 9293 §3.7.4 (the current TCP specification, which obsoletes RFC 793) restates the algorithm from RFC 896 and RFC 1122 §4.2.3.4: - If there is **unacknowledged data** on the connection (`SND.NXT > SND.UNA`), the sender **buffers all new user data**, regardless of the **PSH** bit, - until the **outstanding data is acknowledged**, - or until it can send a **full-sized segment** (`Eff.snd.MSS` bytes). - If nothing is outstanding, a small segment is sent immediately. The consequences follow directly: 1. On a quiet connection the first small write is **not delayed at all**. 2. While that segment is unacknowledged, later small writes pile up in the send buffer. 3. When the ACK returns, everything buffered goes out together as one larger segment. 4. Bulk data is unaffected: full-sized segments are never held by Nagle (they remain subject to the receiver's window and the congestion window). The effect is **at most one small segment in flight per round-trip time**. On a LAN with a sub-millisecond RTT the coalescing is barely noticeable as long as acknowledgements come back promptly; on a long path it saves many packets. RFC 9293 keeps the requirement levels from RFC 1122: a TCP implementation **SHOULD** implement Nagle (SHLD-7), and there **MUST** be a way for an application to **disable it on an individual connection** (MUST-17). RFC 1122 gives the reason for the escape hatch: some applications, such as real-time display updates, need small segments streamed out at the maximum rate. ## TCP_NODELAY `TCP_NODELAY` is the portable socket option through which applications exercise that MUST-17 switch. With it set: | Aspect | Nagle on (default) | `TCP_NODELAY` set | |---|---|---| | Small write, nothing outstanding | sent immediately | sent immediately | | Small write, data outstanding | held until ACK or full segment | sent immediately | | Packets for many tiny writes | coalesced | one segment per write (roughly) | | Latency of a small write | held until the ACK returns (longer if the peer delays it) | none added by TCP's coalescing | | Header overhead | low | potentially high | Setting `TCP_NODELAY` does **not** disable flow control, congestion control or delayed acknowledgement on the peer; it only stops the local sender from holding small data while earlier data is unacknowledged. ## When to disable it, and when not to Disabling Nagle is right for latency-sensitive, small-message traffic: interactive protocols, request/response RPC where each message is written once and awaited, real-time control or game state updates. The standard advice has a precondition, though: **the application should hand TCP whole messages**. An application that writes a message in many tiny pieces and sets `TCP_NODELAY` recreates the tinygram problem Nagle was built to prevent. The cleaner fix is usually to assemble a message in a user-space buffer (or use a gather write) and issue one send. Nagle interacts badly with the receiver's **delayed acknowledgements**: if the receiver holds its ACK waiting for data to piggyback on, the sender's held data waits too. RFC 9293 Appendix A.3 notes that common operating systems enable both by default, that request/response applications can suffer, that a modification to Nagle exists and is implemented in some systems (without being added to the standard), and that many applications simply disable Nagle. ## Common misreadings - Nagle is **not a timer**; RFC 896 explicitly rejects timers. - The **PSH** bit does not override it; RFC 9293 says the data is buffered regardless of PSH. - It does **not** delay the first small write on an idle connection. - It is a **sender** algorithm; the receiver-side coalescing is delayed ACK, a separate mechanism.

  • Does setting the PSH flag on a TCP send force Nagle to release buffered small data?
    No. RFC 9293 states that under Nagle the sender buffers user data regardless of the PSH bit, until outstanding data is acknowledged or a full-sized segment can be sent. PSH is a hint to deliver data to the receiving application promptly, not a record marker and not an override for the sender's coalescing.
  • Why is setting TCP_NODELAY not a universal fix for latency in a TCP protocol?
    It removes Nagle's hold, but if the application writes each message as many tiny pieces, every piece becomes its own segment with 40 or more bytes of header, recreating the small-packet problem. The better design gives TCP whole messages in one send; then `TCP_NODELAY` removes the remaining wait without multiplying packets.
  • Why did RFC 896 reject a fixed hold timer for coalescing small TCP writes?
    No single limit fits every path: a timer short enough for a fast LAN is too short to prevent congestion on a path with a multi-second round trip, and one long enough for that path frustrates LAN users. Pacing by acknowledgements adapts automatically to each connection's round-trip time.

A lift that leaves at once when the building is quiet, but while its previous trip is still unconfirmed keeps boarding passengers until the confirmation arrives or the car is full.

saying these in an interview costs you the question

  • Nagle's algorithm waits a fixed 200 ms before sending any small segment.
  • Setting the PSH flag makes Nagle send the buffered data immediately.
  • Nagle delays even the first small write on an idle connection.
  • TCP_NODELAY turns off delayed ACK on the peer as well.
  • Nagle's algorithm is a receiver-side mechanism for batching acknowledgements.
open as a page

Why does a TCP request/response client that writes each request as a small header, then a small body, then reads the reply see every call take about 40 ms longer than the round trip, and how do you fix it?

level: seniorimportance: must knowfreq 38%

basics

~20 s

Nagle's algorithm holds the second small write until the first is acknowledged, while the server delays that ACK because it has nothing to send yet. Each call waits out the delayed-ACK timer; send each request in one write or set TCP_NODELAY.

open as a page

What does TCP keepalive do on an idle connection, and is it enabled by default?

level: juniorimportance: should knowfreq 32%

basics

~20 s

TCP keepalive probes a connection idle for a set interval: an ACK means the peer still has it; an RST or repeated silence means it is gone. It is optional, off by default, and the idle default is at least two hours.

open as a page

Why does a long-lived TCP connection through a stateful firewall hang and fail on the first request after an idle period, and should you fix it with TCP keepalive or an application heartbeat?

level: seniorimportance: should knowfreq 30%

basics

~20 s

The firewall's idle timer expired and it deleted the connection's state without telling either endpoint, so later segments are dropped or reset. Send traffic more often than that timer: TCP keepalive with a short idle time, or an application heartbeat.

open as a page

What is silly window syndrome in TCP, and how do the sender and the receiver each avoid it?

level: middleimportance: nice to knowfreq 14%

basics

~20 s

Silly window syndrome is a stable TCP pattern of tiny window openings filled by tiny segments. RFC 9293 requires both sides to avoid it: receivers withhold small window increases, senders wait for a worthwhile amount to send.

open as a page