skip to content

In TCP, what does Nagle's algorithm do with small application writes, and what changes when an application sets TCP_NODELAY on the socket?

level: middleimportance: must knowfreq 48%

answer

  1. one small segment in flight
  2. tinygrams: 1 data byte, 40 header bytes
  3. wait for the ACK or a full MSS
  4. per-connection off switch the RFC requires

basics

~20 s

Nagle's algorithm makes a TCP sender hold new small data while earlier data is still unacknowledged, sending it when the ACK arrives or a full-sized segment accumulates. TCP_NODELAY disables it, so small writes go out immediately.

solid answer

~40 s

Nagle's algorithm (RFC 896, now RFC 9293 §3.7.4) fights the **small-packet problem**: one byte of data carried in a 41-byte IPv4 packet. If a TCP sender has unacknowledged data outstanding (`SND.NXT > SND.UNA`), it buffers further small user data, regardless of the PSH bit, until that data is acknowledged or it has a full-sized segment (`Eff.snd.MSS` bytes) to send. When nothing is outstanding, a small write goes out at once. So at most one small segment is in flight per round trip, and small writes are coalesced. RFC 9293 says a TCP SHOULD implement it and MUST let an application disable it per connection. `TCP_NODELAY` is the portable socket option that does that: small segments are sent immediately, trading more packets and header overhead for lower latency.

code

pseudocode · 16 lines
pseudocode
on application_write(data):
    queue.append(data)
    try_send()

try_send():
    while queue not empty:
        if len(queue) >= EFF_SND_MSS:
            send_segment(take(queue, EFF_SND_MSS))   # full segment: never held by Nagle
        elif SND_NXT == SND_UNA or nodelay:
            send_segment(take_all(queue))            # nothing outstanding, or Nagle disabled
        else:
            return                                   # small data waits for an ACK

on ack_received():
    advance SND_UNA
    try_send()

go deeper

for a junior

Remember the one-liner: while earlier data is unacknowledged, a TCP sender holds new small writes; TCP_NODELAY switches that off for one connection.

for a middle

Explain the exact condition from RFC 9293: unacknowledged data outstanding, buffer until ACK or a full MSS, regardless of PSH, and why that means at most one small segment per round trip.

for a senior

Show judgement about when TCP_NODELAY helps: it suits whole-message request/response traffic, but an application that dribbles tiny writes needs its writes batched, not just Nagle disabled.

for a principal

Frame the trade-off as latency versus packet and header efficiency, and argue for message-oriented write paths in shared libraries rather than per-service socket-option folklore.

## The problem Nagle solved John Nagle described the **small-packet problem** in RFC 896 (1984). A user typing into a remote terminal produced one TCP segment per keystroke: **1 byte of data plus 40 bytes of IPv4 and TCP header**, a 41-byte packet and a 4000% overhead. On a lightly loaded network that is merely wasteful; on a congested one those tiny packets (often called *tinygrams*) add to the congestion, cause loss and retransmission, and can push throughput so low that connections are aborted. Earlier systems fixed this with a timer: hold small data for a fixed 200-500 ms hoping more arrives. RFC 896 points out the flaw: no single time limit suits both a fast LAN and a congested path with a multi-second round trip. Nagle's answer is **adaptive and timer-free**: let the acknowledgements pace the sender. ## The rule, as the specification states it RFC 9293 §3.7.4 (the current TCP specification, which obsoletes RFC 793) restates the algorithm from RFC 896 and RFC 1122 §4.2.3.4: - If there is **unacknowledged data** on the connection (`SND.NXT > SND.UNA`), the sender **buffers all new user data**, regardless of the **PSH** bit, - until the **outstanding data is acknowledged**, - or until it can send a **full-sized segment** (`Eff.snd.MSS` bytes). - If nothing is outstanding, a small segment is sent immediately. The consequences follow directly: 1. On a quiet connection the first small write is **not delayed at all**. 2. While that segment is unacknowledged, later small writes pile up in the send buffer. 3. When the ACK returns, everything buffered goes out together as one larger segment. 4. Bulk data is unaffected: full-sized segments are never held by Nagle (they remain subject to the receiver's window and the congestion window). The effect is **at most one small segment in flight per round-trip time**. On a LAN with a sub-millisecond RTT the coalescing is barely noticeable as long as acknowledgements come back promptly; on a long path it saves many packets. RFC 9293 keeps the requirement levels from RFC 1122: a TCP implementation **SHOULD** implement Nagle (SHLD-7), and there **MUST** be a way for an application to **disable it on an individual connection** (MUST-17). RFC 1122 gives the reason for the escape hatch: some applications, such as real-time display updates, need small segments streamed out at the maximum rate. ## TCP_NODELAY `TCP_NODELAY` is the portable socket option through which applications exercise that MUST-17 switch. With it set: | Aspect | Nagle on (default) | `TCP_NODELAY` set | |---|---|---| | Small write, nothing outstanding | sent immediately | sent immediately | | Small write, data outstanding | held until ACK or full segment | sent immediately | | Packets for many tiny writes | coalesced | one segment per write (roughly) | | Latency of a small write | held until the ACK returns (longer if the peer delays it) | none added by TCP's coalescing | | Header overhead | low | potentially high | Setting `TCP_NODELAY` does **not** disable flow control, congestion control or delayed acknowledgement on the peer; it only stops the local sender from holding small data while earlier data is unacknowledged. ## When to disable it, and when not to Disabling Nagle is right for latency-sensitive, small-message traffic: interactive protocols, request/response RPC where each message is written once and awaited, real-time control or game state updates. The standard advice has a precondition, though: **the application should hand TCP whole messages**. An application that writes a message in many tiny pieces and sets `TCP_NODELAY` recreates the tinygram problem Nagle was built to prevent. The cleaner fix is usually to assemble a message in a user-space buffer (or use a gather write) and issue one send. Nagle interacts badly with the receiver's **delayed acknowledgements**: if the receiver holds its ACK waiting for data to piggyback on, the sender's held data waits too. RFC 9293 Appendix A.3 notes that common operating systems enable both by default, that request/response applications can suffer, that a modification to Nagle exists and is implemented in some systems (without being added to the standard), and that many applications simply disable Nagle. ## Common misreadings - Nagle is **not a timer**; RFC 896 explicitly rejects timers. - The **PSH** bit does not override it; RFC 9293 says the data is buffered regardless of PSH. - It does **not** delay the first small write on an idle connection. - It is a **sender** algorithm; the receiver-side coalescing is delayed ACK, a separate mechanism.

  • Does setting the PSH flag on a TCP send force Nagle to release buffered small data?
    No. RFC 9293 states that under Nagle the sender buffers user data regardless of the PSH bit, until outstanding data is acknowledged or a full-sized segment can be sent. PSH is a hint to deliver data to the receiving application promptly, not a record marker and not an override for the sender's coalescing.
  • Why is setting TCP_NODELAY not a universal fix for latency in a TCP protocol?
    It removes Nagle's hold, but if the application writes each message as many tiny pieces, every piece becomes its own segment with 40 or more bytes of header, recreating the small-packet problem. The better design gives TCP whole messages in one send; then `TCP_NODELAY` removes the remaining wait without multiplying packets.
  • Why did RFC 896 reject a fixed hold timer for coalescing small TCP writes?
    No single limit fits every path: a timer short enough for a fast LAN is too short to prevent congestion on a path with a multi-second round trip, and one long enough for that path frustrates LAN users. Pacing by acknowledgements adapts automatically to each connection's round-trip time.

A lift that leaves at once when the building is quiet, but while its previous trip is still unconfirmed keeps boarding passengers until the confirmation arrives or the car is full.

saying these in an interview costs you the question

  • Nagle's algorithm waits a fixed 200 ms before sending any small segment.
  • Setting the PSH flag makes Nagle send the buffered data immediately.
  • Nagle delays even the first small write on an idle connection.
  • TCP_NODELAY turns off delayed ACK on the peer as well.
  • Nagle's algorithm is a receiver-side mechanism for batching acknowledgements.