skip to content

TCP

Handshake, byte sequencing with sliding windows, and congestion control turn lossy IP packets into an ordered stream. Interviewers lean on it because slow or stalled connections trace back here.

on this pageshow

questions

page 1 of 2

In TCP, what problem does congestion control solve, and why is the receiver's advertised window alone not enough to prevent it?

level: juniorimportance: must knowfreq 58%

answer

  1. the path, not the peer
  2. router queues overflow and drop
  3. congestion collapse
  4. a second, sender-side window
  5. the smaller of two windows

basics

~20 s

TCP congestion control stops senders from overloading the network path between two hosts. The receiver's advertised window only protects the receiver's buffer, so the sender keeps its own congestion window and sends no more than the smaller of the two.

solid answer

~40 s

The receiver advertises `rwnd` to say how much buffer *it* has free, but it knows nothing about the routers in between. A fast receiver behind a slow bottleneck link would let a sender push data that queues and drops at that router, and the retransmissions add even more load: the spiral RFC 896 named **congestion collapse**. So the sender keeps a second, private limit, the congestion window `cwnd`, which it grows while data is acknowledged and cuts when it sees a congestion signal: a lost segment or, with ECN (RFC 3168), a Congestion Experienced mark. RFC 5681 states the rule: never send data beyond the highest acknowledged sequence number plus `min(cwnd, rwnd)`. RFC 9293 makes slow start, congestion avoidance and exponential RTO backoff mandatory for every TCP.

go deeper

for a junior

Recall that the receiver's window protects the receiver and the congestion window protects the network, and that the sender uses the smaller of the two.

for a middle

Explain congestion collapse as a feedback spiral of drops and retransmissions, and name the signals a sender uses to infer congestion: loss and ECN marks.

for a senior

Be ready to say which window is binding in a real transfer, for example a fast receiver behind a slow bottleneck versus a fast path to a slow-reading application.

for a principal

Frame congestion control as a shared-resource contract among independent senders: why endpoint inference won over router-supplied rates, and what that costs in probing and latency.

## Two different things can be overwhelmed A TCP sender can overrun two separate resources, and TCP protects them with two separate mechanisms: - **The receiver.** Its socket buffer has finite space. The receiver tells the sender how much it can still accept by advertising a **receive window** (`rwnd`) in every segment. That is **flow control**, and it is the receiver's own statement about itself. - **The network path.** Between the two hosts sit links and routers with their own capacity and their own queues. The slowest link on the path, the **bottleneck**, decides how fast data can really move. The receiver cannot see that link, so `rwnd` says nothing about it. A receiver on a fast local network can advertise a large window while the path to it crosses a much slower link. If the sender trusted `rwnd` alone, it would send at whatever rate the receiver's buffer allowed, far faster than the bottleneck can forward. ## What congestion collapse is RFC 896 described what happens when senders ignore the path, calling it **congestion collapse**: 1. Senders inject more data than the bottleneck can forward, so the router's queue fills. 2. Once the queue is full, the router drops arriving packets. 3. The senders see missing acknowledgments and retransmit, which adds more packets to an already overloaded link. 4. The link stays busy carrying duplicates and data that will be dropped further along, so the useful throughput falls even though the link is fully loaded. The network is busy, but little of the work is useful. Congestion control exists to keep the total load near what the path can carry, so that this spiral never starts. ## The congestion window RFC 5681, the current Standards Track specification of TCP congestion control, gives the sender a state variable, the **congestion window** (`cwnd`). Its definition is the key rule: > a TCP MUST NOT send data with a sequence number higher than the sum of the highest acknowledged sequence number and the minimum of `cwnd` and `rwnd`. | | `rwnd` (receive window) | `cwnd` (congestion window) | |---|---|---| | Protects | the receiver's buffer | the network path | | Who decides it | the receiver | the sender | | Carried on the wire? | yes, in the TCP header's window field | no, it is local sender state | | Changes when | the receiving application reads or falls behind | ACKs arrive, or congestion is detected | | Owning mechanism | flow control | congestion control | In RFC 5681 both windows are measured in **bytes**. Whichever is smaller is the binding limit at that moment: early in a connection `cwnd` usually is, while a slow-reading application makes `rwnd` the limit instead. ## How the sender learns about congestion Classic TCP gets no rate from the network. It infers congestion at the endpoints from two signals: - **Loss.** A segment that goes missing is taken as evidence that a queue overflowed. Loss shows up as duplicate acknowledgments or as a retransmission timeout. - **ECN marks.** With Explicit Congestion Notification (RFC 3168), a router doing active queue management can set the Congestion Experienced mark on a packet instead of dropping it. The receiver echoes it back with the TCP `ECE` flag, and the sender answers with `CWR` once it has reduced its window. RFC 3168 requires the reaction to be essentially the same as the reaction to a single dropped packet. RFC 9293 says a TCP SHOULD implement ECN. Neither signal says how much capacity exists. The sender has to **probe**: grow `cwnd` while things go well, and cut it when a signal arrives. ## The algorithms that manage cwnd RFC 9293 (MUST-19) requires every TCP to implement **slow start**, **congestion avoidance** and **exponential backoff of the RTO**, precisely to avoid creating congestion collapse. RFC 5681 specifies four intertwined algorithms: 1. **Slow start** grows `cwnd` quickly from a small initial window. 2. **Congestion avoidance** grows it by about one segment per round trip once it nears the last known safe size. 3. **Fast retransmit** resends a segment after three duplicate ACKs. 4. **Fast recovery** cuts the window without restarting from one segment. RFC 9293 also allows alternative algorithms, such as CUBIC (RFC 9438), provided they conform to the IETF's congestion-control guidance. ## Common confusions - Flow control and congestion control are different: one protects a host, the other the network. - `cwnd` is never sent in a header; only `rwnd` is. - Congestion control is not optional in a conforming TCP, even though reliability alone could be achieved by retransmission.

  • Does a router tell a TCP sender how much bandwidth is available?
    No. Classic TCP infers congestion only at the endpoints: a lost segment, or, with ECN (RFC 3168), a Congestion Experienced mark that the receiver echoes back with the `ECE` flag. Neither signal carries a rate, so the sender still probes by growing `cwnd` and cuts it when a signal arrives.
  • If `rwnd` is 64 KB and `cwnd` is 20 KB, how much unacknowledged data may the sender have outstanding?
    At most 20 KB. RFC 5681 limits the sender to the highest acknowledged sequence number plus `min(cwnd, rwnd)`. Here the path is the constraint; if the receiving application stopped reading and `rwnd` fell below 20 KB, the receiver would become the constraint instead.

The receiver's window is the size of the car park at the destination; the congestion window is the driver's own judgement of how many cars the road can take right now. A huge empty car park does not widen a jammed road, so each driver meters cars onto the road by watching for jams.

saying these in an interview costs you the question

  • Flow control and congestion control are the same mechanism under two names.
  • The receiver's advertised window tells the sender how fast the network path is.
  • Routers send TCP an explicit rate saying how fast it may transmit.
  • Congestion control is optional because retransmission already makes TCP reliable.
  • The congestion window is advertised by the receiver in the TCP header.
open as a page

In TCP, what problem does flow control solve, and how is it different from congestion control?

level: juniorimportance: must knowfreq 62%

basics

~20 s

TCP flow control keeps a fast sender from overrunning the receiver's buffer: the receiver advertises how many more bytes it can accept (rwnd). Congestion control protects the network through the sender's cwnd; the sender obeys the smaller.

open as a page

In TCP, what does each segment of the three-way handshake (SYN, SYN-ACK, ACK) carry, and why are three segments needed?

level: juniorimportance: must knowfreq 82%

basics

~20 s

The SYN carries the client's initial sequence number, the SYN-ACK carries the server's and acknowledges the client's, and the final ACK acknowledges the server's. Three segments let each side confirm the other's starting number and reject stale duplicate SYNs.

open as a page

How does a TCP sender detect that a segment was lost, and what are the two ways it decides to retransmit it?

level: juniorimportance: must knowfreq 55%

basics

~20 s

A TCP sender infers loss two ways: its retransmission timer (RTO) expires with data still unacknowledged, or a third duplicate ACK arrives and triggers fast retransmit. Fast retransmit usually fires first; the timer is the slower backstop.

open as a page

In a TCP header, what do the SYN, ACK, FIN, RST, PSH and URG control flags each signal?

level: juniorimportance: must knowfreq 58%

basics

~20 s

SYN synchronizes sequence numbers to open a connection, ACK says the acknowledgment field is valid, FIN means the sender has no more data, RST aborts the connection, PSH asks for prompt delivery, and URG marks the urgent pointer as significant.

open as a page

How does TCP use sequence numbers and acknowledgment numbers to turn unreliable IP delivery into a reliable, ordered byte stream?

level: juniorimportance: must knowfreq 62%

basics

~20 s

TCP numbers every byte it sends. The receiver's acknowledgment number names the next byte it expects, confirming all earlier bytes at once; sequence numbers let the receiver reorder and drop duplicates, and the sender resends whatever stays unacknowledged.

open as a page

In TCP, what is the difference between a port and a socket, and how can one listening port serve thousands of clients at once?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A TCP port is a 16-bit number on a host; a socket is an IP address plus a port. A connection is a pair of sockets, so one server port holds many connections differing by client address or port.

open as a page

In TCP, what is the difference between closing a connection gracefully with FIN and aborting it with RST?

level: juniorimportance: must knowfreq 68%

basics

~20 s

FIN is an orderly close of one direction: data queued before it is still delivered, the peer acknowledges it, and the other direction stays open. RST aborts the whole connection at once, discarding queued data and state.

open as a page

In TCP, how do slow start and congestion avoidance each grow the congestion window, and how does ssthresh decide which one runs?

level: middleimportance: must knowfreq 55%

basics

~20 s

In slow start a TCP sender adds up to one SMSS per ACK of new data, roughly doubling cwnd each round trip; in congestion avoidance it adds about one SMSS per round trip. Slow start runs while cwnd is below ssthresh.

open as a page

How does TCP's sliding window work: what moves its left and right edges, and how much new data may the sender transmit?

level: middleimportance: must knowfreq 50%

basics

~20 s

TCP's sliding window is the byte range a sender may have in flight. Acknowledgments move the left edge (SND.UNA); the receiver's advertised window, which grows as its application reads, sets the right edge. Usable = SND.UNA + SND.WND - SND.NXT.

open as a page

In TCP, what does Nagle's algorithm do with small application writes, and what changes when an application sets TCP_NODELAY on the socket?

level: middleimportance: must knowfreq 48%

basics

~20 s

Nagle's algorithm makes a TCP sender hold new small data while earlier data is still unacknowledged, sending it when the ACK arrives or a full-sized segment accumulates. TCP_NODELAY disables it, so small writes go out immediately.

open as a page

In TCP, what does the MSS option announce, how is its value derived, and what is assumed when a SYN omits it?

level: middleimportance: must knowfreq 45%

basics

~20 s

The MSS option announces the largest TCP payload the sender of the SYN can receive, normally its MTU minus fixed IP and TCP headers (1460 on a 1500-byte IPv4 link). Without it, the peer assumes 536 for IPv4 or 1220 for IPv6.

open as a page

A TCP client with ISN 1000 connects to a server with ISN 8000, sends 200 bytes, then a FIN: which sequence and acknowledgment numbers appear?

level: middleimportance: must knowfreq 46%

basics

~20 s

The client's SYN uses 1000, so the server acknowledges 1001; data occupies 1001-1200 and is acknowledged with 1201; the FIN uses 1201 and is acknowledged with 1202. The server's SYN uses 8000, so the client acknowledges 8001.

open as a page

In the TCP socket API, what do socket, bind, listen, accept and connect each do, and where does the three-way handshake actually happen?

level: middleimportance: must knowfreq 60%

basics

~20 s

socket creates an endpoint, bind names its local address and port, listen makes it a passive open in LISTEN, and connect sends the SYN. The stack completes the handshake and queues it; accept hands it over as a new socket.

open as a page

In a TCP close, which side ends up in TIME-WAIT, what does that state protect against, and why does it last 2×MSL?

level: middleimportance: must knowfreq 55%

basics

~20 s

The side that sends the first FIN (the active closer) enters TIME-WAIT. It waits 2×MSL so it can re-acknowledge a retransmitted final FIN and so stray segments of this connection expire before the same four-tuple is reused.

open as a page

Why does a TCP request/response client that writes each request as a small header, then a small body, then reads the reply see every call take about 40 ms longer than the round trip, and how do you fix it?

level: seniorimportance: must knowfreq 38%

basics

~20 s

Nagle's algorithm holds the second small write until the first is acknowledged, while the server delays that ACK because it has nothing to send yet. Each call waits out the delayed-ACK timer; send each request in one write or set TCP_NODELAY.

open as a page

What does TCP keepalive do on an idle connection, and is it enabled by default?

level: juniorimportance: should knowfreq 32%

basics

~20 s

TCP keepalive probes a connection idle for a set interval: an ACK means the peer still has it; an RST or repeated silence means it is gone. It is optional, off by default, and the idle default is at least two hours.

open as a page

During a TCP bulk transfer, how does the sender's congestion window respond to three duplicate ACKs compared with a retransmission timeout, and why do they differ?

level: middleimportance: should knowfreq 42%

basics

~20 s

On the third duplicate ACK a TCP sender halves ssthresh from FlightSize, retransmits and continues in fast recovery at about half its old window. On a retransmission timeout it also halves ssthresh but drops cwnd to one segment and slow-starts again.

open as a page

A TCP receiver's application stops reading its socket; what happens as the advertised window reaches zero, and why must the sender keep probing?

level: middleimportance: should knowfreq 38%

basics

~20 s

When a TCP receiver's buffer fills, it advertises a zero window and the sender stops sending new data. The sender then sends periodic zero-window probes, because the ACK that reopens the window carries no data and is not retransmitted if lost.

open as a page

Why does TCP not start every connection's sequence numbers at zero, and how does RFC 9293 say the initial sequence number should be chosen?

level: middleimportance: should knowfreq 35%

basics

~20 s

A fixed start lets delayed segments from an earlier connection on the same ports be accepted; a predictable one lets off-path attackers forge segments. RFC 9293 sets ISN = M + F(four-tuple, secret key): a 4-microsecond clock plus a keyed pseudorandom offset.

open as a page

When a TCP client sends a SYN, how do open, closed and silently filtered ports respond, and how does each attempt end?

level: middleimportance: should knowfreq 50%

basics

~20 s

An open port answers SYN-ACK. A closed port on a live host sends a RST acknowledging the SYN, failing the attempt within one round trip. A silent filter answers nothing, so the client retries its SYN with backoff until it gives up.

open as a page

Which TCP states do the client and the server pass through while a connection opens, and when can each side first send data?

level: middleimportance: should knowfreq 48%

basics

~20 s

The client goes CLOSED, SYN-SENT, ESTABLISHED; the server goes LISTEN, SYN-RECEIVED, ESTABLISHED. The client is established one round trip after its SYN and can send then; the server is established about half a round trip later.

open as a page

Why can't TCP use a fixed retransmission timeout, and how does RFC 6298 compute the RTO from measured round-trip times?

level: middleimportance: should knowfreq 32%

basics

~20 s

Round-trip times differ hugely between paths and drift over time, so a fixed timeout is too slow or spuriously early. RFC 6298 sets RTO = SRTT + max(G, 4 x RTTVAR), with gains 1/8 and 1/4, rounded up to 1 s.

open as a page

A captured TCP SYN has a data offset of 10; how long is its header, and what do its option bytes typically carry?

level: middleimportance: should knowfreq 32%

basics

~10 s

Data offset counts 32-bit words, so 10 means a 40-byte header: 20 fixed bytes plus 20 bytes of options, typically MSS (4), SACK-permitted (2), timestamps (10), one NOP and window scale (3).

open as a page

Which ACKs and SACK blocks does a TCP receiver that has acknowledged byte 1000 send when the first of the next three 1000-byte segments is lost?

level: middleimportance: should knowfreq 36%

basics

~10 s

Each later segment triggers an immediate duplicate ACK of 1001, carrying SACK blocks 2001-3001 and then 2001-4001. The receiver buffers bytes 2001-4000 undelivered; when 1001-2000 arrives, it acknowledges 4001 at once.

open as a page

What are the three IANA port-number ranges for TCP and UDP, and where does a client's ephemeral source port come from?

level: middleimportance: should knowfreq 40%

basics

~20 s

RFC 6335 splits the 16-bit space into System ports 0-1023, User ports 1024-49151 and Dynamic ports 49152-65535. A client's ephemeral port is chosen by its own stack at connect time from a locally configured range, ideally unpredictably (RFC 6056).

open as a page

In the socket API, what does a non-blocking TCP connect or accept return when it cannot finish at once, and how does the program learn the outcome?

level: middleimportance: should knowfreq 28%

basics

~10 s

A non-blocking connect sends the SYN and returns EINPROGRESS; the socket turns writable once the handshake resolves, and SO_ERROR says how. A non-blocking accept with nothing queued returns EAGAIN or EWOULDBLOCK.

open as a page

What is a TCP half-close, and what can each endpoint still do after one side has sent its FIN?

level: middleimportance: should knowfreq 32%

basics

~10 s

A TCP half-close ends one direction only. The side that sent FIN promises no more data but keeps receiving; the peer, in CLOSE-WAIT, may keep sending until it closes too.

open as a page

Why does loss-based TCP congestion control such as Reno or CUBIC underuse a lossy wireless link yet cause bufferbloat on a deep-buffered link, and how does BBR differ?

level: seniorimportance: should knowfreq 26%

basics

~20 s

Loss-based TCP treats every drop as congestion: random wireless losses make it cut its window needlessly, and on deep buffers it keeps growing until the queue overflows. BBR instead models bottleneck bandwidth and minimum RTT and paces to that model.

open as a page

A TCP bulk transfer over a 1 Gbit/s path with a 100 ms round trip never exceeds about 5 Mbit/s, with no packet loss; how do the receive window and window scaling explain this?

level: seniorimportance: should knowfreq 30%

basics

~20 s

A TCP connection moves at most one receive window per round trip. An unscaled 65,535-byte window over 100 ms caps it near 5.2 Mbit/s; filling this 12.5 MB bandwidth-delay product needs window scaling (shift 8 or more) and a matching buffer.

open as a page

showing 1–30 of 46