skip to content

In `ss -tln` output on Linux, the Recv-Q and Send-Q columns mean something different for a row in LISTEN state than for one in ESTAB. Explain both readings, and what a persistently non-zero Recv-Q on a listening socket tells you about the service.

level: seniorimportance: should knowfreq 40%

answer

  1. units change with socket state
  2. bytes here, connections there
  3. the kernel already finished its half
  4. application is not calling accept
  5. climbing overflow counter confirms it

basics

~20 s

On an ESTAB row the columns are bytes: unread received data and unacknowledged sent data. On a LISTEN row they are connection counts: Recv-Q is completed connections waiting to be accepted, Send-Q the accept-queue limit.

solid answer

~50 s

On an established socket the two columns are byte counts — Recv-Q is data that arrived and the application has not read, Send-Q is data written but not yet acknowledged by the peer. On a **listening** socket they change meaning entirely: `ss` reports the current accept-queue depth in Recv-Q and that queue's configured maximum in Send-Q, so a row like `LISTEN 129 128` means 129 finished connections are queued and the limit is 128 — the queue is at or over its ceiling. A persistently non-zero Recv-Q there means the kernel completed the handshakes but the application is not calling `accept()` fast enough: blocked worker threads, an exhausted pool, a long stop-the-world pause, or a thread stuck on a slow downstream call. That points the investigation at the process, not the network. I'd corroborate it with the kernel's listen-overflow counter via `nstat`.

code

bash · 10 lines
bash
# Current depth vs configured ceiling on the listening socket
ss -tln 'sport = :8080'
# State  Recv-Q Send-Q Local Address:Port
# LISTEN 129    128          0.0.0.0:8080

# Sample it rather than trusting one reading
watch -n1 "ss -tln 'sport = :8080'"

# Corroborate with the kernel's listen-queue counters
nstat -az TcpExtListenOverflows TcpExtListenDrops

go deeper

for a junior

Know that Recv-Q and Send-Q exist and that on a normal established connection they count bytes not yet read and not yet acknowledged. Recognising that the meaning changes for listening sockets already puts you ahead.

for a middle

Explain both readings precisely and identify which number is the current depth and which is the ceiling. Be able to say that a queued connection is already established from the client's point of view.

for a senior

Use it to redirect an investigation: a non-zero listening Recv-Q means the application is not accepting, so you go looking at threads, pools and pauses, and corroborate with the kernel's listen-overflow counter rather than a single sample.

for a principal

Own the tradeoff between queue depth and fail-fast behaviour across services: how much queueing you want before shedding load, where admission control belongs, and how the accept queue interacts with upstream retry and timeout budgets.

## Two columns, two meanings `ss` prints Recv-Q and Send-Q for every socket, and the units change with the socket's state. This is the single most misread pair of numbers in Linux network triage. **On an established socket (`ESTAB`)** they are byte counts: - **Recv-Q** — bytes that have arrived in the receive buffer and the application has not yet read. A large, growing value means the reader is slower than the sender. - **Send-Q** — bytes the application has written that have not yet been acknowledged by the peer. A large, growing value means the data is not getting out: a stalled peer, a congested or lossy path, or a receiver advertising a zero window. **On a listening socket (`LISTEN`)** they are connection counts: - **Recv-Q** — how many fully established connections are sitting in the accept queue right now, waiting for the application to take them. - **Send-Q** — the maximum depth of that queue for this socket, fixed when the service called `listen()` (and capped by a kernel-wide limit). ```bash ss -tln # State Recv-Q Send-Q Local Address:Port # LISTEN 129 128 0.0.0.0:8080 # LISTEN 0 511 0.0.0.0:80 ``` The first row is the interesting one. The second is a healthy listener: nothing queued, a generous ceiling. ## What a non-zero listening Recv-Q actually means It means the kernel has already done its half of the work. The three-way handshake completed, the connection is real and the client believes it is connected — it is simply waiting in a queue because the server process has not called `accept()` to pick it up. Whatever is wrong is therefore **above the socket layer, inside the application**: - Every worker thread is blocked on something slow — a database call, a lock, a downstream HTTP request with no timeout. - The thread or connection pool is sized too small for the arrival rate. - The process is paused: a long garbage-collection pause, heavy swapping, or CPU starvation from a noisy neighbour. - A single-threaded accept loop is doing work per connection that belongs on a worker. A brief non-zero value under a burst is normal and healthy — the queue exists precisely to absorb bursts. What matters is whether it is *persistent* or repeatedly climbing, which you see by sampling: ```bash watch -n1 "ss -tln 'sport = :8080'" ``` The client-visible symptom is distinctive and worth naming: connections that succeed instantly but then sit silent for seconds before the first byte of response, while the application's own logs show fast request handling. The application never sees the waiting time, because its clock starts at `accept()`. ## When the queue is full Once the accept queue is at its limit, new completed connections cannot be added and the kernel drops them; depending on configuration the client sees either a stall and a retry or an outright failure. Because the drop happens below the application, nothing appears in the service log — which is exactly why a candidate who only reads application logs concludes "the app is fine" while users are timing out. The kernel does count these events, and that counter is the corroboration you want: ```bash nstat -az TcpExtListenOverflows TcpExtListenDrops ``` A `TcpExtListenOverflows` value that climbs while you watch confirms that the accept queue overflowed rather than that you caught one unlucky sample. `nstat -az` prints absolute values; running `nstat` without `-a` shows the delta since the last run, which is often more useful. ## Reading the established rows in the same breath While you are there, the `ESTAB` rows tell a complementary story. Many connections with a large Recv-Q against your own service means your process is not reading fast enough. A large Send-Q on connections to a downstream service means your data is not being acknowledged — that is a path or peer problem, not a local one. `ss -tim` adds the TCP internals and socket memory to these rows if you need to go deeper. ## The fixes, and their order The reflex fix is to raise the backlog. That is worth doing when the limit is plainly too small for the arrival rate, but it should be second: a bigger queue only buys more time before the same wall, and it *increases* the latency clients suffer while queued. The primary fix is to make the application accept and serve faster — unblock the workers, add capacity, put timeouts on downstream calls so threads return, or scale out. Raise the queue when you need headroom for legitimate bursts, not as a way to hide a service that cannot keep up. The interview-grade summary: a non-zero Recv-Q on a `LISTEN` row redirects the whole investigation from the network to the process, which is a large saving on a live incident.

  • Clients report multi-second delays but the application's own request timings look fast. How does that fit what `ss` shows?
    It fits exactly: the application's timer starts when it calls `accept()`, so time spent queued before that is invisible to it. A non-zero Recv-Q on the `LISTEN` row accounts for the missing seconds. That gap between client-observed and server-observed latency is the tell, and it is why you check the listening socket rather than trusting the service's own numbers.
  • You confirm the accept queue is overflowing. Why is raising the backlog not automatically the right fix?
    A deeper queue admits more waiting connections, so clients wait longer before being served rather than failing fast — you convert refusals into latency. It is the right move when the ceiling is genuinely too small for normal bursts, but if the service cannot keep up with sustained arrivals the queue only delays the same wall. Fix accept throughput first, size the queue second.
  • How would you distinguish a one-off burst from a sustained accept-queue problem?
    Sample rather than take a single reading: watch the `LISTEN` row for a minute and see whether Recv-Q drains to zero between bursts. Then check whether the kernel's listen-overflow counter is increasing over the same window with `nstat`. A drained queue and a static counter mean a healthy absorbed burst; a queue that never empties and a climbing counter mean sustained overload.

saying these in an interview costs you the question

  • Reading listening Recv-Q as bytes of unread data
  • Assuming a full queue appears in the application's log
  • Raising the backlog as the first and only fix
  • Blaming the network when the handshake already completed
  • Judging it from one sample instead of watching it

context