skip to content

A Go CLI hangs forever in a net.Conn Read against a dead peer, with the read deadline cleared after the handshake. How do you diagnose and fix it?

level: seniorimportance: should knowfreq 48%

answer

  1. SIGQUIT dumps every goroutine stack
  2. the state line says IO wait, with a duration
  3. a dead host sends neither FIN nor RST
  4. clearing the deadline made the read unbounded
  5. size the deadline from the heartbeat interval

basics

~20 s

Send SIGQUIT to dump goroutine stacks: the reader sits in IO wait inside a socket read for minutes. With no deadline, a peer that dies without sending FIN or RST leaves TCP silent forever. Fix it by setting a read deadline before every read.

solid answer

~50 s

Diagnose it with `kill -QUIT` on the process: the goroutine dump shows the reader in `IO wait` with a duration attached, and a `net` frame under your wire layer — proof it is parked in the socket, not spinning or deadlocked on a mutex. The cause is that clearing the deadline after the handshake left an unbounded read, and when a peer disappears without closing — power loss, a firewall dropping flow state — no FIN or RST ever arrives, so the local TCP stack has nothing to report and the read waits indefinitely. The fix is to set a read deadline before every read in the steady-state loop, sized from the protocol's heartbeat interval rather than a request latency budget, plus a write deadline. TCP keep-alives are a useful backstop but detection depends on OS settings and can take many minutes, so they do not replace the deadline. When a deadline fires mid-frame, drop the connection and reconnect rather than resuming.

code

text · 5 lines
text
goroutine 12 [IO wait, 51 minutes]:
internal/poll.(*FD).Read(...)
net.(*conn).Read(...)
main.(*wireClient).readFrame(0xc0000a4000)
	/src/wire/client.go:88 +0x9c

go deeper

for a junior

Know that a read with no deadline can block forever, and that sending SIGQUIT to a hung Go program prints every goroutine's stack so you can see where it is stuck.

for a middle

Explain the IO wait state and its printed duration in a goroutine dump, and why a half-open connection produces no error: with no FIN or RST arriving, there is nothing for the socket to report.

for a senior

Show the full remediation — a deadline before every read, sized from heartbeats, a write deadline too, keep-alives as backstop only — and state clearly that a mid-frame timeout means reconnect rather than resume.

for a principal

Set the standard: which layer owns liveness for long-lived connections, what the CLI's overall upper bound is, and whether the protocol gains a heartbeat — a wire-format change every consumer inherits.

## The symptom The CLI prints nothing, consumes no CPU, and never exits. It is not a livelock and not a busy loop — it is a goroutine parked on a socket that will never produce a byte. ## Getting proof, not a guess Go hands you the answer for free. Send `SIGQUIT` (`kill -QUIT <pid>`, or Ctrl-\ in a terminal): the runtime dumps every goroutine's stack and aborts. The interesting lines are the header of each goroutine: ``` goroutine 12 [IO wait, 51 minutes]: ``` Two facts come straight out of that. The state is `IO wait`, so the goroutine is blocked waiting for the network poller — not on a channel (`chan receive`), not on a lock (`sync.Mutex.Lock`), not running. And the runtime prints how long it has been in that state, so "51 minutes" removes any doubt that this is slowness. Under it you see `internal/poll` and `net` frames beneath your own wire-layer function, naming the exact read. If you cannot signal the process — it is inside a container with no shell, say — the same dump is available by other routes: `GOTRACEBACK=all` controls how much is printed, and a service can expose goroutine stacks over an HTTP debug endpoint. But for a CLI, SIGQUIT is the whole investigation. ## Why TCP does not tell you This is the part candidates most often get wrong. TCP is silent when idle. If the peer process crashes but its host stays up, the host's kernel sends an RST and you find out. If the **host** dies — power cut, kernel panic, a cloud instance terminated — or a stateful firewall or NAT quietly forgets the flow, nobody sends anything. Your side holds a connection that looks perfectly healthy and waits for data that will never come. This is a half-open connection, and the only way to discover it is to send something and notice you get nothing back. So an unbounded `Read` on a long-lived connection is not a small risk. It is a guaranteed hang the first time the peer disappears ungracefully, which for a connection held open for hours is a matter of when, not if. ## The bug in the code The shape is almost always the same, and it comes from good intentions: ``` c.SetReadDeadline(time.Now().Add(10 * time.Second)) if err := handshake(c); err != nil { return err } c.SetReadDeadline(time.Time{}) // "this stream is long-lived, no timeout" for { frame, err := readFrame(c) // unbounded ... } ``` Someone bounded the handshake correctly, then cleared the deadline because the steady-state stream is idle for long stretches and a ten-second deadline fired constantly. Clearing it swapped a noisy failure for a silent one. ## The fix **Set a read deadline before every read.** Deadlines are absolute instants that do not renew themselves, so the call belongs at the top of each loop iteration, covering one frame. **Size it from the protocol, not from a latency budget.** A stream that is idle by design needs a deadline derived from how often the peer is contractually required to say something — a heartbeat or ping interval, with enough room for two missed beats. If the protocol has no heartbeat, add one, or send your own ping when the read deadline approaches. A deadline is only meaningful if silence actually means "broken". **Bound writes too.** A `Write` blocks when the peer stops reading and the send window fills; `SetWriteDeadline` is the same fix in the other direction, and a wire layer that bounds only reads still has a way to hang. **Keep TCP keep-alives as a backstop, not the mechanism.** Go's `net.Dialer` enables TCP keep-alives on the connections it creates, and they will eventually kill a half-open connection, but how long depends on the probe interval and retry count in the OS, which can add up to many minutes. Application deadlines are the control you actually own. **Decide what a timeout means for the stream.** For a length-prefixed protocol, a deadline that fires after part of a frame has been consumed leaves the stream misaligned — the next read would treat body bytes as a length header. Treat that as fatal: close the connection and reconnect. Only a timeout while idle between frames is safe to shrug off. ## Making the failure loud next time Two cheap habits keep this from recurring. Log at the wire layer when a read deadline fires, with the connection's age and idle time, so a hang becomes a visible reconnect instead of silence. And give the CLI a global upper bound so a stuck exchange ends the program with a diagnosable error rather than an unkillable prompt — a user who has to reach for Ctrl-C is a user who cannot tell your bug from the network's.

  • Why does the operating system not report that the peer is gone?
    TCP sends nothing while a connection is idle. A crashed process still gets an RST from its host's kernel, but a dead host, a terminated instance or a firewall that drops flow state sends nothing at all. The connection is half-open: healthy locally, nonexistent remotely, and discoverable only by probing.
  • How would you choose the read deadline for a stream that is idle most of the time?
    Derive it from the protocol's heartbeat interval — roughly two missed beats — not from a request latency target, which would fire constantly on an idle stream. If the protocol has no heartbeat, add one, or send a ping as the deadline approaches. A deadline only means something when silence genuinely implies breakage.
  • Would enabling TCP keep-alives alone have fixed this?
    Not adequately. Go's `net.Dialer` already enables keep-alives on the connections it creates, and they do eventually tear down a half-open connection, but the probe interval and retry count come from OS settings and can take many minutes. Keep them as a backstop; the application deadline is the bound you control.
  • The deadline now fires halfway through reading a frame. What should the wire layer do?
    Close the connection and reconnect. The bytes already consumed are gone, so the stream is misaligned and the next read would parse body bytes as a length header. Only a timeout that consumed nothing — fired while idle between frames — is safe to treat as recoverable on the same connection.

saying these in an interview costs you the question

  • Assumes TCP always notices when the peer disappears
  • Clears the deadline because the stream is long-lived
  • Bounds only the handshake and not the steady-state reads
  • Relies on TCP keep-alives as the primary timeout
  • Resumes reading a framed stream after a mid-frame timeout
  • Blames the network without taking a goroutine dump
  • Bounds reads but leaves writes unbounded