A program uses channels exclusively and holds no locks at all, yet it hangs. What kinds of channel usage cause that, and how do you diagnose and prevent it?
answer
- wait-for cycle; resource = a matching partner
- unmatched send: caller timed out, sender parked
- producer returned without closing -> consumer waits
- mutual send before receive on unbuffered = hang
- buffers postpone, topology fixes; task-count growth = leak
basics
~20 sChannel operations block, so tasks can wait in a cycle just like lock holders do. Typical causes: a send with no receiver ever arriving, a receive with no sender, two tasks sending to each other over unbuffered channels, and a loop waiting for a stream nobody closes.
solid answer
~60 sChannels remove data races, not blocked-waiting cycles. A send blocks until a partner is ready (or buffer space exists), so a task waiting for a communication that will never occur is deadlocked exactly like a task waiting for a lock that is never released. The recurring shapes: - **Unmatched send**: a value is sent on a channel nobody ever receives from — the receiver returned early, crashed, or was never started. - **Unmatched receive**: a consumer waits on a channel whose producer already finished without closing it, so the end of stream is never signalled. - **Cyclic topology**: A sends to B before receiving from B while B does the mirror image; with unbuffered channels neither send can complete. - **Full buffers all round**: every task is blocked in a send on a full channel, so nobody reaches its receive. - **Self-communication**: a task sending on a channel it is itself supposed to drain. Diagnose by dumping task stacks: deadlocked tasks are parked at a send or receive, which names the channel. Prevent with an acyclic flow direction, closing streams, and putting every blocking operation in a select with a cancel or timeout arm.
code
text · 9 lines# leaky
reply = channel(0)
send(work, Job(payload, reply))
select:
case r = receive(reply): use(r)
case receive(timer(1s)): return # worker later blocks in send(reply, r) forever
# fixed: the responder can always deposit and move on
reply = channel(1)go deeper
Know that send and receive block, so a send with no receiver (or a receive with no sender) hangs, and that two tasks sending to each other can both get stuck.
Explain it as a wait-for cycle where the resource is a matching partner, and name the missing-close and full-buffer variants.
Focus on partial deadlock in production: timed-out callers stranding responders, growing task counts, stack dumps as the diagnostic, cancel arms on every blocking operation.
Argue for structural guarantees — acyclic data flow, scoped task lifetimes so every spawned task has an owner that cancels it, capacity-1 replies — so that liveness is a property of the topology rather than of reviewer vigilance.
## Why a lock-free program still hangs Deadlock is not about locks; it is about a cycle in a wait-for graph. With mutexes the resource is a lock. With channels the resource is a **matching partner**: a send waits for a receiver or for buffer space, a receive waits for a sender. If you draw an edge from each blocked task to the task that could unblock it, a cycle in that graph is a hang, and the graph does not care what primitive created it. The difference is that channel deadlocks are usually structural — a property of the topology and the order of operations — rather than timing-dependent, so they often reproduce every run instead of once a week. ## The recurring shapes **Unmatched send.** The commonest one in real code. A task sends a result on a reply channel, but the caller took a timeout branch and moved on; nobody will ever receive, so the sender parks forever. Note that this frequently is not a *global* hang: the rest of the program works, and you simply accumulate one parked task per timed-out request until memory runs out. Making the reply channel buffered with capacity one is the standard fix, because the abandoned reply can then be dropped without blocking anyone. **Unmatched receive.** A consumer loops receiving until end-of-stream, but the producer returned without closing the channel. The consumer waits forever for a value that will never come. **Cyclic topology.** Two stages that each send to the other before receiving. On unbuffered channels, both sends block and neither reaches its receive. A: send(toB, x); v = receive(fromB) B: send(toA, y); w = receive(fromA) -> both parked in send; classic communication deadlock **Buffer exhaustion.** The same cycle appears at scale: in a ring or feedback pipeline, once every channel is full every task is blocked sending and none reaches its receive. Buffers postponed the deadlock; they did not remove it. **Self-communication.** A task sends on a channel that it alone drains later in the same sequential path — it can never reach the drain. ## Diagnosis The fastest tool is a dump of all task stacks. Blocked tasks are parked precisely at a channel operation, which names both the operation and the channel, so a cycle is usually visible by inspection. Some runtimes detect the special case where *every* task is asleep and abort with a deadlock message, but that only catches whole-program deadlock; the common partial deadlock — a handful of leaked tasks while the service keeps serving — is silent. Watch the count of live tasks over time: monotonic growth with steady load is the signature. ## Prevention **Make the flow acyclic.** A pipeline with one direction of data cannot have a communication cycle. If feedback is genuinely needed, break it with a dedicated task that owns the loop and never blocks in both directions at once. **Bound every wait.** Every long-lived send and receive goes inside a select with a cancel arm, and often a timeout arm. A task that can always observe cancellation cannot be permanently stranded. **Give replies room.** Request/reply channels get capacity one so a responder is never blocked by a caller that gave up. **Close streams.** Closing is how consumers learn there will be no more values; a producer that returns without closing strands them. **Do not treat buffering as a fix.** Increasing capacity because a hang went away hides the topology bug and reintroduces it under load. Use capacity for throughput, and fix cycles structurally. **Own the lifecycle.** Whoever starts a task is responsible for having a path that ends it. Structured spawning — where a scope waits for its children and cancels them on exit — turns leaked tasks into an impossible state rather than a review question. ## The contrast with lock deadlock Lock deadlock is fixed by a global acquisition order. Channel deadlock is fixed by topology and by making every wait interruptible, because there is no ordering to impose — the resource is another task's willingness to communicate.
- How would you confirm a suspected channel deadlock in a running service?Dump the stacks of all tasks and look for ones parked at a send or receive; the frame names the channel, and the set of blocked tasks usually reveals the cycle directly. Also trend the live task count and channel occupancy: a steadily rising task count under steady load points to leaked tasks blocked on unmatched sends rather than a full-program deadlock, which most runtimes would abort on.
- Does adding buffer capacity ever legitimately fix a deadlock?It legitimately fixes exactly one case: a bounded, known number of outstanding messages, such as a reply channel with capacity one so a responder never blocks on an absent caller. For cyclic topologies it only raises the number of in-flight messages needed to hang, so the deadlock returns under load. Treat a capacity increase that makes a hang disappear as evidence of a structural bug.
saying these in an interview costs you the question
- Claiming a program without locks cannot deadlock.
- Bumping channel capacity until the hang stops and shipping that as the fix.
- Assuming the runtime will always detect and report deadlock — it typically detects only all-tasks-asleep, not leaked tasks.
- Applying lock-ordering advice to channels; the fix is acyclic topology plus cancellable waits.
- Forgetting that a caller that timed out leaves the responder blocked on an unbuffered reply channel.