skip to content

A bulk uploader's in-flight permit count sits pinned at its semaphore limit with nothing finishing — how do you find the cause?

level: seniorimportance: should knowfreq 35%

answer

  1. count what goes out against what comes back
  2. stuck, not slow
  3. an early return skips a defer never reached
  4. dump the stacks and look for the send

basics

~20 s

That shape means permits were taken and never returned. Look for an acquire whose release is not deferred immediately after it, since an early error return above the defer never registers it. Confirm with the goroutine profile.

solid answer

~50 s

A count pinned at the limit with zero completions is a leak, not slowness: a slow dependency still finishes work, just later. Take the goroutine profile and read the stacks — a leak shows as many goroutines parked in a channel send at the acquire site and nothing at all inside the guarded work, whereas a slow downstream shows goroutines sitting in the network read. Then read the acquiring function for exit paths above the release: the classic bug is `sem <- struct{}{}` followed by a few error checks that `return` before `defer func(){ <-sem }()` is reached, so on the error path the permit is never given back and the effective limit ratchets down by one each time. The fix is to register the release as the very next statement after a successful acquire, and to make that pairing something a reviewer can see in one line rather than trusting every future branch.

code

go · 10 lines
go
func (u *Uploader) put(part Part) error {
	u.sem <- struct{}{}
	rc, err := part.Open()
	if err != nil {
		return err // permit never returned
	}
	defer func() { <-u.sem }()
	defer rc.Close()
	return u.send(rc)
}

go deeper

for a junior

Recall the failure mode by name: a permit taken and not returned is gone permanently, and enough of them stop everything. Know that the release belongs in a defer registered right after the acquire.

for a middle

Explain why a defer below an early return never runs, and why that makes the bug survive review. Be able to describe how the effective limit ratchets down one permit at a time.

for a senior

Walk the diagnosis end to end: distinguish stuck from slow by whether anything completes, take the goroutine profile from the wedged process, read the stacks for a parked acquire with no active workers, then find the exit path above the release.

for a principal

Argue for the structural fix and the signal: one helper that pairs acquire with release so the pattern cannot be half-copied, plus an outstanding-permit gauge alerted on when it is pinned with zero completions, so the failure is named rather than restarted away.

## Reading the symptom first The gauge is an application counter: incremented on a successful acquire, decremented on release. Pinned at the limit with zero completions is a distinctive shape, and it is worth separating from its neighbours before touching any code. - **Slow dependency**: the count sits high, but work still completes — latency is up, throughput is down, and both are non-zero. - **Saturation with a correct semaphore**: the count oscillates against the limit as permits go out and come back. - **Leak**: the count climbs monotonically, sticks at the limit, and completions go to zero and stay there. Nothing recovers it, including waiting; only a restart clears it, which is why it is often misdiagnosed as a memory problem. Monotonic-and-stuck is the tell. Permits are being held by code that is no longer running. ## Confirming it with the goroutine profile The goroutine profile — from `runtime/pprof`, or over HTTP at `/debug/pprof/goroutine` if the process serves the pprof endpoints — dumps every live goroutine's stack, grouped by identical stacks with a count. This is the right diagnostic here for a specific reason: the process is doing nothing, so a CPU profile is almost empty and tells you nothing, and a heap profile shows what was allocated rather than where execution is stuck. What a leak looks like in that dump: - A large count of goroutines with an identical stack, blocked in a channel send, at the line where the semaphore is acquired. - **Nothing** inside the guarded work — no goroutine in the upload path, no goroutine in a network read. That second observation is what distinguishes a leak from congestion. If the permits were genuinely in use, the profile would show as many goroutines inside the guarded work as there are permits. Zero of them, with the count pinned, means the holders have finished and taken their permits with them. ## Finding the missing release Now read the acquiring function looking for exit paths between the acquire and the release. The canonical bug: ```go func (u *Uploader) put(part Part) error { u.sem <- struct{}{} rc, err := part.Open() if err != nil { return err // permit never returned } defer func() { <-u.sem }() defer rc.Close() return u.send(rc) } ``` The release *is* deferred — which is why it survives review — but the `defer` statement sits **below** an early return. A deferred call only runs if its `defer` statement actually executed; the error path never reaches it. Each failed `part.Open()` permanently consumes one permit. With a limit of eight, the ninth such failure wedges the uploader forever, and the errors it returns look like ordinary transient failures rather than the cause of a total stall. The correction is one line moved: ```go func (u *Uploader) put(part Part) error { u.sem <- struct{}{} defer func() { <-u.sem }() // registered before anything can fail rc, err := part.Open() if err != nil { return err } defer rc.Close() return u.send(rc) } ``` ## Other ways permits go missing - **A release inside branches rather than deferred.** Works until a new branch forgets. - **A panic in the guarded work that a `recover` elsewhere swallows.** Unwinding runs deferred calls in that goroutine, so a deferred release is safe; a non-deferred one is not. - **Releasing on a path that never acquired.** The mirror image: a receive from an empty semaphore channel blocks the releaser forever, or, with a weighted semaphore, panics for releasing more than is held. The gauge then reads *below* zero-ish rather than pinned, but the stall is just as complete. - **Two different exit points, each with its own release.** Under an error path that hits both, the count drifts the other way and the limit stops limiting. ## Making the class of bug impossible The durable fix is structural rather than a patched branch. Pair the acquire and the release in one place so no call site can half-copy the pattern — a small helper that acquires and hands back a release function, called as an acquire followed immediately by a deferred call to what it returned. Then the review rule is a single, checkable one: an acquire is always followed on the next line by the deferred release, with no statement in between that can return or panic. It is also worth keeping the gauge. A counter of outstanding permits, exported and alerted on when it stays at the limit while completions are zero, turns a total stall that a restart hides into a fault you can name in seconds. ## What not to do Raising the limit is the reflex and it is wrong: a leak consumes whatever you grant it, so a bigger limit only buys time proportional to how many extra permits you added. Restarting clears the count and hides the evidence. Diagnose it while it is stuck — the goroutine profile from the wedged process is the whole story.

  • In a goroutine dump, what distinguishes a leaked permit from ordinary saturation?
    Saturation shows as many goroutines inside the guarded work as there are permits, plus the queue parked on the acquire. A leak shows the parked queue with nothing at all inside the guarded work — the holders have already returned. The absence of active workers is the discriminator.
  • Why is raising the limit the wrong first response?
    A leak consumes whatever you grant it, so a larger limit only postpones the stall in proportion to the permits added, and it does so while making the failure rarer and harder to catch. It also widens real concurrency against a dependency that may already be complaining about load.
  • What if a release runs on a path that never acquired?
    With a channel semaphore, the receive blocks on an empty buffer and the releasing goroutine hangs; with a weighted semaphore, releasing more than is held panics. Either way the accounting is corrupt, so the review rule cuts both ways: exactly one release per successful acquire.
  • How do you make this class of bug hard to reintroduce?
    Pair the acquire and release in one helper that hands back a release function, so a call site is an acquire followed immediately by a deferred call to it. The review rule becomes checkable at a glance: nothing that can return or panic may sit between the acquire and the deferred release.

Tokens that go out and never come back. The table ends up empty, so nobody can start, even though nobody is working.

saying these in an interview costs you the question

  • Blames the downstream service without checking whether anything completes
  • Raises the limit instead of finding the missing release
  • Registers the release defer below the error checks
  • Thinks a blocked channel send eventually times out by itself
  • Restarts the process and calls the incident resolved
  • Reaches for a CPU profile when nothing is running