skip to content

In a Go service, what does exporting len(jobs) on a bounded intake channel actually measure?

level: middleimportance: nice to knowfreq 34%

answer

  1. a count of what, exactly?
  2. the workers are holding items too
  3. it stops moving once it is full
  4. unbuffered channels always report zero
  5. pair depth with block time and shed count

basics

~20 s

It measures how many items are sitting in that channel's buffer at the instant you read it — nothing else. It excludes items already taken by workers and producers blocked waiting to send, it is stale immediately, and it is always zero for an unbuffered channel.

solid answer

~50 s

`len(jobs)` on a channel returns the number of elements currently in its buffer, and `cap(jobs)` returns the buffer size, so `len/cap` is a saturation gauge for the intake queue. Its limits matter more than its value. It does not count work in flight inside the workers, and it does not count producers already parked on a full send, so a saturated pipeline reports exactly `cap` no matter how bad the overload is — the gauge saturates before the system does. It is an instantaneous sample that is out of date the moment it returns, and on an unbuffered channel it is always zero even with a hundred blocked senders. Export it as a ratio with `cap`, but pair it with the signals that keep resolution past saturation: time spent blocked in the enqueue path, a counter of shed or timed-out events, and the age of the oldest queued item.

code

go · 7 lines
go
// jobs was created as make(chan Job, 256)
func queueDepth(jobs chan Job) (queued, capacity int) {
	return len(jobs), cap(jobs) // buffered now; the fixed ceiling
}

// saturation = float64(queued) / float64(capacity)
// pins at 1.0 under overload, so export block time and shed count too

go deeper

for a junior

Recall that len on a channel counts buffered elements and cap gives the buffer size, and that both are fixed-cost reads. Know that an unbuffered channel always reports a length of zero.

for a middle

Explain what the number excludes — in-flight items and parked senders — and why it saturates at cap. Be able to say what you would export beside it: enqueue block time, a shed counter, queueing delay.

for a senior

Interpret the gauge in an incident: pinned at cap means consumer-bound, oscillating means bursty, zero with rising latency means the bottleneck is elsewhere. Explain why depth alone is a poor thing to page on.

for a principal

Decide what this service actually publishes as its saturation signal and what the alerting threshold means in business terms. Depth is a number; the thing worth committing to is queueing delay and how many events you are prepared to shed.

## What the builtins actually return for a channel For a channel, `len(ch)` is the number of elements queued in its buffer and `cap(ch)` is the buffer's size. For `jobs := make(chan Job, 256)`, `cap(jobs)` is 256 forever, and `len(jobs)` moves between 0 and 256. Both are cheap, non-blocking reads, which is why they are tempting as a metric. ## Why it is a useful gauge Queue depth is the most direct in-process statement of "are the consumers keeping up". A pipeline whose intake sits near zero has spare capacity; one that sits near `cap` is consumer-bound and its producers are being stalled right now. Because the ceiling is known, the *ratio* `len/cap` is comparable across deployments and across differently sized queues, which a raw count is not. It is also a leading indicator relative to latency: the queue fills before end-to-end latency visibly degrades. ## The four things it does not tell you **It saturates.** Once the intake is full, the gauge reads `cap` whether the service is 10% behind or 500% behind. All the information about *how far* behind lives in the producers that are now parked on the send — and those are invisible to `len`. **It excludes work in flight.** An item a worker has received but not finished is out of the buffer and not yet done. With 32 workers each holding an item, the real amount of unfinished work is `len(jobs) + 32`, and if the workers are slow that hidden portion can dominate. **It is instantaneous and immediately stale.** By the time the value reaches your metrics pipeline it describes a moment that has passed. A scrape every 15 seconds over a queue that fills and drains in milliseconds will sample almost anything. A high-water mark or an occupancy histogram carries far more than a periodic point sample. **It is zero for an unbuffered channel.** `make(chan Job)` has no buffer, so `len` is 0 even when many goroutines are blocked sending. Reading a persistent zero as "healthy" here is exactly backwards: it may be a permanent rendezvous stall. ## What to export alongside it - **Enqueue block time.** Measure how long the producer spent in the send. This keeps resolving after depth pins at `cap`, and it is the number that maps to user-visible latency. - **A shed counter.** Every event abandoned because the queue would not take it, counted and exported. Depth tells you the queue is full; this tells you what that cost. - **Oldest-item age.** Timestamp items on entry and record the age of the one being dequeued. This is the queueing delay, and unlike depth it is meaningful in seconds rather than in items. - **Worker occupancy.** How many of the N workers are busy, so you can distinguish "queue full because workers are saturated" from "queue full because workers are blocked on a dependency". ## Reading the gauge in an incident Depth pinned at `cap` for minutes means consumers are the bottleneck and every producer send is now blocking; the fix is downstream capacity or shedding, not a bigger buffer. Depth oscillating between 0 and `cap` means the arrival rate is bursty and the buffer is doing its job absorbing bursts. Depth at zero with rising latency means the bottleneck is not the intake at all — look inside the workers or further downstream. Depth slowly climbing under constant input means drain rate is below arrival rate, and if the queue were unbounded that same signal would have been heap growth instead. ## A caution about how you sample it Calling `len` from a metrics goroutine is safe and lock-free, but do not build control logic on it — deciding whether to send based on a previous `len` read is a race, because the value can change between the check and the send. Depth is for humans and dashboards; flow control belongs in the send itself.

  • Why does queue depth stop being informative under heavy overload?
    It saturates at `cap`. Once the buffer is full, the excess pressure lives in producers parked on the send, and `len` cannot see them, so a 10% overload and a 500% overload report the same number. Enqueue block time and the shed counter keep resolving after depth has pinned, which is why they belong on the same dashboard.
  • Would you use len(ch) to decide whether to send?
    No. The value can change between the read and the send, so any decision based on it is racy — the queue can fill or drain in between. Depth is an observability signal. Flow control belongs in the send itself, where blocking, cancellation and a deadline give you a decision that is actually atomic with the enqueue.
  • What does a persistently zero depth on an unbuffered intake channel mean?
    Nothing reassuring. An unbuffered channel has no buffer, so `len` is always 0 regardless of how many senders are parked waiting for a receiver. To see pressure there you need the enqueue block time or the goroutine profile showing goroutines blocked in a channel send.

saying these in an interview costs you the question

  • Reads a low depth as proof the pipeline is healthy
  • Forgets items already taken by workers are unfinished work
  • Expects len on an unbuffered channel to count blocked senders
  • Uses a previous len reading to decide whether to send
  • Responds to a permanently full queue by enlarging the buffer