skip to content

A metrics daemon sends into make(chan Sample, 50000) and its sink is slow: no errors, but samples land minutes stale. What is the buffer doing?

level: seniorimportance: should knowfreq 42%

answer

  1. a buffer buys time, not throughput
  2. queued samples are ageing in place
  3. capacity divided by the rate deficit
  4. measure arrival minus capture time
  5. steady state means one buffer of staleness

basics

~20 s

The buffer is converting a throughput shortfall into latency. It absorbs the gap between production and consumption only until it fills, so samples queue and age instead of failing. Capacity buys time; it never raises the sink's rate.

solid answer

~50 s

A buffer absorbs *bursts*, not a sustained rate mismatch. If the daemon captures 1000 samples a second and the sink drains 800, the buffer swallows 200 a second and fills after `50000/200`, about four minutes — after which the sender blocks anyway. What you bought is a queue whose contents are, at steady state, a full capacity old by the time they arrive. Because `ch <- s` has no return value, none of that shows up as an error; the only symptoms are a producer that occasionally parks and data that is stale. The way to see it is to stamp each `Sample` at capture and measure sink-arrival minus capture time — end-to-end lag — rather than watching for failures. The fix is on the consuming side: parallelise or batch the sink, aggregate rather than forward every sample, and size the capacity to a burst you can name (a second or two of production), not to the outage you hope to ride out.

code

go · 11 lines
go
type Sample struct {
	Name  string
	Value float64
	At    time.Time // stamped at capture
}

// in the goroutine draining the channel, once the sink write has succeeded
lag := time.Since(s.At)
if lag > 30*time.Second {
	log.Printf("sample %s delivered %v after capture", s.Name, lag)
}

go deeper

for a junior

Understand the basic shape: a channel buffer holds values until someone takes them, so if nobody keeps up, values wait there and arrive late rather than being lost.

for a middle

Be able to do the arithmetic out loud: capacity divided by the rate deficit is how long until the buffer fills, and capacity divided by the drain rate is the staleness once it has.

for a senior

Show that you instrument age rather than errors, that you fix the consuming side before touching the number, and that you can say what should happen when the channel is full.

for a principal

Take the position that freshness is the service level for a telemetry pipeline, and that a capacity constant silently trading freshness for completeness is a design choice that must be made deliberately and reviewed.

## The arithmetic nobody does before picking the number A channel buffer is a queue, and a queue drains only if the consumer's average rate is at least the producer's. Write it out for the daemon: - production: 1000 samples/second - sink drain: 800 samples/second - deficit: 200 samples/second - capacity: 50000 The buffer absorbs the deficit for `50000 / 200 = 250` seconds. After roughly four minutes it is full, and the sending goroutine blocks on `ch <- s` exactly as it would have with no buffer at all. The buffer did not fix the mismatch; it postponed the visible consequence and, in the meantime, filled itself with data that was ageing. And this is the part that matters for a metrics daemon specifically: once the buffer is full and staying full, **every sample delivered is one full buffer old**. At 800/second draining a 50000-deep queue, a sample waits about 62 seconds before it even reaches the sink. Dashboards go quietly wrong — not empty, not erroring, just describing the past. ## Why nothing alerts Go gives you no hook here. A send is a statement, not a call with a result: `ch <- s` cannot return "the buffer is full" the way an enqueue API might return an error. The channel does not drop, does not overwrite and does not log. The observable effects are indirect: a producing goroutine that spends time parked, and output that is late. If your monitoring watches error rates and send counts, all of it stays green — you are sending everything, and everything eventually arrives. Capacity also cannot be turned down at runtime: it is fixed at `make`, so there is no dial to twiddle while you investigate. ## The diagnostic: measure age, not volume Give each sample a capture timestamp, and at the sink measure how long ago that was. End-to-end lag is the signal that directly encodes the failure: - lag flat and small — the buffer is doing its job, absorbing jitter. - lag climbing steadily — production exceeds consumption right now; the buffer is filling and you can extrapolate when it will be full. - lag high and *flat* — steady state with a full buffer. The plateau value is roughly capacity divided by drain rate, which is the clearest possible confirmation that the buffer, not the network, owns the delay. A blocked sender is visible in a goroutine profile too — the producing goroutine parked in a channel send — but that only tells you it is stuck now, whereas lag tells you how wrong the data has been all along. ## Fixing it The fix is never a bigger number. In rough order: 1. **Raise consumption.** Batch samples into fewer sink writes, run several sender goroutines against the sink, or reduce per-write overhead. This is the only change that alters the arithmetic. 2. **Reduce production.** Aggregate at capture — a gauge does not need every raw sample forwarded; pre-aggregate per interval and send summaries. 3. **Decide explicitly what happens when the channel is full.** Blocking the sampler and dropping stale samples are both defensible for telemetry, but the choice should be written down in the code, not implied by a capacity constant. For a metrics pipeline, fresh data usually beats complete data — which makes silently queueing for minutes the worst of the options. 4. **Then size the buffer**, to a burst you can name: enough for a garbage-collection pause or a brief sink hiccup, typically a second or two of production. Since capacity is fixed at `make`, that number should come from a measurement, not from an unused zero. ## The general lesson A buffer's honest purpose is to decouple two components in *time* so a brief stall on one side does not stall the other. Used that way, small capacities are enough. Used as insurance against a consumer that is simply too slow, a buffer becomes a mechanism for hiding the problem and degrading the product at the same time: it removes the back-pressure signal that would have told you at once, and it delivers old data while looking healthy. An unbuffered channel, by contrast, would have stalled the sampler immediately and made the slow sink obvious on day one — which is why unbuffered is the better default until you can state what the buffer is for.

  • Why does raising the capacity from 50000 to 500000 not fix this?
    Because the drain rate is unchanged. A buffer empties only if consumption is at least production on average; with a 200/second deficit the bigger buffer just takes ten times longer to fill, and once full the steady-state staleness is ten times worse. You have made the same failure quieter and the data older.
  • Would an unbuffered channel have surfaced this sooner?
    Yes. With `make(chan Sample)` the sampler stalls the instant the sink slows, so missed sampling intervals or a producer parked in a send show up immediately, on day one rather than after a launch. The cost is that every ordinary hiccup on the sink side also stalls sampling — which is exactly the tradeoff a small buffer is meant to buy off.
  • If buffering is still the right call, how do you pick the number?
    From a burst you can name: how long a pause you want to absorb times the production rate — a garbage-collection pause or a brief sink retry, so typically a second or two of samples. Then check the memory cost, since capacity times element size is allocated at `make` and cannot be changed later.
  • What would you monitor so this is caught next time?
    The age of delivered data, as sink-arrival minus capture timestamp, with an alert on a percentile rather than on errors. Error rate and send throughput both look healthy in this failure; staleness is the only signal that moves, and it is also the quantity the consumers of the metrics actually care about.

A bigger in-tray does not make you read faster. It lets people keep dropping paper in for longer, and guarantees that whatever you finally pick up describes a situation from an hour ago.

saying these in an interview costs you the question

  • Raises the capacity and calls it fixed
  • Says a full channel means the channel is broken
  • Measures only error rate, never data age
  • Claims buffering increases consumer throughput
  • Assumes a slow sink would have thrown an error