skip to content

A metrics daemon builds every StatsD line with fmt.Sprintf and GC CPU climbs with traffic. How do you cut the allocations per line?

level: seniorimportance: should knowfreq 40%

answer

  1. the string is the middleman
  2. who owns the bytes between calls
  3. strconv has an Append family
  4. reset the length, keep the capacity
  5. count allocations per line before and after

basics

~20 s

Stop producing strings. Keep one reusable byte slice per emitter, reset it with buf = buf[:0], build the line with append and strconv.AppendInt, and write those bytes straight out. That removes the boxing, the format scan and two string copies.

solid answer

~50 s

First I would price it: allocations per line from a benchmark run with `-benchmem`, times the measured emit rate, against the process's allocation rate. If it is real, the Sprintf version is paying a box per argument, a run-time format scan, a result string, then a second copy back to `[]byte` for the write — four or five allocations for twenty bytes. The fix is to own the buffer: `buf = buf[:0]`, then `append` the literal parts and `strconv.AppendInt` the numbers, then write the slice. That is zero allocations per line once the buffer reaches its steady-state capacity. The buffer must be per-goroutine or from a `sync.Pool`, never shared. `fmt.Appendf` into the same buffer is the middle ground: it drops the string copies but keeps the boxing. I would test the fast rendering against the readable one for byte-identical output.

code

go · 2 lines
go
line := fmt.Sprintf("%s:%d|c\n", name, delta)
w.Write([]byte(line)) // the conversion copies the whole line again

go deeper

for a junior

Know that appending into an existing byte slice avoids creating a new string each time, and that strconv has Append functions built for exactly that.

for a middle

Enumerate what one Sprintf line allocates — a box per argument, the result string, and usually a copy back to bytes — and write the append-based replacement correctly, including resetting the slice length.

for a senior

Show the whole loop: measure, change the smallest thing, prove the bytes are identical to the readable version, and confirm the allocation rate actually fell. Say plainly when you would not bother.

for a principal

Ask whether the daemon should emit less rather than format faster. Sampling, batching and pre-rendered prefixes remove the cost by a factor and cost no readability, so they come before any per-verb rewrite.

### Inventory the cost of one line The daemon renders something like `api.requests:1|c` for every counter update and writes it out. The obvious implementation is: ```go line := fmt.Sprintf("%s:%d|c\n", name, delta) w.Write([]byte(line)) ``` Count what that costs per line: 1. `name` is boxed into an `any`. A string box is a pointer-shaped pair; the string header itself has to be somewhere the interface can point at, so this generally allocates. 2. `delta` is boxed too — another small allocation unless it happens to be a byte-sized value. 3. The format string is scanned byte by byte, verbs are matched to arguments, and each argument is dispatched on. 4. The scratch buffer is copied into a new `string`. 5. `[]byte(line)` copies it *again*, because `io.Writer` takes bytes and strings are immutable. That is four or five allocations for a line of about twenty bytes. Multiply by the emit rate. At a thousand lines a second nobody notices; at a million, the allocation rate alone will show up as garbage-collector CPU, and the CPU cost rises linearly with traffic — which is exactly the symptom described. ### The rewrite Stop producing strings. Own a byte slice, reset its length, append into it, and write it: ```go func appendCounter(buf []byte, name string, delta int64) []byte { buf = append(buf, name...) buf = append(buf, ':') buf = strconv.AppendInt(buf, delta, 10) return append(buf, "|c\n"...) } ``` `strconv.AppendInt(dst []byte, i int64, base int) []byte` writes the digits straight into the caller's slice and returns the extended slice. `strconv` has the same shape for the other kinds — `AppendQuote`, `AppendFloat`, `AppendBool`. Nothing is boxed, no format string is interpreted, and no string is ever created. After the first few lines the buffer has reached its steady-state capacity and subsequent calls allocate nothing at all: this is the version that reaches zero allocations per line. ### Buffer ownership is the part people get wrong - **Reset the length, keep the capacity.** `buf = buf[:0]` before each line. Forgetting it does not crash; it silently concatenates every line ever emitted and grows without bound. - **Do not share it.** One buffer per emitting goroutine, or a `sync.Pool` if emitters come and go. A shared buffer without synchronisation is a data race, and the race detector will only find it on a run that actually interleaves. - **Do not let the writer keep it.** `io.Writer` implementations are required not to retain the slice they are handed, so writing and then reusing is legal — but if you hand the bytes to something that buffers them for later, it must copy. ### The middle ground If you want to keep the readability of verbs, `fmt.Appendf(buf, format, args...)` formats into your slice and returns it. It removes the result string and the second copy, but it is still a variadic `...any`, so the boxing remains and the format string is still scanned per call. Expect it to remove roughly half the allocations, not all of them. `fmt.Fprintf(w, ...)` writing directly into a `strings.Builder` or straight to the connection gets the same partial win — it never materialises the intermediate string that `Sprintf` returns. ### Prove it, in both directions Correctness first: keep the readable implementation and add a table test asserting the appended bytes are identical to it, including negative deltas, empty names and values around the small-integer boundary. A fast formatter that quietly emits the wrong separator is worse than a slow one. Then the win: a benchmark run with `-benchmem` gives `allocs/op` and `B/op` for both renderings. That number times the measured emit rate is the argument you take to review — a ratio on its own proves nothing about the service. ### And know when not to Everything above buys you allocations per line. If the daemon emits a few hundred lines a second, the rewrite removes nothing measurable and costs every future reader a paragraph of explanation. Emitting fewer lines — sampling, batching several counters per packet, pre-rendering the constant prefix of a metric name once at construction — is usually the larger and cheaper win, and it keeps `fmt` where it reads best.

  • Does fmt.Appendf get you all the way to zero allocations?
    No. It removes the result string and lets you reuse your own slice, but its parameter is still a `...any`, so every argument is boxed and the format string is still scanned on each call. Expect it to remove roughly half the allocations. The `strconv.Append*` form is what takes a simple line to zero.
  • Where should the reusable byte slice live?
    With the goroutine that emits, or in a `sync.Pool` if emitters come and go. It must never be shared without synchronisation. It is safe to reuse it immediately after a write, because `io.Writer` implementations are required not to retain the slice they are given — but anything that buffers your bytes for later has to copy them.
  • What would make you decide not to do this at all?
    A low emit rate. Allocations per line times lines per second, compared against the service's total allocation rate, decides it. If the daemon emits a few hundred lines a second the rewrite removes nothing and costs every future reader an explanation. Emitting less — sampling, batching, pre-rendering the constant prefix — is usually the bigger win anyway.
  • How do you keep the hand-rolled formatter honest?
    Keep the readable implementation as the reference and add a table test asserting the appended bytes match it exactly, covering negative values, empty names and the boundaries. Speed is worthless if the separator quietly changes; a benchmark proves throughput, only a comparison test proves the output.

saying these in an interview costs you the question

  • Rewrites the formatter without measuring the emit rate
  • Shares one scratch byte slice across goroutines
  • Forgets buf = buf[:0] so the slice grows forever
  • Assumes fmt.Appendf removes the argument boxing
  • Converts the finished string back to bytes on every write