skip to content

Your Go log-shipping agent gzips rotated files before upload. How do you choose the compression level, and who can overrule you?

level: principalimportance: should knowfreq 30%

answer

  1. the default is already a choice
  2. measure on real lines, not a synthetic file
  3. two numbers: bytes out, CPU per gigabyte
  4. the CPU and the saving have different owners
  5. make it config so 3am can change it

basics

~20 s

Choose the level from a benchmark over real log lines that records bytes out and CPU per gigabyte, not from a default. The CPU lands on the service owner's node and the saving on someone else's bill.

solid answer

~50 s

The knob is `gzip.NewWriterLevel(w, level)`, from `gzip.BestSpeed` (1) through `gzip.DefaultCompression` (-1, roughly level 6) to `gzip.BestCompression` (9), plus `gzip.NoCompression` and `gzip.HuffmanOnly`; anything else returns an error rather than being clamped. I would not pick from folklore: I benchmark the agent's own encoder with `-benchmem` over a representative sample of our log lines and report two numbers, output bytes and CPU-seconds per gigabyte. On typical text, moving from 1 to 6 buys a lot of ratio cheaply and 6 to 9 usually buys a few percent for several times the CPU. Then it is a cost decision with two owners: the agent shares a node with the service, so the CPU is the service owner's latency risk, while the saving shows up on someone else's storage and egress bill. I make the level configuration with a measured default, so it can be changed without a release, and I reuse writers via a `sync.Pool` with `Reset` so the choice is not swamped by allocation.

code

go · 15 lines
go
var gzPool = sync.Pool{New: func() any {
	// BestSpeed is a valid constant, so this cannot fail.
	zw, _ := gzip.NewWriterLevel(io.Discard, gzip.BestSpeed)
	return zw
}}

func compress(dst io.Writer, src io.Reader) error {
	zw := gzPool.Get().(*gzip.Writer)
	defer gzPool.Put(zw)
	zw.Reset(dst)
	if _, err := io.Copy(zw, src); err != nil {
		return err
	}
	return zw.Close()
}

go deeper

for a junior

Know that gzip has selectable levels between fastest and smallest, and that you pick one with gzip.NewWriterLevel rather than accepting the default without thought.

for a middle

Be ready to name the constants and their numeric values, to say that an invalid level errors rather than clamps, and to describe the shape of the ratio-versus-CPU curve.

for a senior

Show how you measure on representative data, reuse writers with a sync.Pool and Reset, and make the level a validated configuration value with a stated CPU posture.

for a principal

Own the tradeoff as a cost allocation across two budgets, bring the shared numbers that let it be decided rather than argued, and say who is entitled to overrule the level in each direction.

## The knob, precisely `compress/gzip` exposes the level through `gzip.NewWriterLevel(w io.Writer, level int) (*gzip.Writer, error)`. Valid values are `gzip.HuffmanOnly` (-2), `gzip.DefaultCompression` (-1), `gzip.NoCompression` (0), and 1 through 9, where `gzip.BestSpeed` is 1 and `gzip.BestCompression` is 9. Anything outside that range returns a `nil` writer and an error — the constructor does not clamp, which is a small mercy, because it means a level read from configuration is validated at startup rather than silently rounded. `gzip.NewWriter(w)` is exactly `NewWriterLevel(w, gzip.DefaultCompression)` with the error dropped, and the default behaves like level 6. ## Why the default is a decision, not an absence of one Every agent that calls `gzip.NewWriter` has chosen level 6; it has just chosen it without measuring. For a shipper that is often the wrong end of the curve. The characteristic shape on log text is that levels 1 to about 5 climb steeply in ratio for modest CPU, and levels 7 to 9 pay several times the CPU for single-digit percentage gains, because the compressor spends far longer searching for matches in its window. So the method is: take a real sample — a full rotation from a busy node, not a synthetic file of repeated lines, which flatters every level — and write a benchmark that runs the agent's actual encode path at each candidate level. Record output bytes and nanoseconds per operation, and derive CPU-seconds per gigabyte of input, which is the unit a capacity conversation needs. `-benchmem` also tells you what the writers are costing in allocation. ## The two budgets, and why they have different owners This is the part that makes it a judgment call rather than a tuning exercise. The CPU is spent on the node where the workload runs, next to the service that produced the logs, and it competes with that service for cores. The saving lands as fewer bytes stored and fewer bytes moved — a line on an infrastructure bill that a different person owns. So the service owner sets the level and the budget holder can overrule it, in both directions. A finance-driven push to level 9 is legitimate if the CPU headroom exists; a platform push down to level 1 is legitimate if the shipper is measurably stealing latency from the request path. What is not legitimate is either side deciding unilaterally with no shared number, which is why the deliverable is a small table of level against bytes and CPU rather than an opinion. Before reaching for a higher level, check whether the input deserves compression at all. Data that is already compressed — archives, images, anything the producer gzipped once already — gains nothing and costs full CPU, and the right level for it is `gzip.NoCompression` or no compression layer at all. Equally, if the destination re-compresses on ingest or the storage tier is already compressed, spending level 9 to save bytes that get re-encoded downstream is pure waste. ## Making the choice operable - **Configuration, not a constant.** The number belongs in the agent's config with a measured default. At 3 a.m., when a shipper is pinning a core on a node whose service is timing out, dropping to `gzip.BestSpeed` must be a config change, not a build. - **Validate on load.** Because `NewWriterLevel` returns an error for an out-of-range value, construct one writer at startup so a bad config fails immediately with a clear message rather than on the first rotation. - **State the constraint you are engineering to.** 'The shipper stays under a defined fraction of one core at peak rotation volume' is a testable posture; 'use a sensible level' is not. - **Amortise the writer.** A `gzip.Writer` carries a compression window and Huffman tables — a substantial allocation per instance. Keep them in a `sync.Pool` and call `Reset` on each use so a per-file writer allocation does not dominate the CPU you just spent choosing a level. - **Instrument the ratio.** Emit compressed and uncompressed byte counts per upload. When the ratio moves, either the log format changed or something upstream is already compressing, and both are worth knowing before the bill arrives. ## What a strong answer sounds like Name the constants and the fact that the constructor validates rather than clamps; describe measuring on real data with two numbers; place the CPU cost and the byte saving on different owners' budgets; make the level configurable and the posture explicit; and mention the cases where the correct level is none at all.

  • What does gzip.NewWriterLevel do with a level of 12?
    It returns a `nil` writer and a non-nil error naming the invalid level, rather than clamping to 9. That makes it a useful validation point: construct one writer while loading configuration and the process fails at startup with a clear message instead of at the first rotation, hours later, on a node nobody is watching.
  • How would you argue against a request to move the agent from level 1 to level 9?
    With the measurement, not an opinion: show bytes saved per gigabyte against CPU-seconds added per gigabyte at current volume, convert both to the units each owner cares about, and state the latency headroom the service has left. If the ratio gain is a few percent and the CPU triples, the request is cheap to decline; if the node is idle and egress dominates, it is cheap to accept.
  • When is the right compression level none at all?
    When the input is already compressed, when the destination re-compresses on ingest, or when the storage tier compresses transparently — in all three cases you pay full CPU for a ratio near one. It is also the answer when the shipper is CPU-starved and the network is not, since the point of the layer is to trade the scarcer resource for the cheaper one.

It is like choosing a courier's packing density: whoever loads the van pays in time, whoever books the freight saves on space, and neither can decide alone without a shared measurement.

saying these in an interview costs you the question

  • Picks level 9 because smaller output is obviously better
  • Quotes a ratio without any measurement of its own data
  • Hardcodes the level so changing it needs a release
  • Ignores that the CPU and the byte saving hit different budgets
  • Compresses data that is already compressed