One logger is emitting thousands of lines a second in production. How do you cut that volume without hiding the failure?
answer
- Noise costs more than storage does
- Cap it at the call site
- Never drop anything silently
- Emit a count of suppressed lines
- Sample whole requests, not single lines
basics
~20 sRate-limit or deduplicate at the call site rather than deleting the line: keep the first occurrence, count the rest, and emit a summary saying how many were suppressed. A silent drop is worse than the noise it removed.
solid answer
~40 sCut the volume at the call site and make the cut visible. Cap a logger or a message template at a few lines per window, or deduplicate by exception fingerprint, and when the window closes emit a line saying how many occurrences were suppressed and from where — plus a counter in the metrics stream so the loss shows up on a dashboard. A silent drop is indistinguishable from the system going quiet, which is exactly the ambiguity you cannot afford mid-incident. If you must sample rather than cap, decide once per request and keep every line for the selected fraction; an independent per-line decision shreds each request into fragments. And check severity first: a WARN on a normal path is a defect to fix, not volume to manage.
code
pseudocode · 13 linesonLog(template, args):
window := limiter.windowFor(template)
if window.emitted < 5:
window.emitted += 1
write(template, args)
else:
window.suppressed += 1
onWindowClose(window):
if window.suppressed > 0:
write(WARN, "suppressed {} further lines from {} in the last {}s",
window.suppressed, window.template, window.durationSeconds)
counter("log_lines_suppressed_total", window.suppressed, template = window.template)go deeper
Know that log volume is not free and that a line written on every iteration of a hot loop is a problem. Be able to say why you would look for the loudest logger before asking anyone for more storage.
Explain the mechanics of suppression: a cap per message template per window, a count of what was dropped, and a summary line so the gap is visible. Know that argument formatting is paid before the write, not after it.
Show that you have made this call under pressure: which logger you capped, what you preserved, how you proved the drop was safe, and how you argued the case for or against leaving verbose detail on in production afterwards.
Own the fleet-wide position. Decide what the platform guarantees about log throughput, who pays when one team's verbosity crowds out everyone else's telemetry, and whether verbose detail is a per-service choice or a platform-granted budget.
## What the volume actually costs A logger emitting thousands of lines a second is usually described as a cost problem, and the invoice is the least interesting part of it. On a cheese-ageing inventory platform, one humidity-reconciliation service logged a WARN per sensor read across 47 vaults; at 384 sensors each and a two-second reconcile loop, that came to roughly 18,400 lines a second. Three things happened long before anyone looked at a bill: - **The request path got slower.** Each line costs an argument-formatting pass and a handful of allocations. When the appender's bounded queue filled, the implementation blocked the calling thread rather than dropping, and p99 latency moved from 41 ms to 610 ms. - **The collection tier lost data.** That service's share of ingest jumped from 4.6 GB/day to 63 GB/day, outran what the shared collection tier would accept, and lines were shed — including lines belonging to the other forty services using it. - **Nobody could read anything.** The incident nobody could explain from the existing dashboards was in the logs the whole time, three thousand identical WARN lines away from the one that mattered. The ordering matters because it decides what you are optimising. Cutting volume is not primarily an economy measure; it is how you keep logs usable and stop logging from becoming the failure itself. ## Suppress, but report the suppression One rule separates a fix from a cover-up: **a drop has to leave a trace.** A logger that quietly stops emitting is indistinguishable, on an incident timeline, from a system that stopped doing the thing — and that ambiguity is most expensive at exactly the worst moment. What a suppression mechanism should preserve: 1. **The first occurrence, in full.** The first line carries the detail; the thousandth carries none. 2. **A count of what was dropped, and over what window**, emitted as a line of its own when the window closes: "suppressed 17,326 further lines from this logger in the last 10s". 3. **The identity of what was collapsed** — the logger name and, better, the message template or an exception fingerprint — so the count is attributable to something. 4. **A counter in the metrics stream**, labelled by logger or template, so suppression is visible on a dashboard rather than only inside the logs it is thinning. A spike in suppressions is itself worth noticing. 5. **Independent budgets per template.** If the cap is keyed on the message template, a genuinely different failure is a different template and is not silenced by the noisy one's allowance. ## Choosing the axis you cut on | Approach | Keeps | Loses | Use when | |---|---|---|---| | Cap per message template per window | First occurrences of every distinct message | Repeat detail inside a burst | One or a few loggers are loud | | Deduplicate by exception fingerprint | One example per distinct failure, with a count | Per-occurrence context | The same stack repeats endlessly | | Sample by request identity | Complete narratives for a known share of traffic | Everything about unselected requests | Volume is spread evenly across the service | | Independent per-line coin flip | A fixed fraction of lines | Coherence: every request becomes fragments | Almost never | The last row is the mistake worth naming out loud. Deciding per line looks equivalent to deciding per request and is not: you keep the third line of one request and the seventh of another and can reconstruct neither. Hashing a request or trace identifier and keeping every line for the selected fraction costs the same and leaves you with stories rather than fragments — and it lines up with whatever share of traces was retained, so the two correlate instead of missing each other. Two things suppression must not become. It is not a substitute for fixing severity: a WARN emitted on a path that is normal and expected is a defect in the code, and the fix is to lower or delete the line rather than to build machinery that tolerates it. And a cap keyed on severity alone is dangerous, because the first real ERROR then arrives inside the same budget as the noise it is competing with. ## The honest argument about verbose levels in production **For.** Some incidents are not reproducible. Detail already being captured is the only detail you get, and turning DEBUG on afterwards captures the next occurrence rather than the one under investigation. For a small deterministic slice — one replica, or a hashed fraction of requests — the cost is bounded and known in advance, and the payoff on the rare unexplainable incident is large. **Against.** That cost is paid on every request, by every instance, permanently, for a benefit that materialises a handful of times a year. It concentrates risk in the logging path, where a full queue converts a degradation into an outage. And the volume erodes the very thing it was bought for: a reader who has to skim fifty times as much material is slower, not better informed. The defensible position is neither extreme. Run INFO by default, make verbose reachable per logger at runtime, and keep a small number of deliberately-chosen high-value lines permanently on the paths that history says you always end up wanting. Choose that list from real incidents rather than from whatever was easiest to add.
- Why is an independent per-line sampling decision worse than sampling by request identity?Because a coin flip per line shreds every request's story: you keep line three of one request and line seven of another and can reconstruct neither. Hashing a request or trace identifier and keeping every line for the selected fraction gives complete narratives for a known share of traffic at the same volume, and it lines up with whatever share of traces was retained so the two correlate.
- How would you tell, after the fact, that suppression hid something you needed?By making suppression measurable. Every suppressed line increments a counter labelled by logger or message template, so what you dropped sits on a dashboard beside what you kept. If an incident timeline has a gap, the counter says whether the logger went quiet or the suppressor did — and a spike in suppressions is itself a signal worth watching.
- When is deleting the log line outright the right answer instead of rate-limiting it?When the line has no reader. A WARN emitted on a path that is normal and expected is a defect in the code, and the fix is to remove it or lower it to DEBUG. Rate-limiting is for lines that genuinely matter but arrive in bursts; it must not become a way to live comfortably with severity that was wrong in the first place.
Good rate-limiting is a smoke alarm that stops shrieking but keeps reporting "still smoking, 4,000 times" — quieter without pretending the fire went out.
saying these in an interview costs you the question
- Drops noisy lines silently with no record of the loss
- Thinks the only cost of verbose logging is storage
- Samples each line independently and expects coherent requests
- Caps by severity, so the first real failure is dropped too
- Says DEBUG everywhere is fine because disks are cheap
- Ignores that a blocking appender pushes cost onto request threads