In a log collection pipeline, which records do you discard or sample before ingest, and what makes that irreversible?
answer
- Rank by volume, then by who reads it
- Prefer the weakest lever that works
- Key sampling on the request identifier
- Count what you discard, per rule
- A sampled count is not a count
basics
~20 sDiscard or sample the classes with high volume and no demonstrated readers: repetitive health-endpoint access lines, duplicated stack traces, oversized fields. It is irreversible because those records never reach storage, so version the rules and count what you discard.
solid answer
~50 sChoose the class from data, not intuition: rank producers by bytes per day, then cross that against what any human search or alert rule has touched in the last quarter. The reliable candidates are repetitive success records a metric already summarises, readiness-endpoint access lines, chatty third-party client output, and oversized fields inside otherwise useful records. Then apply the weakest lever that solves it — truncate a field, pre-aggregate a class into a counter at the edge, rate-limit a repeating record, sample, and only discard outright when nothing reads it. Key any sampling decision on the request or trace identifier so a kept request keeps all of its lines. It is irreversible because the record never exists downstream, so keep the rules in version control with dates, emit a per-rule discard counter, and carry the sampling rate with the survivors so nobody reads a sampled count as a true count.
code
pseudocode · 13 linesif record.source == "readiness-access":
discard(rule="probe-access") # increments a per-rule discard counter
if size(record.response_headers) > 2048:
record.response_headers = truncate(record.response_headers, 2048)
if record.severity == "ERROR" or record.duration_ms > 850:
keep(record) # never eligible for sampling
else if hash(record.request_id) % 40 != 0:
discard(rule="success-sample")
else:
record.sample_rate = 40 # travels with the survivor
keep(record)go deeper
Know that not every log record is worth storing and that a pipeline can discard records before they are ever searchable. Be able to name an obvious candidate, such as access lines for a health-check endpoint.
Explain the ladder from truncating a field, through pre-aggregating and rate-limiting, to sampling and discarding outright, and say why a sampling decision must be made per request rather than per line.
Show that you rank streams by volume and by demonstrated readers before cutting, that every rule emits a discard counter, and that rules are versioned and dated so an absent record is never ambiguous.
Own the policy: who may add a discard rule, which classes are never eligible, and how the platform proves what it is throwing away. Be ready to argue for a short unindexed archive as the price of making a destructive default reversible.
Discarding and sampling at collection is the only class of lever that reduces every part of a logging bill at once, because it acts before the records are accepted, processed, stored or scanned. It is also the only lever that destroys data, so it deserves the discipline that implies: chosen from measurement, applied at the weakest strength that works, and made visible. ## Choosing the class from data, not intuition Rank producers by bytes per day, then rank the same producers by demonstrated readers — streams touched by a human search or named by an alert rule in the last quarter. The interesting cell is high volume with no readers. | Class | Why it is large | Weakest sufficient lever | What is lost | |---|---|---|---| | Access lines for readiness and health endpoints | one per probe, per replica, per interval | discard at the collection tier | nothing a probe-success metric does not carry | | Successful high-volume request lines | one or more per request | sample, keyed on the request identifier | per-request detail on successes only | | A repeating stack trace from one live bug | one per occurrence, unbounded | rate-limit per fingerprint, keep a count | repeat detail beyond the first few | | High-frequency third-party client output | chatty by default, not tunable per call | discard by logger source | third-party internals while debugging | | Oversized fields (bodies, header blocks) | kilobytes inside an otherwise useful record | truncate the field, keep the record | the tail of one field | ## The ladder of loss Take the highest rung that solves the problem: 1. **Truncate** a field — the record survives intact apart from one tail. 2. **Pre-aggregate** — replace a class of line with a counter or a latency distribution at the edge, keeping the shape and losing the individual records. 3. **Rate-limit per fingerprint** — keep the first few occurrences of a repeating record per interval, plus a count of the rest. 4. **Sample** — keep a representative, complete subset. 5. **Discard** — keep nothing but a counter. A team that starts at rung five because it is the simplest thing to configure usually discovers afterwards that rung one would have removed most of the volume. ## Sampling that does not shred a request Sampling each line independently produces a corpus in which no single request is complete: you have the third and the seventh line of a failure and nothing between them, which is worse than not having the request at all, because it still looks like data. Key the decision on a request or trace identifier — hash it into buckets — so every record belonging to one request is kept or discarded together, and so independent collectors reach the same verdict with no coordination between them. Two rules always travel with it: - **Exempt the interesting classes.** Records at error severity, requests above a latency threshold, and anything belonging to a stream under a retention obligation are kept at full rate regardless of the sampling rule. - **Record the rate.** A sampled count is not a count. If one success in forty is kept, every query that counts records must multiply by forty, which means the rate has to travel with the surviving records or with a companion counter of discards. Otherwise a graph silently understates by a factor of forty, and changing the rate moves a business-looking line for no business reason. ## What makes it reversible, and what does not The decision is destructive the moment it is made. What you can control is how recoverable a mistake is. - **Push the rule as late as you can while still avoiding the charge.** If the cost is levied on bytes accepted at the front door, the rule must run before hand-off — but a central collection tier reconfigured in one place beats per-host shipper configuration rolled across a fleet. - **Count every discard, per rule.** A pipeline that discards silently is one where a broken shipper and a working discard rule look identical, and a service's logs can vanish for a week before anyone notices. - **Keep the rules in version control, dated.** Without that, an absent record is ambiguous: did the service not log it, or did the pipeline eat it? "What were we discarding on the fourteenth" has to be answerable. - **Consider a short unindexed archive.** Routing discarded classes to compressed raw blobs for five to seven days, with no processing and no search structures, converts an irreversible discard into a reversible one for exactly the period in which you would notice the mistake, at a small fraction of the indexed cost. It is a second store to operate, under the same deletion rules. - **Classify before you cut.** Records you are obliged to retain, and records that are the only evidence a control executed, are never eligible. ## A worked example A vinyl-record marketplace ships 18.4 GB a day. Ranking by producer shows 7.5 GB is readiness-endpoint access lines from 109 replicas polled twice a second, and a further 3.2 GB is one HTTP client writing a full response header block per call. Neither has been queried in ninety days. Discarding the first and truncating the second's header field removes about 10 GB a day from ingest, processing, storage and query simultaneously — shipped with a per-rule discard counter and a five-day unindexed archive of the discarded class while the team builds confidence.
- Why sample on a hash of the request identifier rather than deciding per line?Per-line sampling leaves no request complete: you get the third and seventh lines of a failure and nothing between them, which reads like data but explains nothing. Hashing the identifier makes the decision identical for every record of one request and stateless, so independent collectors agree without coordinating.
- A dashboard counts log records from a stream sampled one in forty. What has to change?The count must be scaled by the sampling rate, which means the rate has to travel with the surviving records or with a companion counter of discards. Otherwise the graph understates by forty times, and a later change to the rate moves a business-looking line for no business reason.
- Which records would you refuse to sample, whatever their volume?Anything under a retention obligation, anything that is the only evidence a control actually executed, and records at error severity. Low-volume streams are also poor candidates: sampling them saves nothing measurable while destroying the completeness that makes them useful during an investigation.
saying these in an interview costs you the question
- Samples every log line independently, leaving no request complete
- Discards a stream without checking whether an alert rule queries it
- Reads a count from sampled records as if it were the true count
- Keeps discard rules only in live configuration, undated and unversioned
- Discards silently, so a broken shipper looks identical to a working rule
- Applies sampling to records the organisation is obliged to retain