If sampling policies keep only errors and slow traces, how do you still answer whether a request's behaviour is normal?
answer
- Only-the-bad has no denominator
- Keep a random baseline stratum too
- Record why each trace was kept
- Weight only known probabilities
- Counting questions belong to metrics
basics
~20 sKeep a small random baseline alongside the policy-driven keeps, and record why each trace was kept so consumers know which ones can be weighted into an estimate. Rates and distributions should come from unsampled metrics, not from the trace store.
solid answer
~50 sA corpus selected for being bad has no denominator: percentiles over it mean nothing, dependency maps over-represent retry and fallback paths, and nobody can say what normal looks like because no normal trace was kept. The reconciliation is **stratified keeping**. Run a small uniform random baseline stratum regardless of outcome, alongside the policy strata for errors, slow requests and rare attributes, and give each stratum its own budget so an error storm cannot evict the baseline. Record on every kept trace why it was kept and with what probability, so a consumer can weight the baseline by inverse probability into an unbiased estimate and count the policy strata separately rather than folding them in. The deeper move is to stop asking the trace store to count at all: derive rates, error ratios and latency distributions from a signal computed over every request before sampling. Metrics say how often; traces show what one looked like.
code
pseudocode · 10 lineschoose(trace):
if trace.has_error: return keep(reason = "error", probability = unknown)
if trace.duration > SLOW: return keep(reason = "slow", probability = unknown)
if trace.debug_forced: return keep(reason = "forced", probability = unknown)
if hash(trace.id) % 800 == 0: return keep(reason = "baseline", probability = 1/800)
return drop()
# aggregating over kept traces:
# population estimates -> baseline only, each weighted by 1 / probability
# incident counts -> policy strata, counted, never weightedgo deeper
Understand that if a system only stores the requests that went wrong, it cannot show what a healthy request looks like, so there is nothing to compare a suspicious trace against.
Be able to explain why an average or a percentile computed over deliberately-selected slow traces is meaningless, and why a small random slice kept regardless of outcome restores the comparison.
Show how you would operate it: separate budgets per stratum, provenance recorded on each kept trace, alerting on realised keep rates, and pushing rate and distribution questions onto metrics computed over all requests.
Own the whole allocation. Decide which strata exist and what each is worth, who may add a policy, how bias is documented so nobody builds analysis on an incident corpus, and where you accept that rare failures stay invisible.
## A trace store answers two different kinds of question Traces get asked **forensic** questions ("show me the request that failed") and **analytic** questions ("what does a normal booking look like, which dependency dominates, has the fan-out changed"). Outcome-driven sampling policies optimise the first and destroy the second, because a corpus of only-the-bad has no denominator. What breaks, concretely: - a latency percentile computed over a set selected for being slow is not a percentile of anything; - a service dependency map built from failure traces over-represents retries, fallbacks and error paths that normal traffic never takes; - "this looks wrong" has nothing to compare against, because nobody can show what right looks like; - engineers calibrate their intuition on a corpus that is one hundred percent incidents. The ferry-timetable booking platform ran exactly this policy: keep every server error, keep everything over 1.8 seconds, keep nothing else. During an incident nobody could explain from the existing dashboards, the store held 61,400 slow traces and not one fast one, so no one could say whether the sixth downstream call visible in them was new or had always been on the path. ## Strata, budgets and recorded provenance The standard reconciliation is to stop thinking about "the sample" and start thinking about **strata**, each with its own budget: 1. a **baseline stratum** — a small uniform random ratio kept regardless of outcome, and the only part that can be reasoned about statistically; 2. **policy strata** — errors, duration over a threshold, a named tenant, a rare attribute combination; 3. a **forced stratum** — requests explicitly flagged for debugging. Three rules turn strata from a nice idea into something usable: - **Record on each kept trace why it was kept and with what probability.** A consumer has to be able to tell a one-in-eight-hundred baseline trace from a deterministic error keep. Without that field the corpus is unweightable and every aggregate over it is wrong. - **Weight only what has a known probability.** Inverse-probability weighting turns the baseline stratum back into an unbiased estimate of the population: a trace kept at one in eight hundred stands for eight hundred requests. A trace kept because it errored was selected with probability one given the error; folding it into the same estimate inflates the failure rate. - **Budget each stratum separately.** When policies share one pool, a latency regression that trips the slow-request policy fleet-wide consumes the whole allowance and silently evicts the baseline, so the comparison set disappears exactly when it is needed. The price is small. At 2,560 requests a second, a one-in-eight-hundred baseline is about 3.2 traces a second, under two percent of that platform's trace bill. ## Better still: stop asking the trace store to count The deeper move is to route the counting questions somewhere else. Rates, error ratios, throughput and latency distributions should be computed from a signal derived from every request before any sampling decision is taken, then aggregated. The division of labour becomes clean: | Question | Answerable from a policy-only corpus? | Where it should come from | | --- | --- | --- | | Show me a failed booking | yes | policy stratum | | What is p99 checkout latency? | no | metrics over all requests | | What does a normal fan-out look like? | no | baseline stratum | | Is this dependency new? | misleading | baseline stratum | | How often does this failure happen? | no | metrics over all requests | Metrics say how often and how bad; traces say what one of them looked like. Teams that make this split stop needing the trace corpus to be representative for most questions, which is fortunate, because for the rare-failure questions it never will be. ## The part that is organisational - Policy lists accrete and nobody prunes them. Each policy is a standing claim on the budget and deserves review like any other cost line. - Alert on each stratum's realised keep rate. A policy moving from rare to common is a signal in its own right, and it is also the failure mode that eats the budget. - Publish which strata exist, so an engineer reading something built on traces knows whether the number in front of them is an estimate or a count of incidents. - Accept the honest limit. For a failure that happens a handful of times a day, no affordable baseline will contain an example — which is why the policy strata exist. The answer is both, with provenance recorded, not a choice between them.
- How do you turn a sampled trace corpus back into an unbiased estimate?By recording the probability with which each trace was selected and weighting it by the inverse when aggregating: a trace kept at one in eight hundred stands for eight hundred requests. That only works for traces chosen by a known random rule. A trace kept because it errored was selected with probability one given the error, so folding it into the same estimate inflates the failure rate. Weight the baseline stratum; count the policy strata separately.
- A policy keeping every request over a latency threshold suddenly consumes the whole trace budget. What do you do?Give each stratum its own rate limit instead of letting policies share one pool, so a latency regression cannot evict the baseline and the error stratum. Then alert on the realised keep rate per stratum: a policy moving from rare to common is a signal worth having in its own right. Widening the threshold comes second, once you know which behaviour actually changed.
- Is there a case for accepting a biased corpus and not fixing it?Yes, when the trace store is explicitly the incident tool and every counting question is already answered by metrics computed over all requests. What makes that safe is honesty: publish that the corpus is outcome-selected, so nobody builds a dashboard on it. The failure is not bias, it is undocumented bias that people quietly treat as a sample of production.
A hospital that files only the charts of patients who died can describe every death in detail and still cannot say what recovery looks like.
saying these in an interview costs you the question
- Computes latency percentiles from policy-kept traces
- Thinks keeping every error is a complete strategy
- Folds deterministic keeps into weighted population estimates
- Lets all policies share one undifferentiated budget
- Builds dependency maps from an error-only corpus