skip to content

You must choose a trace sampling strategy for a fleet of services that call one another. How do you keep the decision consistent across a whole trace, how do you decide between deciding at the entry point and deciding later in a collector, and how do you set the rate?

level: principalimportance: should knowfreq 35%

answer

  1. one decider at the root, everyone inherits
  2. consistency requires identical algorithm and rate at all roots
  3. head = cheap, blind to outcome; post-hoc = outcome-aware, must buffer
  4. rate from budget, plus per-route bias
  5. dropped at head can never be recovered later

basics

~20 s

Decide once at the entry point and have everyone else inherit, using parent-based sampling everywhere. Head sampling is cheap and blind to outcomes; deciding after the trace completes, in a buffering collector, can keep errors and slow traces but costs buffering and complexity. Rate follows an ingestion budget, not a round number.

solid answer

~60 s

Two rules first. **One decider**: the entry point makes the call, every other service runs a parent-based sampler and inherits, so traces are never half-collected. **Uniform configuration**: if root decisions are taken in more than one place — several ingress services, or asynchronous entry points — they must use the same trace-id-derived algorithm and rate, or the fleet disagrees about the same trace. Then the placement question. **Head** sampling in the SDK decides at the first span, before anything is known: cheapest by far, since dropped data never leaves the process, but it cannot favour errors or slow requests. **Post-hoc** sampling in a collector sees the assembled trace and can keep every error and outlier, at the price of receiving and buffering everything, routing all spans of a trace to the same instance, and holding a decision window. Rate comes from a budget: ingestion allowance divided by traffic, biased up on low-volume critical paths and down on health checks. Keep 100% in the environments where volume is small enough not to matter.

go deeper

for a junior

Not expected to own this; know that one service should decide and the rest inherit, and that sampling trades coverage for cost.

for a middle

Explain parent-based inheritance, trace-id determinism for agreement, and why head sampling cannot favour errors.

for a senior

Compare head and post-hoc placement on cost, buffering and routing constraints, and derive a rate from a budget with per-route bias.

for a principal

Own it as fleet policy: named root deciders with identical configuration, a defended budget, rules centrally governed, effective rate published as telemetry, and span-derived metrics insulated from rate changes.

## The two questions that actually matter Sampling strategy reduces to *where* the decision is made and *what it is allowed to know when it is made*. Everything else is arithmetic. ## Consistency is a correctness property A trace sampled by half its services is worse than one not sampled at all: it looks complete, and the gap is indistinguishable from a real fault. The mechanism that prevents it is parent-based sampling everywhere, so the decision travels with the context and only the root decides. That leaves the roots. A fleet usually has several: each ingress service, plus every asynchronous entry point — a scheduled job, a consumer of an external queue. All of them must use the same rate and the same trace-id-derived algorithm; deterministic derivation is what makes independent processes agree without coordination. Two ingress tiers with different rates yield partial traces whenever a request crosses between them. A related nuance: because the derivation depends on the trace id, sampling is uniform over traces, not over customers, endpoints or error rates. Everything you care about that is rare is rare in the sample too. ## Head versus post-hoc **Head** (in the SDK, at the first span): the decision is made before the outcome exists. Its advantage is economic and structural — dropped spans are never serialised, never sent, never received, so cost falls at the source and no component needs to buffer. Its limit is absolute: you cannot ask for 'all traces with an error' because at decision time there is no error. **Post-hoc** (in a buffering collector, after the trace is assembled): the decision sees the finished trace and can keep every error, every latency outlier, and a small sample of the boring remainder. That is exactly the population humans want. It costs: every span must be transmitted and received even though most are discarded, the buffering layer must hold traces for a decision window (so late spans past that window are missed), and all spans of a trace must reach the same decision instance, which constrains how you scale and route the collection tier. Configuring that tier belongs to the collector layer; the strategic choice is yours. A common composition is a coarse head sample to cut the obvious bulk, plus a post-hoc layer for the surviving population. Reason about it as a multiplication — the head rate bounds what any later stage can possibly keep — and remember that once head sampling has dropped a trace, no downstream cleverness can bring it back. ## Setting the rate Work from a budget, not a habit: 1. Take the ingestion or storage allowance you can defend and divide by traffic to get a baseline rate. 2. Bias per route. Health checks and readiness probes deserve zero. A low-traffic, high-value endpoint can stay at 100% and cost nothing. A hot endpoint at 1% still yields plenty of examples per minute. 3. Check absolute counts, not just percentages: a rate is only adequate if it produces enough traces per minute on your *rarest important* path to be useful during an incident. This is where uniform ratios fail and per-route rules earn their complexity. 4. Keep development and staging at 100% — volume is trivial and the debugging value is highest. ## Custom root samplers Route- or attribute-aware decisions require a custom sampler at the root, using only start-time inputs. Two constraints: it must be fast, since it runs on every span start, and it should stay deterministic in the trace id so it remains consistent if more than one service ever roots a trace. Keep the rules few and centrally owned, because sampling rules are effectively a query the fleet has committed to answering, and per-team drift makes fleet-wide comparisons meaningless. ## What sampling must not be used for Not a privacy control — kept traces keep everything. Not back-pressure — that is the processors' and exporters' job. Not a substitute for instrumentation coverage — a broken hop is missing at 100% sampling too. And not something to change quietly: a rate change silently rewrites the denominator of anything computed from traces, so any span-derived metric must be produced before sampling or explicitly scaled. ## The observable you owe yourself Whatever you choose, publish the effective rate as telemetry, and monitor traces kept per minute on the paths you care about. Sampling policy drifts, gets overridden per service, and is discovered to be wrong exactly when you are trying to use it.

  • Two ingress services both root traces, one at 10% and one at 1%. What goes wrong?
    Requests that enter through one tier and cross into the other produce inconsistent decisions, so traces come back partial — spans present on one side of the hop and absent on the other, which reads as a fault rather than as configuration. Root decisions must share one rate and one deterministic algorithm; vary the rate by route through a rule the roots all apply, not by service.
  • Someone computes request rate from span counts and the sampling rate is halved. What breaks?
    Every span-derived quantity silently halves, and the change looks like a traffic drop. Metrics should be produced from instrumentation before the sampling decision, or from a component that scales by the known rate. Otherwise sampling policy becomes an invisible input to your dashboards, and any future rate change becomes an incident.

Head sampling is deciding which letters to keep before opening them; post-hoc sampling is opening every letter and filing the interesting ones. The second finds far more, and you pay the postage on all of it.

saying these in an interview costs you the question

  • Setting different sampling rates per service instead of deciding once at the root
  • Expecting head sampling to preferentially keep errors or slow traces
  • Assuming a post-hoc decision layer removes the cost of transmitting the data it discards
  • Picking a round percentage without checking traces-per-minute on the rarest important path
  • Changing the rate without accounting for anything computed from span counts

context