skip to content

Explain how tail-based sampling works in the OpenTelemetry Collector — the decision buffer, policies, and late-arriving spans — and what it costs compared with deciding at span creation time.

level: seniorimportance: must knowfreq 44%

answer

  1. Buffer by trace id, wait for idle, then decide once for the whole trace
  2. decision_wait ≈ 30s; num_traces caps the buffer, eviction is silent
  3. Policies OR together: errors + slow + a probabilistic baseline
  4. All spans of a trace must hit one instance → routing_key: traceID
  5. Derive RED metrics BEFORE the sampler

basics

~20 s

The tail_sampling processor buffers spans by trace id for a wait window, then evaluates policies (latency, error status, attribute match, probabilistic) against the assembled trace and keeps or drops the whole trace. It costs memory and requires trace affinity.

solid answer

~50 s

Head sampling decides when the root span starts, before anything interesting has happened, so it is cheap but blind — you cannot say "keep every trace that errored". Tail sampling defers the decision. The `tail_sampling` processor buffers spans in memory keyed by trace id; when a trace has been idle for `decision_wait` (30s by default) it evaluates the configured policies and emits or drops **the whole trace**. Policies include `status_code`, `latency`, `string_attribute`/`numeric_attribute`, `probabilistic`, `rate_limiting`, `span_count` and OTTL conditions, combined with `and`/`composite`; typically you keep all errors, all slow traces, and a small percentage of the rest. The costs are real: memory proportional to traces in flight (`num_traces` caps the buffer), a `decision_wait` delay before anything reaches the backend, and the hard requirement that every span of a trace reach the same Collector instance. Spans arriving after the decision are handled by the late-span path and may be dropped, so long-running traces need care. Derive RED metrics with a connector **before** the sampler so aggregates stay accurate.

code

text · 11 lines
text
processors:
  tail_sampling:
    decision_wait: 30s
    num_traces: 50000
    policies:
      - { name: errors,   type: status_code,  status_code: {status_codes: [ERROR]} }
      - { name: slow,     type: latency,      latency: {threshold_ms: 2000} }
      - { name: baseline, type: probabilistic, probabilistic: {sampling_percentage: 1} }

pipeline order:  otlp → k8sattributes → [spanmetrics connector] → tail_sampling → batch → otlp/vendor
                                          ^ metrics computed on 100% of traffic

go deeper

for a junior

Know the distinction: head samples at trace start and cannot see outcomes; tail waits and can keep errors and slow traces.

for a middle

Describe the buffer, decision_wait, the common policy types, and that the verdict covers the whole trace.

for a senior

Own the operational cost model — memory by concurrent traces, delayed visibility, silent buffer eviction, trace-affinity routing — and put metric derivation before the sampler.

for a principal

Decide policy strategy and budget: what fraction of spend goes to errors versus baseline coverage, whether an always-on slice is needed for live views, and how the sampling tier's scaling events are constrained to keep decisions stable.

## Why the decision is deferred Sampling exists because storing every span of every request is unaffordable at scale. Head sampling makes the decision at trace start — a ratio applied to the trace id, propagated so the whole trace agrees. It is cheap, stateless and consistent, but the decision is made before the request has done anything: you cannot condition on "it failed", "it took four seconds", or "it touched the payments path", because none of that is known yet. Tail sampling exists to condition on exactly those things. ## The mechanism The `tail_sampling` processor keeps an in-memory map from trace id to the spans seen so far. A trace's clock starts on its first span. When no new span has arrived for `decision_wait` (default 30 seconds), the processor evaluates every configured policy against the assembled trace and produces a single verdict for the entire trace — sampling is all-or-nothing per trace, which is what makes the retained traces usable rather than full of holes. Capacity is bounded by `num_traces`, the maximum number of traces held simultaneously (default in the tens of thousands); `expected_new_traces_per_sec` pre-sizes internal structures. When the buffer is full, the oldest traces are evicted — which shows up as unexplained missing traces, not as an error, so it must be watched. ## Policies Each policy has a name, a type and its own parameters. The common set: - `status_code` — keep traces containing a span with `ERROR`. - `latency` — keep traces whose duration exceeds a threshold. - `string_attribute` / `numeric_attribute` / `boolean_attribute` — keep by attribute value, e.g. a specific tenant, endpoint or feature flag. - `probabilistic` — keep a percentage, the baseline for representative traffic. - `rate_limiting` — cap spans per second regardless of other policies, a safety valve. - `span_count` — keep unusually large traces (a fan-out gone wrong). - `ottl_condition` — arbitrary expressions over span or resource fields. - `and` / `composite` — combine policies; `composite` also lets you allocate a rate budget across sub-policies in priority order. Policies are evaluated as an OR by default: if any policy says sample, the trace is kept. That is why the typical configuration reads "all errors, all slow, plus 1% of everything else" — the union produces both complete failure coverage and an unbiased baseline. ## The three costs **Memory.** You are holding every span of every in-flight trace for the wait window. Budget roughly traces-per-second × `decision_wait` × average bytes per trace, and remember that attribute-heavy spans move that average a lot. This is why the sampling tier is sized on concurrent traces, not on request rate. **Latency to the backend.** Nothing appears in the backend until the wait expires. Thirty seconds of delay is fine for post-hoc debugging and wrong for a live tail view, which is why some teams run a separate always-on pipeline for a small unsampled slice. **Affinity.** The processor can only decide from what it holds, so **every span of a trace must reach the same Collector process**. Behind a normal load balancer that is false, and the failure is silent: policies evaluate against fragments. The standard remedy is a stateless routing layer whose `loadbalancing` exporter is configured with `routing_key: traceID`, hashing spans onto a consistent instance of the sampling layer. ## Late spans A span that arrives after its trace's decision has been made cannot retroactively change it. Depending on configuration the processor either drops such spans or applies the recorded decision for a limited retention period. Long-running traces — a job whose root span lasts minutes, a streaming request, a workflow spanning retries — are the pathological case: the trace looks idle, the window expires, a decision is taken, and the remainder is orphaned. Extending `decision_wait` multiplies the memory cost, so long-lived work is usually modelled as several shorter traces joined by span links instead. ## Interaction with derived metrics If you compute request-rate, error-rate and duration metrics from spans, compute them **before** sampling. Metrics derived after a 1% sampler describe 1% of your traffic; unless every consumer knows and corrects for the ratio — and error-biased policies make the ratio non-uniform — your dashboards are simply wrong. Ordering the pipeline as receivers → enrichment → span-metrics connector → `tail_sampling` → exporters gives accurate aggregates over 100% of traffic while storing a fraction of the traces. ## Choosing between head and tail Head sampling is right when the volume reduction is the whole point and uniform sampling is acceptable: it costs nothing, adds no latency, and needs no shared state. Tail sampling is right when the rare events are the valuable ones. Many estates run both — a modest head ratio to cut raw volume at the source, then tail policies on what survives — but be explicit that head sampling has already thrown away errors you will never see at the tail.

  • Your tail-sampled traces are missing spans from one downstream service. What do you investigate?
    First, whether that service's spans reach the sampling instance holding the trace — if routing is not trace-id-affine, its spans land elsewhere and are decided separately. Second, timing: if the service reports late, its spans may arrive after `decision_wait` has expired and be discarded as late spans. Third, whether that service uses a different propagation format so its spans carry a different trace id entirely, in which case it is not a sampling problem at all.
  • How do you keep dashboards accurate once you sample at 1%?
    Compute the aggregates from unsampled data. Put a span-metrics connector (or equivalent aggregation) before the sampler so request counts, error counts and duration histograms cover 100% of traffic, and let sampling apply only to the stored exemplar traces. The alternative — scaling metrics by the sampling ratio — breaks as soon as policies are non-uniform, and error-biased policies always are.
  • When would you keep head sampling instead of moving to tail sampling?
    When you mainly need volume reduction and uniform coverage, when you cannot afford the memory or the trace-affinity routing tier, or when the added delay before traces become visible is unacceptable. Head sampling is stateless, free and consistent across services; tail sampling buys selectivity and charges memory, latency and topology complexity for it.

saying these in an interview costs you the question

  • Believing tail sampling can be done independently on each Collector replica behind a normal load balancer
  • Thinking the decision applies per span rather than per trace
  • Computing service-level metrics after the sampler and then wondering why traffic dropped 99%
  • Raising `decision_wait` to cover long-running traces without accounting for the memory multiplier
  • Assuming late spans retroactively flip a decision already made

context