Explain the built-in trace samplers in the OpenTelemetry SDK — AlwaysOn, AlwaysOff, TraceIdRatioBased and ParentBased. Which is the default, what information does a sampler see when it runs, and what do the OTEL_TRACES_SAMPLER and OTEL_TRACES_SAMPLER_ARG environment variables set?
answer
- sampler runs at span start, sees start-time inputs only
- ratio derived from trace id ⇒ services agree
- ParentBased delegates only at the root
- default = ParentBased(AlwaysOn)
- OTEL_TRACES_SAMPLER + _ARG
basics
~20 sAlwaysOn and AlwaysOff record all or nothing. TraceIdRatioBased keeps a fraction derived from the trace id, so every service decides alike. ParentBased follows the incoming decision and delegates only at the root; the default is ParentBased wrapping AlwaysOn. OTEL_TRACES_SAMPLER picks one, OTEL_TRACES_SAMPLER_ARG parameterises it.
solid answer
~50 sA sampler runs **at span start** and returns drop, record-only, or record-and-sample. Its inputs are the parent context, the trace id, the span name and kind, the links, and the attributes supplied at creation — nothing discovered later. - **AlwaysOn / AlwaysOff** — unconditional; fine for development or for a service you want silent. - **TraceIdRatioBased(p)** — computes the decision from the trace id itself. Because the trace id is shared, independent services running the same ratio reach the same verdict, which is what stops traces coming out half-sampled. - **ParentBased(delegate)** — if there is a parent, honour its sampled flag; only for a root span consult the delegate. This is the composition that makes a trace all-or-nothing. The SDK default is ParentBased(AlwaysOn): keep everything, respect upstream. `OTEL_TRACES_SAMPLER` selects by name (`always_on`, `always_off`, `traceidratio`, `parentbased_always_on`, `parentbased_traceidratio`, plus vendor options); `OTEL_TRACES_SAMPLER_ARG` supplies the parameter, e.g. the ratio.
go deeper
Name the four samplers with a one-line description each and state the default, ParentBased with AlwaysOn.
Explain the sampler's start-time inputs, why the trace-id derivation gives consistency, and what the two environment variables set.
Discuss the coverage cost, why errors cannot be favoured at head time, and when a custom sampler is justified.
Frame the fleet policy: who decides at the root, uniformity as a correctness requirement, and where a post-hoc decision layer belongs.
## What sampling decides, and when Head sampling is the SDK deciding, at the instant a span starts, whether it will be recorded and exported. The decision has three possible outcomes: - **DROP** — the span is not recorded; it still gets a valid span context so children and downstream services can be parented and the trace stays structurally coherent. - **RECORD_ONLY** — recorded in-process (visible to anything reading spans locally, e.g. deriving metrics from them) but not exported. - **RECORD_AND_SAMPLE** — recorded and passed to the processors for export. The sampler's inputs are fixed: parent context, trace id, name, kind, initial attributes, links. It cannot see duration, status, or anything that happens later — which is the structural reason head sampling can never mean 'keep all the errors'. ## The built-ins **AlwaysOn** samples everything. It is the right default while a system is small, and it is what the SDK's default wraps. **AlwaysOff** samples nothing. Useful to silence a service without removing instrumentation, keeping propagation intact. **TraceIdRatioBased(p)** derives the verdict deterministically from bits of the trace id compared against a threshold. Two properties matter. First, the *same trace* yields the same verdict everywhere, so services configured with the same ratio agree — sampling is consistent without any coordination. Second, it is uniform over trace ids, not over customers or endpoints: 1% of everything, including 1% of your rare endpoints. **ParentBased(root, …)** is a wrapper, not a decision. If the span has a parent, it honours that parent's sampled flag (with separately configurable behaviour for remote-sampled, remote-not-sampled, local-sampled and local-not-sampled parents); if there is no parent, it consults the delegate. This is what makes traces all-or-nothing rather than a scatter of unrelated fragments: exactly one service — the first one — decides, and everyone else complies. The SDK's default sampler is ParentBased(AlwaysOn). Read literally: honour the incoming decision, and if you are the root, keep it. ## Configuring it Environment configuration uses two variables: - `OTEL_TRACES_SAMPLER` — a name: `always_on`, `always_off`, `traceidratio`, `parentbased_always_on`, `parentbased_always_off`, `parentbased_traceidratio`, plus implementation-specific options such as a remotely-configured sampler or vendor samplers. - `OTEL_TRACES_SAMPLER_ARG` — the parameter for the chosen sampler; for the ratio samplers a number in [0,1] (`0.05` = 5%), for remote samplers a small key/value string carrying an endpoint and poll interval. Using a bare `traceidratio` rather than the parent-based variant is a classic mistake: each service then decides independently, and although the trace-id derivation makes them *agree* when the ratio is identical, any service configured with a different ratio produces partial traces. Wrap in ParentBased and vary only the root. ## The trade-off Dropping is what makes tracing affordable, and what it costs is coverage: at 1% the trace you want during an incident probably does not exist. Mitigations belong to different layers — biasing rates by route or importance via a custom sampler, or deferring the decision to a collector that can see the finished trace. Both are strategy decisions; the mechanics above are what you configure in the SDK. ## Custom samplers The interface is small and implementing it is legitimate: keep everything on a low-traffic critical endpoint, drop health checks entirely, sample by a start-time attribute. Two rules: it must be fast, because it runs on every span start; and it must be deterministic with respect to the trace id if you want consistency across services. ## What sampling is not A dropped span still carries a valid context and still propagates. Sampling is not a security or privacy control, it is not a rate limiter for your backend's ingestion (the exporter and processors handle back-pressure), and it never removes attributes from spans that were kept.
- Why can a head sampler not be configured to 'always keep traces that contain an error'?The decision is made when the first span starts, before anything has failed and before any downstream service has even been called. Error status is only known at span end, and whether any span in the trace errored is only known once the whole trace is complete. Keeping error traces therefore requires deciding after the fact, in a component that buffers the finished trace.
- What is the practical difference between a dropped span and a span that was never created?A dropped span still yields a valid span context, so its children get a coherent parent chain and outgoing requests still carry trace context with the sampled flag clear — downstream services see a consistent decision. A span that was never created leaves a real gap, and if the instrumentation is missing entirely the context may not propagate at all, orphaning everything downstream.
saying these in an interview costs you the question
- Configuring a bare ratio sampler on every service instead of the parent-based variant, then wondering about partial traces
- Believing head sampling can preferentially keep errors or slow requests
- Thinking a dropped span stops context propagation
- Assuming a ratio sampler samples uniformly per endpoint or per customer rather than per trace id
- Setting different ratios on different services with no parent-based wrapper