In production, sampling probability is 0.1 and users report 'traces are missing.' Explain how sampling, context propagation, and the exporter interact, and what you'd verify.
answer
- 0.1 = 90% dropped by design
- head decision at root, propagated flag, honored downstream
- unsampled still propagates ids → logs still correlated
- export failures separate: endpoint/path/flush
- tail sampling in Collector keeps errors/slow traces
basics
~20 sAt 0.1 only ~10% of traces are exported by design — that's expected, not a bug. The sampling decision is made once at the trace root and propagated downstream, so all services agree. Verify the endpoint, that ids still appear in logs (proving propagation works), and consider raising probability or using tail sampling in the collector.
solid answer
~40 smanagement.tracing.sampling.probability=0.1 is a head-based sampler: at each new trace root a random ~10% are marked sampled, and that decision rides the propagation header (traceparent sampled flag / B3) so every downstream service exports consistently — you never get half a trace. So 90% 'missing' is correct behavior. Before calling it broken I'd confirm: (1) trace/span ids still show in logs — if yes, propagation and the tracer work, only export is filtered; (2) the exporter endpoint (management.otlp/zipkin.tracing.endpoint) is reachable and the path is right (/v1/traces vs /api/v2/spans); (3) the backend isn't dropping spans. If they need more coverage without cost blowing up, move to tail-based sampling in the OTel Collector (sample after seeing errors/latency) and keep app-side probability high, or raise probability for specific routes. Unsampled traces still propagate ids for log correlation.
code
java · 21 lines// application.properties
// Head-based probabilistic sampling (10%):
// management.tracing.sampling.probability=0.1
// Export target must match the bridge + be reachable:
// management.otlp.tracing.endpoint=http://otel-collector:4318/v1/traces
// Principal move: keep app sampling HIGH, decide in the Collector (tail sampling).
// otel-collector-config.yaml (sketch):
// processors:
// tail_sampling:
// policies:
// - name: errors
// type: status_code
// status_code: { status_codes: [ERROR] }
// - name: slow
// type: latency
// latency: { threshold_ms: 500 }
// - name: baseline-10pct
// type: probabilistic
// probabilistic: { sampling_percentage: 10 }
// => always keep error/slow traces, sample the rest.go deeper
Know sampling.probability limits how many traces are exported; 0.1 means ~10%.
Explain the decision is made once and propagated so services stay consistent; check the endpoint.
Separate sampling from export failures and use logs (ids present) to localize the problem.
Design a tail-sampling/error-aware strategy in the Collector, enforce propagation consistency, and reason about cost vs. coverage and flush guarantees.
## First: distinguish 'not sampled' from 'not working' The single most common misread of tracing in prod: **low sampling looks like breakage**. With `management.tracing.sampling.probability=0.1`, roughly **90% of traces are intentionally never exported**. That is the sampler doing its job to control cost/volume, not a fault. ## How head-based (probabilistic) sampling works - The decision is **head-based**: made **once, at the trace root**, when a new trace starts (no incoming sampled decision). - It's encoded into the **propagation header**: the W3C `traceparent` trailing flags (`-01` sampled, `-00` not) or the B3 `X-B3-Sampled` header. - Downstream services use **parent-based sampling**: they **honor** the upstream decision rather than re-rolling. So a trace is sampled or dropped **consistently end-to-end** — you never get an orphaned half-trace where A exported but B didn't. - **Unsampled traces still propagate ids.** The tracer still assigns trace/span ids and puts them in MDC, so **logs remain correlated** even for traces that never export. This is a key diagnostic signal. ## The export pipeline is separate from sampling Even for sampled spans, export can fail independently: - **Endpoint/path**: OTLP HTTP is `/v1/traces`; Zipkin is `/api/v2/spans`. A wrong path silently drops. - **Reachability**: collector down, DNS, TLS, auth headers. - **Batching**: SDKs batch and flush on an interval / on shutdown; a crash before flush loses the buffer. Short-lived jobs may exit before flush. - **Backend retention/quotas**: Tempo/Zipkin may drop or expire. ## Diagnostic checklist (what I'd actually verify) 1. **Look at logs**: do `[app,traceId,spanId]` ids appear? If **yes**, the tracer and propagation work — the issue is *only* sampling/export filtering, not instrumentation. 2. **Temporarily set probability=1.0** in a canary/staging and reproduce: if traces now appear, it was sampling volume. 3. **Check the exporter endpoint** property and hit it manually; confirm path and that the collector receives OTLP/Zipkin. 4. **Check the collector/backend** for drops, rate limits, or retention windows. 5. **Confirm one bridge + matching exporter** and a consistent **propagation format** across services (mismatched format → broken/duplicate traces, which *also* looks like 'missing'). 6. **Short-lived processes**: ensure graceful shutdown so the batch span processor flushes. ## Fixing coverage without cost blow-up: tail-based sampling Head sampling is blind — it decides before it knows if the request errored or was slow, so it often drops exactly the interesting traces. The principal-level answer: - Keep application-side sampling **high (even 1.0)** and move the decision to a **tail-based sampler in the OpenTelemetry Collector**. Tail sampling buffers a whole trace, then keeps it based on outcome (has an error, exceeded a latency threshold, matches a route), dropping the boring majority. - Alternatively, per-route or rule-based sampling for critical endpoints. - Trade-off: tail sampling needs the Collector to buffer complete traces (memory, and all spans of a trace must reach the same collector instance / use a load-balancing exporter). ## Governance considerations (why this is principal-level) - **Cost vs. observability**: export volume drives collector/storage cost; pick a sampling strategy deliberately, document it. - **Consistency**: enforce one propagation format and one bridge org-wide, or cross-service traces silently fragment. - **SLO alignment**: guarantee error/high-latency traces are always captured (tail sampling or error-based head rules) even if baseline probability is low. - **Flush guarantees** for batch/serverless workloads. ## Gotchas summary - Low probability = expected missing traces, not a bug. - Ids in logs prove propagation works even when nothing exports. - Head sampling can drop error traces; tail sampling in the Collector fixes that. - Wrong endpoint path or missing flush drops even sampled spans.
- Why can head-based sampling be a poor fit for capturing production incidents?The keep/drop decision is made at the trace root before the request outcome is known, so slow or errored requests are dropped at the same 10% rate as normal ones. Tail-based sampling in the Collector buffers the full trace and keeps it based on error/latency, capturing the interesting ones.
- How do you know instrumentation works if almost no traces reach the backend?Check the logs: if trace/span ids appear in the [app,traceId,spanId] pattern, the tracer is creating and propagating context correctly — the shortfall is sampling and/or export, not instrumentation. Setting probability to 1.0 in staging confirms it.
- A downstream service exports spans under a different trace id than its caller. What's a likely cause?Propagation-format mismatch (caller sends W3C traceparent, callee reads B3 or vice versa) so extraction fails and the callee starts a new trace. Align management.tracing.propagation.type across services.
saying these in an interview costs you the question
- Concluding tracing is broken purely because most traces are absent at probability 0.1
- Thinking each service independently re-rolls the sampling decision (leading to half-traces)
- Believing unsampled traces lose log correlation ids
- Assuming head sampling reliably captures error/slow traces
- Overlooking exporter endpoint path or batch-flush-on-shutdown as an export failure mode