Explain head-based sampling and how the sampling decision is propagated between services via the sampled flag.
answer
- decide once at the head
- sampled flag on context
- traceparent last byte, bit 0
- -01 sampled, -00 not
- downstream honors, no re-dice
basics
~20 sThe sample-or-not decision is made once at the start (head) of a trace and stored on the context. It's carried to downstream services in the trace-context headers (e.g. the W3C traceparent flags byte), so every service in the trace makes the same keep/drop choice.
solid answer
~40 sHead-based sampling means the keep-or-drop decision is made a single time, at the trace's origin, by the sampler configured via management.tracing.sampling.probability. That boolean is recorded on the span context as a 'sampled' flag. When the service makes an outgoing call, Micrometer Tracing's propagator writes the trace context into headers — with W3C Trace Context that's the traceparent header, whose final byte encodes flags, and bit 0 is the 'sampled' flag (…-01 = sampled, …-00 = not sampled). The downstream service reads traceparent, sees the flag, and honors it instead of rolling its own dice, so the whole distributed trace is consistently sampled or dropped end-to-end. This gives coherent, gap-free traces but means the decision is committed before you know whether the request errored or was slow.
code
java · 13 lines// W3C traceparent header carrying the decision to the next service:
// version traceId parentSpanId flags
// 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
// ^^ traceFlags byte, bit 0 = sampled
// -01 => sampled (record & export) -00 => not sampled (drop)
// B3 equivalent (Brave default): the decision travels as its own header:
// X-B3-TraceId: 4bf92f3577b34da6a3ce929d0e0e4736
// X-B3-SpanId: 00f067aa0ba902b7
// X-B3-Sampled: 1 // 1 = sampled, 0 = not sampled
// Choose the propagation wire format in Spring Boot:
// management.tracing.propagation.type=W3C (or B3)go deeper
Know the decision is made once and travels in headers.
Name traceparent's flags byte / B3 X-B3-Sampled and that downstream honors it.
Explain the outcome-blindness limitation and cross-format pitfalls.
Discuss trust boundaries (honoring vs overriding upstream flags) and format governance across a fleet.
**Head vs tail sampling.** *Head-based* sampling decides at the **head** (start) of a trace, before the request is processed, whether to record it. *Tail-based* sampling waits until the trace is complete and decides based on outcome (errors, latency). Spring Boot / Micrometer Tracing does **head-based** sampling out of the box; tail-based requires an external collector (e.g. the OpenTelemetry Collector's tail-sampling processor). **Making the decision once.** When a request arrives with no existing trace context, the tracer creates a new trace and asks the `Sampler` (configured from `management.tracing.sampling.probability`) whether to sample. The resulting boolean is stored on the span/trace context as the **sampled flag**. It is *not* re-evaluated for child spans or downstream services. **Propagation — the wire format.** To keep a distributed trace coherent, the decision must travel with the request. Micrometer Tracing uses a **propagator** to inject trace context into outbound headers and extract it from inbound ones. - **W3C Trace Context** (the modern default): the `traceparent` header looks like `00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01` = `version-traceId-parentSpanId-traceFlags`. The last field, `traceFlags`, is a two-hex-digit byte; **bit 0 is the *sampled* flag**. `-01` means sampled, `-00` means not sampled. - **B3** (Zipkin/Brave legacy): either a single `b3` header or multiple headers including `X-B3-Sampled: 1` / `0`, and `X-B3-TraceId` etc. **Downstream honoring.** The receiving service extracts the context, sees the incoming sampled flag, and continues the *same* trace with the *same* decision. So if service A decided "sampled," B, C, and D all record their spans; if A decided "not sampled," nobody downstream records. This is what prevents partial traces where only some hops appear. **Micrometer/Spring specifics.** - The bridge is either `micrometer-tracing-bridge-brave` (Brave `Sampler`, B3 by default) or `micrometer-tracing-bridge-otel` (OpenTelemetry `Sampler`, W3C by default). - Propagation format is configurable via `management.tracing.propagation.type` (`W3C`, `B3`, etc.). - Instrumented clients (`RestTemplate`, `WebClient`, `RestClient`, messaging) inject headers automatically once tracing is on. **Gotchas.** - Because the decision is committed at the head, an **erroring or slow request may be dropped** if it wasn't sampled — head-based sampling is blind to outcome. That's the core limitation and the reason tail sampling exists. - If an upstream sends `-01`, your local `probability` is effectively bypassed for that trace — you honor the upstream decision. Misconfigured or malicious upstreams can therefore force 100% sampling; some setups use a sampler that ignores the incoming flag for untrusted edges. - Mismatched propagation formats between services (one on B3, one on W3C) break context continuity, producing disconnected traces. - The sampled flag is a *boolean per trace*, not per span — you cannot sample individual spans within a sampled trace differently via this mechanism.
- What's the main weakness of head-based sampling?The decision is made before the request runs, so it can't be based on outcome — an errored or slow trace that wasn't sampled is lost. Tail-based sampling (via an external collector) solves this by deciding after the trace completes.
- Which header carries the W3C sampling decision and where exactly?The traceparent header. Its last field, traceFlags (two hex digits), holds the flags byte; bit 0 is the sampled flag — '01' means sampled, '00' means not sampled.
- What happens if two services use different propagation formats?Context extraction fails to find the expected headers, so the downstream service starts a new, disconnected trace. You must align management.tracing.propagation.type across services.
saying these in an interview costs you the question
- Thinking the sampling decision is re-made independently at each service
- Believing sampled=0 means the trace headers aren't sent (they still are; only recording stops)
- Confusing the sampled flag with the trace ID