skip to content

Trace Sampling

Sampling decides which traces are kept, usually by probability at the head of the request, with the decision propagated so a trace is complete or absent. Interviewers ask how you keep tracing affordable without losing the traces you actually need.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

questions

5

What does the management.tracing.sampling.probability property do in Spring Boot, and what is its default value?

level: juniorimportance: must knowfreq 70%

answer

  1. double 0.0–1.0
  2. default 0.1 = 10%
  3. head-based, decided once
  4. 1.0 for local debugging
  5. 0.0 ≠ tracing fully off

basics

~10 s

It sets the fraction of traces that get recorded and exported, from 0.0 (none) to 1.0 (all). The default in Spring Boot is 0.1, meaning about 10% of traces are sampled.

solid answer

~40 s

management.tracing.sampling.probability configures Micrometer Tracing's sampler with the fraction of new traces to record and export, expressed as a double between 0.0 and 1.0. The default is 0.1, so roughly 10% of root traces are kept and the other 90% are dropped to save cost and overhead. The decision is made once per trace at its start (head-based) and then applied consistently to every span in that trace. Setting it to 1.0 samples everything (useful locally or when debugging); 0.0 disables trace export while the tracing infrastructure still runs. It works the same whether the underlying bridge is Brave or OpenTelemetry, since it is a Spring Boot abstraction property, not a vendor-specific one.

code

java · 9 lines
java
// application.properties
// Global sampling fraction: record & export ~10% of traces (the default)
// management.tracing.sampling.probability=0.1

// For local debugging, sample everything so a single request shows up:
// management.tracing.sampling.probability=1.0

// Disable export entirely (context still propagated, nothing recorded):
// management.tracing.sampling.probability=0.0

go deeper

for a junior

Know it's a 0.0–1.0 fraction, default 0.1, and set 1.0 to see traces locally.

for a middle

Explain head-based (decided once per trace) and that it's bridge-independent.

for a senior

Tie the value to cost/volume and note upstream decisions can override local probability.

for a principal

Frame probability as one knob in a broader sampling strategy (head vs tail, per-route rates).

**What sampling is.** Distributed tracing records the path of a request across services as a *trace* made of *spans* (individual timed operations). Recording, storing, and exporting every trace to a backend (Zipkin, Tempo, Jaeger, etc.) is expensive in CPU, memory, network, and storage. *Sampling* is the mechanism that decides which traces to actually record and send. **The property.** In Spring Boot 3.x (which uses Micrometer Tracing), `management.tracing.sampling.probability` is a `double` in the range `[0.0, 1.0]`. It is bound to Micrometer's sampler configuration and controls the probability that any given new trace is sampled (recorded + exported). - `1.0` = sample 100% of traces. - `0.1` = sample ~10% (this is the **Spring Boot default**). - `0.0` = sample nothing (no traces exported), though span context still propagates. **Where it lives.** Put it in `application.properties` / `application.yml`: ```properties management.tracing.sampling.probability=1.0 ``` **Head-based decision.** The sampler decides *once*, at the moment a trace is created (the "head" of the trace), and that boolean decision is recorded on the span context and reused for the whole trace. It is not re-evaluated per span. This keeps a trace whole — you never get a trace that is half-sampled. **Bridge independence.** The property is a Spring Boot abstraction. Whether you include `micrometer-tracing-bridge-brave` or `micrometer-tracing-bridge-otel`, Spring Boot auto-configures the appropriate `Sampler` (Brave `Sampler` or OpenTelemetry `Sampler`) from this single property, so you rarely touch vendor APIs directly. **Gotchas.** - Sampling ≠ tracing on/off. Even at `0.0`, trace/span IDs are still generated and propagated in headers; you just don't export data. To fully disable, exclude the tracing autoconfiguration or don't add a tracer bridge. - The probability applies to *new* (root) traces. For an incoming request that already carries a sampling decision from an upstream service, that decision is normally honored (see head-based propagation), so your local probability may not apply. - 10% default surprises people who set up tracing, hit an endpoint once, and see nothing exported. Bump to `1.0` while developing. **When to change it.** Raise toward `1.0` in dev and low-traffic services; lower it in high-throughput production services to control cost. Choose a value based on request volume × per-trace cost vs. the coverage you need for debugging.

  • You set up tracing, call an endpoint once, and nothing shows up in Zipkin. Why?
    The default probability is 0.1, so a single request has only a ~10% chance of being sampled. Set management.tracing.sampling.probability=1.0 while developing so every trace is exported.
  • Does probability=0.0 turn tracing off completely?
    No. Span/trace IDs are still generated and context is still propagated in headers; it only stops traces from being recorded and exported. To fully disable, remove the tracer bridge or exclude the tracing auto-configuration.

saying these in an interview costs you the question

  • Saying the default is 1.0 / 100%
  • Thinking 0.0 stops trace-context propagation, not just recording
  • Believing the property differs between Brave and OpenTelemetry bridges

context

open as a page

Explain head-based sampling and how the sampling decision is propagated between services via the sampled flag.

level: middleimportance: must knowfreq 60%

basics

~20 s

The sample-or-not decision is made once at the start (head) of a trace and stored on the context. It's carried to downstream services in the trace-context headers (e.g. the W3C traceparent flags byte), so every service in the trace makes the same keep/drop choice.

open as a page

Why not just set sampling probability to 1.0 everywhere? Discuss the cost-vs-coverage trade-off.

level: seniorimportance: should knowfreq 50%

basics

~20 s

Sampling 100% is accurate but expensive: more CPU/memory per request, more network to the collector, and much higher storage/backend cost. Lower probabilities cut cost but risk missing the rare errored or slow request. You balance budget against how much visibility you need.

open as a page

How would you configure a custom sampling strategy in Spring Boot — for example, always sampling a critical endpoint but rate-limiting the rest?

level: seniorimportance: should knowfreq 35%

basics

~10 s

For simple cases set management.tracing.sampling.probability. For custom logic, define your own Sampler bean (Brave's brave.sampler.Sampler or OpenTelemetry's io.opentelemetry.sdk.trace.samplers.Sampler), which overrides the property. Brave offers RateLimitingSampler and per-path samplers you can compose.

open as a page

What are the fundamental limitations of probabilistic head-based sampling, and how would you architect around them at scale?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Head-based sampling decides before the outcome is known, so it randomly drops errors and slow requests and can't guarantee capturing rare problems. At scale you add tail-based sampling in a collector, rate limiting, per-route rates, force-sample-on-debug, and lean on metrics/exemplars so tracing isn't your only signal.

open as a page