skip to content

Why not just set sampling probability to 1.0 everywhere? Discuss the cost-vs-coverage trade-off.

level: seniorimportance: should knowfreq 50%

answer

  1. cost ∝ traffic at 1.0
  2. backend storage dominates
  3. random drops errors too
  4. rate limiting = predictable cost
  5. tail sampling keeps the interesting ones

basics

~20 s

Sampling 100% is accurate but expensive: more CPU/memory per request, more network to the collector, and much higher storage/backend cost. Lower probabilities cut cost but risk missing the rare errored or slow request. You balance budget against how much visibility you need.

solid answer

~50 s

Setting probability=1.0 gives perfect coverage but scales cost linearly with traffic: every request pays serialization and export overhead, the collector and network carry full volume, and the tracing backend stores and indexes everything — which dominates the bill at high QPS. Lower probabilities (say 0.01–0.1) slash that cost but, because Spring's sampling is head-based and random, you're just as likely to drop the one request that errored or was slow as a boring one. The trade-off: high-volume, stable services can sample low and still see representative latency distributions; low-volume or critical services should sample high so rare events aren't lost. Common refinements are per-route/per-service rates, higher sampling on error-prone paths, always-sample on debug headers, or moving the keep-errors decision to tail sampling in an OpenTelemetry Collector so you pay full price only for interesting traces.

go deeper

for a junior

Know 1.0 is expensive and 0.1 is a cost compromise.

for a middle

Explain where cost lands (backend storage) and that random sampling can miss errors.

for a senior

Reason about per-service/per-route rates, rate limiting vs probability, and tail sampling.

for a principal

Design a fleet strategy balancing budget, incident-debuggability, exemplars, and spike safety.

**The two forces.** - **Coverage**: the fraction of real traffic you can actually inspect. Higher probability = more traces retained = better chance of catching a specific bad request and more statistically complete aggregates. - **Cost/overhead**: everything you pay to produce and keep traces. **Where the cost actually lands.** 1. **In-process overhead** — creating spans, capturing timestamps/attributes, serializing, and batching to the exporter costs CPU and heap per sampled request. Unsampled traces are cheap (IDs still generated, but no recording/export). 2. **Network** — exporting spans to a collector/backend consumes bandwidth and connections; at 1.0 this tracks total request volume. 3. **Backend storage & indexing** — usually the biggest line item. Trace backends (Tempo, Jaeger, vendor SaaS priced per span/GB) charge for ingest and retention. At high QPS, 100% sampling can be an order of magnitude more expensive than 10%. **Why low sampling hurts coverage (the head-based catch).** Spring/Micrometer sampling is **head-based and probabilistic** — it decides before the request runs, blind to outcome. So a 1% sampler drops 99% of *everything*, including 99% of your 500s and slow calls. For a bug that happens on 1 in 10,000 requests, low random sampling may never capture it. Random sampling is fine for *aggregate* latency/throughput but weak for *needle-in-haystack* debugging. **How to resolve the tension.** - **Tune per service by volume**: a service doing 10k req/s can sample at 0.01 and still get plenty of exemplars; a service doing 5 req/min should sample near 1.0 or you'll see almost nothing. - **Per-route / per-endpoint rates**: health checks and hot read paths low; checkout/payment paths high. In Brave this is a `RateLimitingSampler` or a custom `Sampler`; in OTel a `ParentBased`/rules sampler or collector config. - **Rate limiting instead of probability**: Brave's `RateLimitingSampler` caps traces/second, giving predictable cost regardless of traffic spikes (probability doesn't — a spike multiplies exports). - **Force-sample on demand**: honor a debug flag/header (or B3 `X-B3-Flags: 1`) so support can capture a specific user's trace at will. - **Tail-based sampling** (in an **OpenTelemetry Collector**, not in Spring): record everything at the edge, buffer full traces, then keep 100% of errors/slow traces and a small % of normal ones. Best coverage of *interesting* traces, but the collector must hold whole traces in memory and you pay full in-process/export cost up front. **Practical defaults.** Spring's 0.1 default is a compromise. For production, pick per-service based on: request volume, per-trace backend cost, retention needs, and how often you rely on traces for incident debugging. Pair modest head sampling with **exemplars** (metrics linking to sampled traces) and good metrics/logs so you're not solely dependent on catching the exact trace. **Gotchas.** - Probability sampling makes cost *proportional to traffic* — budget can blow up during incidents (exactly when traffic and errors spike). Rate limiting or tail sampling is safer there. - Aggregations computed from sampled traces need to know the sampling rate to extrapolate correctly; naive counts undercount. - Setting 1.0 in a high-QPS service 'just to be safe' is the classic way to melt your tracing bill and add latency.

  • How can you keep 100% of error traces without paying to store all traces?
    Use tail-based sampling in an OpenTelemetry Collector: it buffers complete traces and keeps all that errored or exceeded a latency threshold plus a small percentage of normal ones. Spring itself only does head-based, so this lives in the collector.
  • Why might rate limiting be preferable to a fixed probability in production?
    Probability makes export volume scale with traffic, so a spike multiplies cost and overhead right when you can least afford it. A rate limiter (e.g. Brave's RateLimitingSampler) caps traces/second, giving predictable, bounded cost regardless of load.
  • If you sample at 10%, can you still trust your latency dashboards?
    For aggregate distributions, yes — a 10% random sample is representative — as long as consumers account for the sampling rate when extrapolating counts. It's the rare individual bad request that 10% random sampling can miss.

saying these in an interview costs you the question

  • Assuming higher sampling is always better with no cost consideration
  • Thinking probabilistic sampling preferentially keeps errors (it doesn't — it's outcome-blind)
  • Believing tail sampling is a Spring Boot property (it's a collector feature)
  • Setting 1.0 on high-QPS services 'to be safe'

context