What are the fundamental limitations of probabilistic head-based sampling, and how would you architect around them at scale?
answer
- head = outcome-blind, decided too early
- random drops errors equally
- cost peaks with traffic at incidents
- tail sampling keeps errors/slow
- traces are one pillar: metrics + exemplars + logs
basics
~20 sHead-based sampling decides before the outcome is known, so it randomly drops errors and slow requests and can't guarantee capturing rare problems. At scale you add tail-based sampling in a collector, rate limiting, per-route rates, force-sample-on-debug, and lean on metrics/exemplars so tracing isn't your only signal.
solid answer
~50 sThe core limitation is that management.tracing.sampling.probability drives a head-based, outcome-blind decision: it's fixed before the request runs, so errors and high-latency traces are dropped at the same rate as normal ones, and cost scales with traffic — worst during incident spikes. Statistically it's fine for aggregate latency but poor for needle-in-haystack debugging, and it can't express 'keep all errors.' At scale I keep in-app head sampling modest and consistent (parent-based so traces stay coherent), then push outcome-based selection to an OpenTelemetry Collector doing tail sampling — buffer complete traces, keep 100% of errors/slow ones plus a small baseline of normal traffic. I add rate limiting for spike-safe bounded cost, per-route rates (hot reads low, critical flows high), and debug-header force sampling for support. Crucially I treat traces as one pillar: RED metrics with exemplars and structured logs give coverage that sampling alone can't, so a dropped trace isn't a blind spot.
go deeper
Know head sampling can miss errors because it decides too early.
Explain outcome-blindness and that tail sampling in a collector fixes error retention.
Combine head + tail + rate limiting + route rules and know the collector trade-offs.
Design a fleet-wide strategy balancing cost, coverage, spike-safety, trust boundaries, and the metrics/logs/exemplars pillars; understand trace-affinity routing and buffering limits.
**The limitations, precisely.** 1. **Outcome-blindness.** Head-based sampling (all Spring/Micrometer does natively) commits the keep/drop decision at trace start, before status/latency exist. A probabilistic sampler therefore drops errors and slow requests at exactly the same rate as successes. You cannot say 'always keep 5xx traces' with a head sampler. 2. **Rare-event miss.** For a fault occurring at low frequency, low random sampling may never capture the specific offending trace, even though aggregate dashboards look complete. 3. **Cost couples to traffic.** Probability makes export/storage volume proportional to load, so cost (and per-request overhead) peaks during traffic spikes — often coincident with incidents, the worst time. 4. **Uniformity.** A single global probability treats a health check and a payment the same unless you add per-route logic. 5. **Extrapolation burden.** Counts derived from sampled traces must be scaled by the sampling rate; naive consumers undercount. **Architecting around them.** - **Tail-based sampling in the OpenTelemetry Collector.** The decisive move. The `tail_sampling` processor buffers all spans of a trace until it's complete, then applies policies: keep if any span errored, keep if duration > threshold, keep a probabilistic baseline of the rest. This gives near-total coverage of *interesting* traces at a fraction of storage. Trade-offs: the collector must hold whole traces in memory (memory + a decision-wait window), traces must route to a consistent collector instance (load-balancing exporter by trace ID), and you still pay in-app span creation + export up front (you emit everything, the collector discards). - **Consistent, parent-based head sampling in-app.** Keep a modest global rate so distributed traces remain coherent end-to-end; wrap in `parentBased` so downstream honors upstream. Don't try to be clever per-hop. - **Rate limiting** (Brave `RateLimitingSampler`) for bounded, spike-safe cost when you don't run a tail-sampling collector. - **Per-route / per-tenant rates** — de-prioritize health checks and hot read paths; elevate checkout/auth/payment. - **Force-sample-on-demand** — honor a debug flag/header (or B3 flags) so support/on-call can capture a specific user's or request's trace deterministically. - **Don't make tracing the only pillar.** Emit **RED/USE metrics** (Micrometer) always-on and unsampled, attach **exemplars** (metric samples linking to a stored trace) so you can jump from a latency spike to a real trace, and keep **structured logs** correlated by traceId/spanId (Micrometer's MDC propagation). Then a dropped trace is not a coverage hole because metrics and logs still cover it. **Decision framework.** - Low volume or business-critical service → high head rate (near 1.0), maybe no tail sampling needed. - High volume → low head rate is unavoidable for cost; add a tail-sampling collector to still keep errors/slow traces. - Strict, predictable budget / spiky traffic → rate limiting or adaptive sampling over fixed probability. - Regulated/audit needs on specific flows → route-based always-on for those, low elsewhere. **Gotchas at scale.** - Tail sampling needs **trace-complete buffering** and **trace-affinity routing**; a naive fan-out to many collector replicas splits a trace across instances and breaks the error-keep policy. Use the load-balancing exporter keyed on traceId. - Buffering adds a decision-latency window and memory pressure; size it for your slowest traces. - Even with tail sampling you emit all spans from apps, so app-side CPU/network isn't saved — only backend storage is. If app overhead is the concern, you still need head sampling. - Honoring upstream sampled flags across a **trust boundary** lets an external caller force 100% sampling; at the edge you may deliberately re-decide rather than honor.
- Does tail sampling reduce the tracing overhead inside your Spring apps?No. Apps still create and export all spans; the collector decides what to discard. Tail sampling saves backend storage/ingest cost and gives error/latency-aware retention, but in-process CPU, memory, and export bandwidth are unchanged. To cut app-side overhead you still need head sampling.
- Why does tail sampling require special routing across collector replicas?A trace's spans may arrive from many services at different times; the tail processor must see the whole trace to judge errors/latency. All spans of one trace must land on the same collector instance, so you use a load-balancing exporter that routes by trace ID (trace affinity).
- How do you avoid a blind spot when a problematic trace happens to be dropped?Don't rely on traces alone. Emit always-on RED metrics, attach exemplars that link metric spikes to whichever traces were stored, and keep structured logs correlated by traceId/spanId. Metrics and logs cover the request even when its trace wasn't sampled.
saying these in an interview costs you the question
- Claiming head-based sampling can 'keep all errors' (it can't — outcome is unknown at decision time)
- Thinking tail sampling saves in-app overhead (it only saves backend storage)
- Fan-out to many collector replicas without trace-affinity routing
- Treating traces as the sole observability signal instead of pairing with metrics/logs/exemplars
- Assuming a single global probability suffices for all routes and tenants