A tracing backend like Jaeger can't afford to store a full trace for every request in a system doing 500,000 requests per second. Compare head-based and tail-based sampling as strategies for deciding which traces to keep, and give a scenario where head-based sampling would make you miss the exact trace you needed.
answer
- head = decide at root, before outcome known
- tail = decide after full trace assembled
- sampled flag propagated via trace context
- tail needs buffering + completion window
- rare errors get missed by pure head-based sampling
basics
~20 sSampling means only keeping some traces, not all of them, to control cost. Head-based sampling decides at the very start of a request, before anything is known about it, usually just by rolling dice, like keep 1%. Tail-based sampling waits until the whole request is finished, then decides based on what actually happened, like keeping it if it was slow or errored.
solid answer
~60 sHead-based sampling makes the keep or discard decision at the root of the trace, before the request has even executed, typically via a fixed probability such as 1% or a rate limiter, and that decision is propagated to every downstream span via the sampled flag in the trace context, so either the whole trace is recorded or none of it is. It's cheap, requires no buffering, and works at any scale, but it's blind: a request that turns out to be a critical, high-latency error has the same chance of being kept as a boring, fast, successful one. Tail-based sampling defers the decision until the full trace has been assembled, often by buffering all spans for a trace in a collector for some window, then applies rules like always keep traces with errors, or always keep the slowest 1%, which means you can guarantee capturing the traces that actually matter for debugging, at the cost of needing to buffer a much larger volume of span data than you ultimately keep, plus added latency and infrastructure complexity in the collection tier.
go deeper
Should understand that sampling means keeping only some traces to save cost, without needing precise mechanics.
Should be able to explain that head-based decides upfront and tail-based decides after the outcome is known.
Should articulate the buffering and completion-window cost of tail-based sampling and design a scenario showing head-based sampling's blind spot for rare errors.
Should be setting sampling policy and collector architecture, buffering windows, hybrid head-plus-tail strategies, per-endpoint sampling rates, as a cost versus observability trade-off decision at the platform level.
## Why sampling exists at all Sampling exists because recording a complete trace, every span, on every service, for every single request, simply doesn't scale for high-traffic systems: at 500,000 requests per second, even a lean average trace of a few kilobytes would produce well over a terabyte of trace data per second, which is untenable to transport, store, and query, and mostly wasted, since the overwhelming majority of requests are unremarkable successes that nobody will ever look at. Sampling is the decision procedure for choosing which fraction of traces get fully recorded, and the fundamental design question is when that decision gets made. | | Head-based | Tail-based | |---|---|---| | Decision | at the root span, before any downstream work has happened and therefore before anything is known about how the request will turn out | deferring the decision until after the full trace's outcome is known | | Cost | the decision is made blind: a request about to hit a rare, severe bug has exactly the same probability of being sampled as a perfectly ordinary one | operational complexity and resource cost in the collection tier rather than in trace volume itself | ## Deciding at the root Head-based sampling decides at the root span, the very first service that starts the trace, before any downstream work has happened and therefore before anything is known about how the request will turn out. The simplest form is a fixed probability: flip a biased coin with, say, a 1% chance of sample. That decision is then propagated to every downstream service via the sampled flag in the W3C trace-context header, so every service in the call chain either records all its spans, if sampled, or records none, if not, keeping the recording decision consistent across the whole distributed call graph without any service needing to coordinate with any other. The mechanism is deliberately simple: - no service needs to buffer anything, - no collector needs to hold spans waiting for a decision, - and the system's storage and throughput cost is a predictable, fixed fraction of total traffic, which makes head-based sampling trivial to reason about and to scale. ## The blind spot The cost of that simplicity is that the decision is made blind. A request that is going to be perfectly ordinary has exactly the same probability of being sampled as a request that is about to hit a rare, severe bug, time out after ten seconds, or fail with an error in the payments service. At 1% sampling, a bug that occurs in 1 out of every 100,000 requests has only about a 1-in-10,000 chance of ever being captured in a trace at all, which for a genuinely rare but important failure mode means it may effectively never show up in the tracing system, even though it's exactly the kind of thing an engineer would want a trace for. ## Deciding after the outcome Tail-based sampling fixes this by deferring the decision until after the full trace's outcome is known. Instead of deciding at the root, an intermediary component, usually a tracing collector configured to buffer, holds all the spans belonging to one trace ID for some completion window, long enough for the slowest expected request to finish, and only then applies rules such as - keep if any span had an error status, - keep if total duration exceeded 500ms, - or keep a random 1% of everything else as a baseline sample. This guarantees that the traces most valuable for debugging, the errors and the slow outliers, are essentially always captured, regardless of how rare they are, while still controlling overall storage cost by discarding the bulk of unremarkable, fast, successful traces. ## What deferral costs The trade-off is operational complexity and resource cost in the collection tier rather than in trace volume itself. Every span from every service, sampled or not, has to be transmitted to the buffering collector, because you can't know in advance which traces will turn out to deserve keeping; the discarding only happens after the fact. That means the network and collector-ingestion cost looks a lot like 100% sampling even though final storage looks like a small fraction, and the collector needs enough memory to buffer every in-flight trace for the length of the completion window, which is itself a scaling and tuning problem, a window too short and slow traces get evaluated before they finish, a window too long and buffering cost balloons. Tail-based sampling also adds latency to the observability pipeline itself, since a trace isn't finalized in storage until its completion window has elapsed. ## The trace you would have missed A concrete scenario where head-based sampling causes a real miss: an intermittent bug causes 1 in 20,000 checkout requests to fail with a database deadlock, and the team runs head-based sampling at 0.5%. The odds that the specific request an angry customer reported, identified after the fact by its correlation ID, was also one of the ones randomly chosen for full tracing are roughly 1 in 200; in practice, the trace simply doesn't exist, and the on-call engineer is left reconstructing the incident from logs alone. With tail-based sampling and an always keep errored traces rule, that same deadlocked request would have been captured with certainty, because the decision was made after the collector saw that the trace ended in an error, not before. Production systems often combine both: head-based sampling at a modest rate to cap the raw ingestion volume the collector tier needs to handle, layered with tail-based rules on top of what does reach the collector, to bias what gets kept in final storage toward errors and outliers rather than a uniformly random slice.
- How does the sampled flag in the trace context actually enforce a consistent head-based decision across five services?The root service makes the coin-flip once and sets a bit in the trace-context header it sends downstream; every subsequent service reads that bit instead of making its own independent decision, and either records and exports its spans, if set, or does minimal or no recording, if not, so the whole call chain stays consistently sampled-in or sampled-out.
- Why can't you just set head-based sampling to 100% and skip the whole problem?At high request volumes, 100% sampling means transporting, ingesting, and storing a full trace for every single request indefinitely, which scales linearly with traffic and becomes prohibitively expensive in both infrastructure cost and query performance; sampling exists specifically to decouple observability cost from raw traffic volume.
- What's a middle-ground technique between pure head-based and pure tail-based sampling?Rate-limiting or adaptive head-based sampling, which adjusts the sampling probability based on recent traffic patterns, sampling more heavily for low-volume, rarely-called endpoints and less for high-volume ones, is one option; another is combining a modest head-based rate to cap ingestion with tail-based rules layered on top to bias what's kept toward errors and outliers.
Head-based sampling is like a security guard deciding, before anyone even walks in the door, to randomly search 1 in 100 visitors regardless of how they behave. Tail-based sampling is like watching everyone the whole time and only keeping the footage of visitors who actually did something suspicious, which catches every incident but means everyone had to be watched, and briefly recorded, first.
saying these in an interview costs you the question
- Thinks tail-based sampling decides before the request completes
- Doesn't realize head-based sampling can miss rare but important errors entirely
- Assumes sampling only affects storage cost, not ingestion or network cost
- Can't explain how the sampling decision stays consistent across services
- Proposes 100% sampling as a default with no cost discussion