In head-based trace sampling, where is the keep-or-drop decision made, and why is it nearly free?
answer
- Decided before the work happens
- One decision per trace, not per span
- Nothing is buffered anywhere
- Dropped traces are never built
- The verdict rides on the request
basics
~20 sHead-based sampling decides at the start of a trace, before any span is recorded. The first service rolls the dice once and stamps the verdict onto the outgoing request, so a dropped trace costs nothing downstream.
solid answer
~50 s**Head-based** means the decision is taken at the *head* of the trace: the first process to handle the request creates the trace and, in the same breath, decides whether it will be kept. That is one draw against a ratio or a rate limit, taken once per trace rather than once per span or once per service. It is nearly free because a *no* is free everywhere afterwards. Nothing is buffered waiting for the request to finish, so the sampler holds no memory; instrumentation skips building spans and attributes; nothing is serialised, crosses the network, or reaches storage. The price is what it cannot see: the decision precedes the outcome, so it cannot keep a trace *because* it failed. And for the kept traces to be whole rather than shredded, the verdict must travel with the request and every downstream service must obey it instead of drawing again.
code
pseudocode · 12 lineshandle(request):
ctx = extract_context(request)
sampled = ctx.sampled if ctx.present
else sampler.decide(request) # one draw, once per trace
if not sampled:
propagate(trace_id, sampled = false) # all that is left to do
return do_work(request) # no span built, nothing exported
span = start_span(request) # only the kept fraction pays this
propagate(trace_id, sampled = true)
...go deeper
Be ready to say what a trace is and that sampling means not every request is kept. Know that the decision happens once, right at the start, and that the services further along the request follow it rather than deciding again.
Explain the mechanics: one draw per trace in the process that starts it, the verdict carried onward in the propagated context, and the fact that a dropped trace is never constructed rather than built and then binned.
Show what the cheapness buys and costs in production: no buffering and near-zero overhead, against a decision taken before any outcome exists, so failures survive only in proportion to the ratio you chose.
Own the tradeoff across a fleet: where the head sits for each entry point, which routes deserve their own ratios or floors, and whether the money saved by deciding early is worth the incidents you will never have an example of.
## What "head-based" names A **trace** is every span produced by one request as it moves through a system, tied together by a shared trace identifier. **Head-based sampling** takes the keep-or-drop decision at the head of that trace: at the moment the trace is created, in the first process that handles the request, before any of the work being traced has happened. Three properties of that decision point do all the work: - **One decision per trace, not one per span.** A request that fans out into forty spans across six services is a single decision, taken once. - **Taken before the outcome exists.** The sampler sees the request as it arrives — its route, its method, its tenant, whatever attributes are present on arrival — plus a source of randomness. It cannot see the status code, the duration, or whether a downstream call timed out, because none of that has happened. - **Taken by whichever process starts the trace.** Usually the edge service or gateway; for work that does not begin on the request path, the scheduler or consumer that creates the trace is the head. The decision itself is trivial arithmetic: a random draw against a ratio ("keep one in two hundred"), a deterministic function of the trace identifier so that any process computing it reaches the same answer, or a counter against a rate limit. ## Why it costs almost nothing The saving is not that dropped traces are cheap to throw away. It is that they are never built. | Cost line | Deciding at the head | Deciding after the fact | | --- | --- | --- | | Memory held while deciding | none | every span of every in-flight trace | | Added decision latency | none | one wait window per trace | | Span objects for dropped traces | never constructed | constructed, then discarded | | Bytes leaving the process | only kept traces | every span | | Routing constraint | none | all of a trace's spans to one decider | Instrumentation checks the sampled flag before it allocates: no span object, no attribute map, no timestamps, no events, no batching, no serialisation, no export. Past the process the saving compounds — nothing crosses the network, nothing lands in the collection tier, nothing is indexed, nothing is retained, nothing is scanned at query time. A one-in-two-hundred ratio really does cost roughly half a percent of the telemetry bill, rather than the full runtime cost with a small storage bill on the end of it. What does not go away is a small fixed cost per request. The service still extracts the incoming context and injects context into its outgoing calls, because it must be able to continue a caller's trace and must let downstream services see the verdict. Budget a constant per request, plus the sampled fraction of the variable cost. On a host running a dozen containers, that difference decides whether the telemetry path is a rounding error or a co-tenant competing with the application for CPU and outbound bandwidth. ## The one property the kept traces must have Cheap is worthless if what survives is unusable, and head-based sampling has exactly one non-negotiable requirement: **the verdict is reached once and then obeyed everywhere.** The head decides; the decision travels with the request as a flag in the propagated context; every downstream service reads that flag and honours it instead of drawing again. Get that right and a kept trace is complete from the entry point down to the deepest leaf span. Get it wrong — every service rolling its own dice — and you do not get fewer complete traces, you get a store full of disconnected fragments whose parents were dropped upstream, and you pay to keep them. ## What it can never do - It cannot keep a trace *because* it failed, was slow, or took a rare code path. Those facts do not exist when the decision is taken. - Rare events survive only in proportion to the ratio. At one in two hundred you keep half a percent of your errors too, so a failure occurring forty times a day may never appear in the store. - It can, however, be biased with what *is* known at the head: a higher ratio for a business-critical route, a floor for a rarely-called endpoint, an always-keep for requests an internal tool has explicitly marked for debugging. That is how teams recover some signal without paying for buffering. That trade — near-zero overhead against blindness to outcome — is the entire reason the alternative strategy exists, and the reason most mature systems end up running both.
- What does the sampler actually see when it runs at the head?The request as it arrives, and nothing else: the incoming context if there is one, the route or operation name, attributes such as tenant or method, and a source of randomness. It does not see the response status, the duration, the number of downstream calls or any error, because none of those exist yet. Every head-based policy you can write is a function of that pre-execution information.
- If the decision precedes the outcome, how do teams still get their failures into the trace store?By biasing the head decision with what is known at the head, and accepting that the rest is lost. Per-route ratios keep a critical booking path at a far higher rate than a health endpoint, and a caller that already suspects the request — a retry, an internal tool setting a debug flag — can force the decision on. What no head decision can do is keep a trace because it later returned an error.
- Does a one-in-a-hundred ratio mean the application does one percent of the tracing work?Close, but not exactly. Context extraction and injection still happen on every request, because a service must be able to continue a caller's trace and must let downstream services see the verdict. What the other ninety-nine percent skip is span construction, attribute recording, batching, serialisation and export, which dominate the cost. Budget a small constant per request plus one percent of the variable cost.
It is like stamping a ticket at the boarding gate: the check happens once, and everyone further along the journey reads the stamp instead of deciding again.
saying these in an interview costs you the question
- Says the decision is taken per span rather than per trace
- Thinks dropped traces are built and then discarded
- Believes head-based sampling can keep traces because they errored
- Lets each service draw its own random number independently
- Assumes a low ratio removes all per-request tracing cost