In distributed tracing, why do independent per-service sampling decisions break traces, and what does inheriting the caller's decision change?
answer
- Probabilities multiply across hops
- Fragments, not fewer whole traces
- One verdict, carried on the wire
- The edge sets the rate for everyone
- Rate limit bounds volume, ratio tracks traffic
basics
~20 sIndependent decisions multiply: five services each keeping a tenth on their own draw leave complete traces essentially never, and what you store is disconnected fragments. Inheriting the caller's verdict makes one decision at the head bind the whole trace.
solid answer
~50 sEach service drawing its own random number means a trace survives only if every hop independently says keep, so survival is the ratio raised to the number of hops. You still store the full ratio's worth of spans, but almost all are **fragments** whose parent was dropped one hop upstream. The fix is to decide once, at the head, encode the verdict as a flag in the context that propagates with the request, and have every downstream service honour it instead of re-deciding. Three things follow. The edge now owns the effective rate for everyone downstream. A service that starts its own traces — a job, a consumer — is its own head with its own rate. And an inbound keep flag is an instruction to spend money, so at a trust boundary you normally ignore it and treat the edge as the head.
code
pseudocode · 10 linesdecide(incoming_context, request):
if incoming_context.present and caller_is_trusted:
return incoming_context.sampled # inherit; never re-draw
# no context, or an untrusted caller: this process is the head
if request.route in always_keep_routes:
return true
if not global_rate_limiter.allow():
return false
return deterministic_draw(request.trace_id) < ratio_for(request.route)go deeper
Know that a sampling decision is carried with the request rather than made again in each service, and that a trace is only useful when every service that handled the request agreed to keep it.
Be able to do the arithmetic out loud: independent draws multiply, so five hops at a tenth each leave one complete trace in a hundred thousand, while you still pay to store a tenth of all spans as fragments.
Show the operational consequences of inheriting: the edge owns the rate for everyone, a service can have several heads, and an inbound keep flag from outside the trust boundary is a cost-amplification vector you must not honour.
Own the policy for the estate: where heads are allowed to exist, which routes get their own ratios, whether a global rate limit caps the bill, and how you explain to teams that they cannot raise their own visibility unilaterally.
## Why independent decisions shred traces If every service draws its own random number, a trace survives end to end only when every hop independently says keep. The probabilities multiply. | Services on the path | Each keeps 10% on its own draw | Complete traces | | --- | --- | --- | | 2 | 0.1 x 0.1 | 1 in 100 | | 3 | 0.1 to the third | 1 in 1,000 | | 5 | 0.1 to the fifth | 1 in 100,000 | | 8 | 0.1 to the eighth | 1 in 100,000,000 | The important part is not the small number, it is what you get instead. Each service still keeps its own ten percent, so you export, ingest and store ten percent of all spans in the system. What that ten percent contains is overwhelmingly **fragments**: spans whose parent was dropped one hop up, sitting in the store as roots of nothing with children missing below them. You have paid the full price of a ten percent sample and received a corpus you cannot follow a request through. A ferry-timetable booking platform makes it concrete. A booking request enters a gateway and fans through six services, and each team independently configured an eight percent ratio. Complete traces then occur about 2.6 times in ten million requests; at 4,300 requests a second that is roughly one whole trace every quarter of an hour, against the storage bill of an eight percent sample of everything. When an incident arrived that nobody could explain from the existing dashboards, the store held millions of spans and nothing anyone could read end to end. ## What inheriting the caller's decision changes Parent-based sampling takes the decision once, at the head of the trace, and carries it as a flag in the context propagated with the request. Downstream services read it and act on it rather than drawing again. You get whole traces or nothing, and the bill becomes the ratio times a whole trace — which is what everyone thought they were buying in the first place. Four consequences worth being able to state: 1. **The head owns the rate for everyone.** A downstream team cannot raise its own visibility for traffic that arrives already decided. Its levers are a per-route ratio at the edge, a narrow local always-keep override, or metrics and logs for the questions that only need aggregates. 2. **One service can have several heads.** Traffic from the edge inherits a decision; a scheduled job, or a consumer that starts a fresh trace, is itself a head and decides with its own sampler. Those two paths can run at completely different effective rates inside the same process. 3. **An inbound keep flag is an instruction to spend money.** At a trust boundary, honouring it from an unauthenticated caller lets that caller force you to record and store every one of its requests. The usual policy is to ignore the inbound decision at the edge and treat the edge as the head, then inherit freely inside the boundary. 4. **Consistency needs no communication.** Deriving the verdict from the trace identifier rather than a fresh random number lets two processes reach the same answer without exchanging anything, which is what makes "inherit if present, otherwise decide" safe. ## A ratio and a rate limit are not interchangeable | | Probability ratio | Fixed rate limit | | --- | --- | --- | | Volume during a traffic spike | grows with traffic | flat | | Cost predictability | poor | good | | Relationship to traffic | proportional, can be weighted back up | not proportional | | A rarely-called endpoint | may produce nothing for days | guaranteed trickle | | Bias within a burst | none | favours the earliest arrivals | A ratio yields a statistically usable sample: each kept trace stands for a known number of requests, so counts and distributions can be estimated from it. Its weakness is that the bill scales with traffic and a quiet route can stay invisible. A rate limit yields a bounded bill and a floor per route, which is why it is the common outer guard. Its weakness is subtler: inside any interval, the first arrivals win the budget. During a burst — exactly when things are interesting — the kept set is systematically the requests that arrived before the system saturated, not the ones that queued behind them. A rate-limited sample cannot be weighted back into an unbiased estimate of anything. Real configurations combine the two: per-route ratios for representativeness, a global rate limit as a cost ceiling, and an always-keep for a small set of critical or explicitly flagged requests.
- If downstream services inherit the decision, how does a team raise the sampling rate for just its own service?For traffic that arrives already decided, it cannot: the rate is whatever the head chose. The realistic levers are a per-route ratio at the edge for the routes that reach the service, a narrow always-keep override for an explicit debug condition, or metrics and logs for questions that only need aggregates. A service that starts its own traces, such as a scheduled job, is a head and does set its own rate.
- Should a service honour a keep flag arriving from an external caller?Not blindly. The flag is an instruction to spend money, and an unauthenticated client that always sets it can force you to record and store every request it sends. At a trust boundary the usual policy is to ignore the inbound decision and treat the edge as the head, or to honour it only for authenticated internal callers. Inside the boundary, inheriting is exactly what you want.
- When is a fixed rate limit a better choice than a probability ratio?When you are buying predictable cost or a guaranteed floor rather than a representative sample. A few traces per second per route keeps a traffic spike from multiplying the bill and gives a rarely-called endpoint a steady trickle that a ratio would leave empty for days. The price is that the kept set is no longer proportional to traffic, so it cannot be weighted back up into an unbiased estimate.
Six independent gates along one corridor: the odds of walking the corridor end to end are nothing like the odds of passing a single gate.
saying these in an interview costs you the question
- Thinks per-service decisions merely yield fewer complete traces
- Says the sampled flag is recomputed at every hop
- Assumes a downstream service can raise its own trace rate
- Honours an inbound keep flag from any external caller
- Treats a rate limit and a probability ratio as interchangeable