skip to content

What must a tail-based trace sampler buffer before it can decide, and what does that cost?

level: seniorimportance: should knowfreq 44%

answer

  1. Nothing tells you a trace ended
  2. Hold every span for a window
  3. Memory is rate times window
  4. All spans of a trace, one decider
  5. Late spans arrive after the verdict

basics

~20 s

Every span of every in-flight trace, held until a wait window expires, because nothing announces that a trace has finished. The cost is memory proportional to span rate times window, plus routing every span of a trace to the same decider.

solid answer

~50 s

A tail decision needs the outcome — did it error, how slow was it — and that exists only when the request is over. Nothing signals it, so the decider groups spans by trace identifier, holds each group for a **decision window** measured from the first span, then judges whatever arrived. That costs three things. **Memory**, roughly span rate times window times span size, sized for peak whatever the keep rate. **Statefulness**: every span of one trace must reach the same instance, so the tier in front routes by trace identifier and any rollout splits in-flight traces. And **late spans**, arriving after the verdict, are either dropped — possibly the child that made the trace slow — or matched against a cached decision that is itself a memory liability. It does not reduce what the application exports: every span still leaves the process.

code

pseudocode · 13 lines
pseudocode
on_span(span):
    trace = buffer.get_or_create(span.trace_id)   # grouped by trace id
    trace.spans.append(span)

every tick:
    for trace in buffer where now - trace.first_seen >= DECISION_WINDOW:
        verdict = evaluate_policies(trace.spans)  # errored? slow? rare attribute?
        decisions.put(trace.id, verdict)          # cache, so late spans can follow
        emit_if_kept(trace)
        buffer.remove(trace.id)

on_span_for_decided_trace(span):
    verdict = decisions.get(span.trace_id)        # absent -> the span is lost

go deeper

for a junior

Know that deciding after the fact means the spans have to be kept somewhere until the request is finished, and that holding them costs memory that deciding up front does not.

for a middle

Explain the window: spans are grouped by trace identifier and held for a fixed period because nothing marks a trace complete, and the buffer size is roughly arrival rate times that period times span size.

for a senior

Demonstrate the operational consequences: a stateful tier routed by trace identifier, in-flight buffers lost on restart, fragments decided independently after a rebalance, and late spans that are dropped or matched to a cached verdict.

for a principal

Own the economics: the decision tier is sized for one hundred percent of span traffic whatever the keep rate, so argue explicitly about what the extra memory, statefulness and egress buy compared with biasing the decision at the head.

## Why anything has to be buffered A tail decision needs facts that exist only once the request has finished: did anything error, how long did the whole thing take, did it touch this tenant, did that unusual code path run. So the decider must hold the trace until it is finished — and nothing tells it when that is. There is no end-of-trace marker in the data. A trace is over when no further spans will arrive, which is not knowable, only guessable. Every implementation therefore approximates the same way: group arriving spans by trace identifier, hold each group for a **decision window** measured from the first span seen, then evaluate the policies on whatever arrived and decide. ## The memory bill Buffered bytes are roughly **span arrival rate x decision window x average span size**, plus the index that groups spans by trace, plus the allocation churn of building and discarding those groups continuously. Take the ferry-timetable booking platform, whose services run twelve containers to a host. Its decision tier receives about 38,000 spans a second, holds each trace for 45 seconds, and its spans average 1.4 kB serialised: 38,000 x 45 x 1.4 kB is about 2.4 GB resident, before per-trace overhead, before any decision cache, and regardless of how few traces are eventually kept. It is a fixed liability that has to be sized for peak, not for average. What happens when it is undersized is worth knowing: - the tier sheds load by evicting the oldest traces and deciding early on partial data, so slow traces — the ones a latency policy exists to catch — are the first casualties; - or it refuses new traces outright, which biases the kept set toward whatever happened to be buffered already; - either way the loss is silent unless the tier's own eviction and drop counters are being watched. ## Horizontal scaling is where it really hurts A decision needs the whole trace, so **every span sharing a trace identifier must reach the same instance**. That single requirement turns what would be a stateless, trivially scalable tier into a stateful one: 1. Whatever sits in front must route by trace identifier — consistent hashing — rather than round-robin or least-connections. 2. Any change to the instance set reshuffles part of the key space. During a rolling deploy or a scale-out, traces in flight are split across two deciders, and each decides on a fragment. 3. Buffers are memory, so an instance that restarts loses everything it was holding. A rollout costs a decision window of telemetry per instance unless it drains first. 4. Load is uneven by trace rather than by span: one enormous trace lands entirely on one instance. ## Late spans A span arriving after its trace has been decided finds nothing to join. There are two honest options and both cost something: - **drop it**, and the trace you stored may be missing the very child that made it slow; - **follow a cached decision**, keeping a map of trace identifier to verdict for a period after the decision — a second memory liability, with its own expiry, that still fails for spans arriving after the entry expires. Long-running traces make this systemic rather than occasional. Batch jobs, queue consumers and anything with a human in the loop routinely outlast any window you can afford, so they are truncated or missed by design, not by accident. ## What you are buying, honestly | | Deciding at the head | Deciding at the tail | | --- | --- | --- | | Decision inputs | pre-execution attributes only | the whole finished trace | | Memory | none | rate x window | | Decision tier | stateless | stateful, trace-id routed | | Spans leaving the application | only kept traces | all of them | | Late data | not applicable | dropped, or cached decision | That fourth row is the one people miss. Deciding at the tail does not reduce what the application exports, what crosses the network, or what the collection tier ingests — all of it is sized for one hundred percent of spans. Only the backend storage, and the queries over it, get smaller. You are buying keep-on-outcome, and paying in memory, statefulness and egress.

  • How do you choose the length of the decision window?
    From the distribution of end-to-end trace durations, not from taste. Pick a window covering a high percentile of the duration of the traffic you care about, and accept that anything longer is judged on partial data. Longer windows buy completeness and cost memory linearly; shorter ones cut memory and quietly truncate slow traces, which are exactly what a latency policy exists to keep. Measure how often a decided trace later receives more spans.
  • What breaks when the decision tier is scaled out or rolled?
    Two things. Buffers live in memory, so an instance that restarts throws away every trace it was holding, costing a window of telemetry per instance in a rolling deploy. And the mapping from trace identifier to instance changes, so traces spanning the change are split: two deciders each see a fragment and each judges it, producing partial traces or duplicate keeps. Draining before shutdown and stable hashing reduce this, they do not remove it.
  • Does deciding at the tail reduce what the application sends?
    No. Every span has to reach whatever makes the decision, so the export path, the network and the collection tier are all sized for one hundred percent of spans; only backend storage and query cost shrink. That is why teams often put a modest head-based ratio in front of a tail policy: the head cuts what is shipped, the tail chooses what is kept from what arrives.

The difference between checking tickets at the gate and holding every passenger in the terminal until the whole sailing has arrived, then deciding whose journey to write down.

saying these in an interview costs you the question

  • Thinks a signal announces that a trace has finished
  • Assumes tail decisions cut what the application exports
  • Round-robins spans across decision instances
  • Ignores spans arriving after the decision window
  • Sizes the buffer without reference to span rate