skip to content

Tracing and Sampling

How distributed tracing works conceptually — spans, propagation, correlation, and the sampling decisions that make it affordable. Interviewers test these mechanics before any OTel- or Jaeger-specific question.

on this pageshow

questions

16

What does a trace ID on every log line buy you, and what makes it go missing?

level: juniorimportance: must knowfreq 66%

answer

  1. one value shared by every line
  2. turns a guess into an equality filter
  3. generated at the edge, carried on the wire
  4. copied from context into each record
  5. lost outside a request, across threads, or unwired

basics

~20 s

A trace ID turns log search into an exact lookup: one filter returns every line every service wrote for one request. It goes missing when no span is active, when context is lost on an async hop, or when the logger never reads it.

solid answer

~50 s

Without a shared identifier, finding one request's logs means guessing at a time window plus some discriminator you hope was logged. A trace ID makes it a **join key**: one equality filter across every service's logs returns that request and nothing else. The ID is generated once at the first component to see the request, carried across process boundaries in a trace-context header, held in a context-local store inside the process, and — the step people forget — **copied out of that store into each log record at write time**. It disappears three ways: work running outside any request context (startup, scheduled jobs, background compaction) has no ID to copy; an asynchronous handoff to a thread pool, callback or queue consumer loses the context unless it is explicitly carried; and a logger or library that was never wired to the context store writes lines that will never carry it.

code

json · 3 lines
json
{"timestamp":"2026-04-11T09:14:22.418Z","level":"WARN","service":"pricing","trace_id":"a3f9c17b5d2e4088","span_id":"7c14b2af","message":"upstream call retried"}
{"timestamp":"2026-04-11T09:14:22.902Z","level":"ERROR","service":"pricing","trace_id":"a3f9c17b5d2e4088","span_id":"7c14b2af","message":"upstream call failed after 2 retries"}
{"timestamp":"2026-04-11T09:14:23.115Z","level":"ERROR","service":"pricing","trace_id":"","span_id":"","message":"async writeback task failed"}

go deeper

for a junior

Be ready to say what the identifier is for: one value on every log line of one request, so a single filter pulls that request's whole story out of many services' logs. Know that it arrives with the request rather than being made up per service.

for a middle

Explain the mechanics end to end — where the ID is created, how it crosses a process boundary, where it lives inside the process, and the separate step that copies it into each log record. Name the asynchronous handoff as the usual place it is lost.

for a senior

Show that you treat completeness as measurable: track the share of records with no identifier, know which work legitimately has none, and explain why a partially stamped corpus misleads an investigation more dangerously than an unstamped one.

for a principal

Own the argument that correlation is a platform guarantee, not a per-team habit: one generation point, one field name across languages and log producers, and a stated position on what background work and third-party log streams carry instead.

A trace ID (often also called a correlation ID or request ID) is a single opaque value that identifies one logical request and is stamped onto every log record produced while handling it, in every process it touches. It is the cheapest correlation mechanism in observability and the one that most reliably pays for itself. ## What the identifier changes about a log search Without it, investigating one request is reconstruction work. You narrow to a time window, then to a service, then guess at a discriminator that somebody hopefully logged — an account number, an order reference, a URL path. Every filter is a hypothesis, and each one silently drops lines that phrased the same fact differently. On a busy service a two-minute window can hold six figures of requests, so the window alone is nearly useless. With the ID present, the search becomes an equality predicate on a single field. Two properties do the work: - **It is global.** Generated once, unchanged across every hop, so the same value selects lines in the gateway, the service that failed, and the three services behind it. - **It is opaque and effectively unique.** That makes it a perfect log field and a terrible metric dimension — putting it on a metric would create one series per request. Correlation identifiers belong on logs and spans, never as a metric label. ## Where the identifier comes from 1. **Generated** at the first component that touches the request — an edge proxy, an API gateway, or the first instrumented service — if the incoming request does not already carry one. 2. **Propagated** across process boundaries: a header on an HTTP call, a record header or message attribute on a broker, a field on an RPC envelope. 3. **Held** inside the process in a context-local store bound to the unit of work, so any code on that path can read it without it being threaded through every method signature. 4. **Copied** out of that store into each log record at write time, by the logging setup. Step 4 is the one that surprises people. A tracing library knowing the ID is not the same thing as the logs carrying it; something has to bridge the two, and if nobody wired that bridge, traces and logs stay two disconnected worlds. ## The three ways it goes missing | Failure mode | What produces it | What you see in the log store | |---|---|---| | **No active request context** | process startup, scheduled jobs, background compaction, retry workers, work that continues after the response was written | lines with the field absent or empty, often the ones explaining the failure | | **Context lost on an asynchronous hop** | submitting to a thread pool, a callback, publishing to and consuming from a queue, a reactive or coroutine boundary | the first lines carry the ID and the rest do not; a consumer's lines belong to a different ID or to none | | **Logger never wired to the context** | a third-party library logging through its own channel, a proxy or sidecar's access log, the runtime's own error stream | whole categories of lines that never carry the field at all | The first is structural: some work genuinely has no originating request, and the honest fix is to give background work its own identifier rather than to fake a request ID. The second is the classic bug, because the handoff usually still works — only the context is lost, so nothing fails loudly. The third is an integration gap that shows up as a blind spot in exactly the layer that saw the network error. ## Why partial stamping is worse than it looks A corpus where most lines carry the ID and some do not produces **silent partial recall**. You filter on the ID, get twenty-seven lines, see a four-second gap between two of them, and conclude the request was blocked. In reality the work in that gap ran on a pool thread that lost the context and wrote forty-three lines you never saw. Absence of evidence reads as evidence of absence, and it reads that way in the middle of an incident when nobody is inclined to question the search. That is why the useful operational habit is to measure completeness rather than assume it: - Track the **proportion of log records with an empty or absent trace field**, per service, as an ordinary metric, and treat a rise in it as a regression. - After adding an asynchronous path, pick a recent request, filter by its ID, and check that lines appear from every stage you expect. - Accept and document the entries that legitimately cannot carry one — anything written before the request is identified, such as connection setup or routing failures. The payoff for getting this right is not just faster grepping. A complete, ID-stamped log corpus is the hop that every other correlation workflow lands on: whatever gets you to a specific request, logs are usually where you finish reading.

  • Background jobs have no incoming request, so what should go in the correlation field for them?
    Give the job its own identifier generated at the start of each run, and record what kind of run it was. Faking a request ID is worse than leaving it empty, because it pollutes a real request's result set. If the job was triggered by a request — a queued task, say — carry the originating ID in a separate field so you can see both the job's own unit of work and what caused it.
  • Why should a trace ID never become a metric dimension, when it is fine as a log field?
    A metric stores one time series per distinct combination of dimension values, so a per-request identifier creates a new series for every request and never sees a second sample. That explodes the series count and the memory of the metrics store while producing nothing queryable. Logs and spans are record-oriented — they store one entry per event — so a unique field costs one field, not one series.
  • Your logs carry the ID but a request's lines still show gaps. How do you tell a lost context from genuinely idle time?
    Compare the log record against another signal for the same request. If a trace exists, its spans cover the gap and name the operation, which proves work happened. Failing that, check whether the gap coincides with a known asynchronous boundary, and look for unstamped lines in the same service within the same millisecond range — if lines exist there with no ID, the context was dropped rather than the request being idle.

It is the order number on every piece of paper in a warehouse: without it you sort by date and hope, with it you pull the whole order in one motion — and the slip that lost its number is the one you never learn about.

saying these in an interview costs you the question

  • Thinks the tracing library having the ID means the logs automatically have it
  • Puts the request identifier on a metric as a dimension
  • Generates a fresh ID in each service instead of reading the incoming one
  • Assumes a filter on the ID always returns every line the request produced
  • Treats missing IDs on background work as a bug rather than an absent request context
open as a page

In distributed tracing, what identifiers does a span carry, and why is a span an interval rather than a point in time?

level: juniorimportance: must knowfreq 76%

basics

~20 s

A span carries a trace id shared by every span of one request, its own span id, and its parent's span id. It stores a start time and a duration because it measures an operation that lasts, not an instant.

open as a page

In head-based trace sampling, where is the keep-or-drop decision made, and why is it nearly free?

level: middleimportance: must knowfreq 62%

basics

~20 s

Head-based sampling decides at the start of a trace, before any span is recorded. The first service rolls the dice once and stamps the verdict onto the outgoing request, so a dropped trace costs nothing downstream.

open as a page

What can an automatic instrumentation agent see inside a running service, and what is it structurally blind to?

level: middleimportance: must knowfreq 62%

basics

~20 s

An automatic instrumentation agent hooks call sites it has a plugin for - web handlers, database drivers, message clients - so it produces spans at library boundaries only. Domain logic between those boundaries stays one unattributed gap in the trace.

open as a page

In distributed tracing, what must travel with an outbound request for the callee to continue the same trace, and what breaks when one hop drops it?

level: middleimportance: must knowfreq 70%

basics

~20 s

The request must carry the trace id, the id of the span making the call, and whether the trace is being recorded. A hop that drops them makes the callee start a brand-new trace, taking every service below it with it.

open as a page

In distributed tracing, why is a trace assembled at read time, and what does a missing intermediate span do to the tree?

level: middleimportance: should knowfreq 48%

basics

~20 s

Each process exports its spans independently, so the backend stores them individually keyed by trace id and joins them into a tree only when someone queries. A missing intermediate span orphans everything beneath it and silently hides that hop's time.

open as a page

In a metric-to-trace-to-log drill-down, what does each hop require and which breaks first?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Metric to trace needs a deliberately emitted example pointer, because an aggregate has no identity. Trace to log needs the trace identifier present in log records and log retention covering the trace. The metric-to-trace hop breaks first, since nothing produces it by accident.

open as a page

In distributed tracing, why do independent per-service sampling decisions break traces, and what does inheriting the caller's decision change?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Independent decisions multiply: five services each keeping a tenth on their own draw leave complete traces essentially never, and what you store is disconnected fragments. Inheriting the caller's verdict makes one decision at the head bind the whole trace.

open as a page

What must a tail-based trace sampler buffer before it can decide, and what does that cost?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Every span of every in-flight trace, held until a wait window expires, because nothing announces that a trace has finished. The cost is memory proportional to span rate times window, plus routing every span of a trace to the same decider.

open as a page

When an automatic instrumentation agent is already running, which code still earns a hand-written span?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Code the agent leaves as an unattributed gap you need explained, plus attributes carrying domain identity and outcome. Instrument decisions and meaningful units of work, not every function - and prefer an attribute over a new span.

open as a page

What creates instrumentation lock-in across a service fleet, and what would you standardise to reduce it?

level: principalimportance: should knowfreq 40%

basics

~20 s

Lock-in lives in what was built on top of the telemetry, not in the code that emits it. Changing the emitting mechanism is a rollout; re-authoring every dashboard and alert rule that hardcodes an attribute name is the real cost.

open as a page

With a request ID on every log line but no tracing, what can you still not learn?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

You can gather every line for one request, but not the call tree, not who called whom, and not where the time actually went. A flat identifier records membership only; parent links and per-operation durations are what tracing adds.

open as a page

How do an in-process instrumentation agent, kernel-level observation and a sidecar proxy differ in what telemetry each can capture?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Three vantage points, three blind spots. An in-process agent sees application objects and can carry trace context; kernel-level observation sees sockets and syscalls in any language but no application meaning; a sidecar proxy sees only traffic that crosses the network.

open as a page

In a distributed trace, a child span starts before its parent and a queued job is nested under the wrong caller. What causes each?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Clock skew explains the first: start times come from each host's own wall clock. The second is that a span's parent records whatever context was active at creation, which stops matching causality once work is queued or pooled.

open as a page

If sampling policies keep only errors and slow traces, how do you still answer whether a request's behaviour is normal?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Keep a small random baseline alongside the policy-driven keeps, and record why each trace was kept so consumers know which ones can be weighted into an estimate. Rates and distributions should come from unsampled metrics, not from the trace store.

open as a page