What does a trace ID on every log line buy you, and what makes it go missing?
answer
- one value shared by every line
- turns a guess into an equality filter
- generated at the edge, carried on the wire
- copied from context into each record
- lost outside a request, across threads, or unwired
basics
~20 sA trace ID turns log search into an exact lookup: one filter returns every line every service wrote for one request. It goes missing when no span is active, when context is lost on an async hop, or when the logger never reads it.
solid answer
~50 sWithout a shared identifier, finding one request's logs means guessing at a time window plus some discriminator you hope was logged. A trace ID makes it a **join key**: one equality filter across every service's logs returns that request and nothing else. The ID is generated once at the first component to see the request, carried across process boundaries in a trace-context header, held in a context-local store inside the process, and — the step people forget — **copied out of that store into each log record at write time**. It disappears three ways: work running outside any request context (startup, scheduled jobs, background compaction) has no ID to copy; an asynchronous handoff to a thread pool, callback or queue consumer loses the context unless it is explicitly carried; and a logger or library that was never wired to the context store writes lines that will never carry it.
code
json · 3 lines{"timestamp":"2026-04-11T09:14:22.418Z","level":"WARN","service":"pricing","trace_id":"a3f9c17b5d2e4088","span_id":"7c14b2af","message":"upstream call retried"}
{"timestamp":"2026-04-11T09:14:22.902Z","level":"ERROR","service":"pricing","trace_id":"a3f9c17b5d2e4088","span_id":"7c14b2af","message":"upstream call failed after 2 retries"}
{"timestamp":"2026-04-11T09:14:23.115Z","level":"ERROR","service":"pricing","trace_id":"","span_id":"","message":"async writeback task failed"}go deeper
Be ready to say what the identifier is for: one value on every log line of one request, so a single filter pulls that request's whole story out of many services' logs. Know that it arrives with the request rather than being made up per service.
Explain the mechanics end to end — where the ID is created, how it crosses a process boundary, where it lives inside the process, and the separate step that copies it into each log record. Name the asynchronous handoff as the usual place it is lost.
Show that you treat completeness as measurable: track the share of records with no identifier, know which work legitimately has none, and explain why a partially stamped corpus misleads an investigation more dangerously than an unstamped one.
Own the argument that correlation is a platform guarantee, not a per-team habit: one generation point, one field name across languages and log producers, and a stated position on what background work and third-party log streams carry instead.
A trace ID (often also called a correlation ID or request ID) is a single opaque value that identifies one logical request and is stamped onto every log record produced while handling it, in every process it touches. It is the cheapest correlation mechanism in observability and the one that most reliably pays for itself. ## What the identifier changes about a log search Without it, investigating one request is reconstruction work. You narrow to a time window, then to a service, then guess at a discriminator that somebody hopefully logged — an account number, an order reference, a URL path. Every filter is a hypothesis, and each one silently drops lines that phrased the same fact differently. On a busy service a two-minute window can hold six figures of requests, so the window alone is nearly useless. With the ID present, the search becomes an equality predicate on a single field. Two properties do the work: - **It is global.** Generated once, unchanged across every hop, so the same value selects lines in the gateway, the service that failed, and the three services behind it. - **It is opaque and effectively unique.** That makes it a perfect log field and a terrible metric dimension — putting it on a metric would create one series per request. Correlation identifiers belong on logs and spans, never as a metric label. ## Where the identifier comes from 1. **Generated** at the first component that touches the request — an edge proxy, an API gateway, or the first instrumented service — if the incoming request does not already carry one. 2. **Propagated** across process boundaries: a header on an HTTP call, a record header or message attribute on a broker, a field on an RPC envelope. 3. **Held** inside the process in a context-local store bound to the unit of work, so any code on that path can read it without it being threaded through every method signature. 4. **Copied** out of that store into each log record at write time, by the logging setup. Step 4 is the one that surprises people. A tracing library knowing the ID is not the same thing as the logs carrying it; something has to bridge the two, and if nobody wired that bridge, traces and logs stay two disconnected worlds. ## The three ways it goes missing | Failure mode | What produces it | What you see in the log store | |---|---|---| | **No active request context** | process startup, scheduled jobs, background compaction, retry workers, work that continues after the response was written | lines with the field absent or empty, often the ones explaining the failure | | **Context lost on an asynchronous hop** | submitting to a thread pool, a callback, publishing to and consuming from a queue, a reactive or coroutine boundary | the first lines carry the ID and the rest do not; a consumer's lines belong to a different ID or to none | | **Logger never wired to the context** | a third-party library logging through its own channel, a proxy or sidecar's access log, the runtime's own error stream | whole categories of lines that never carry the field at all | The first is structural: some work genuinely has no originating request, and the honest fix is to give background work its own identifier rather than to fake a request ID. The second is the classic bug, because the handoff usually still works — only the context is lost, so nothing fails loudly. The third is an integration gap that shows up as a blind spot in exactly the layer that saw the network error. ## Why partial stamping is worse than it looks A corpus where most lines carry the ID and some do not produces **silent partial recall**. You filter on the ID, get twenty-seven lines, see a four-second gap between two of them, and conclude the request was blocked. In reality the work in that gap ran on a pool thread that lost the context and wrote forty-three lines you never saw. Absence of evidence reads as evidence of absence, and it reads that way in the middle of an incident when nobody is inclined to question the search. That is why the useful operational habit is to measure completeness rather than assume it: - Track the **proportion of log records with an empty or absent trace field**, per service, as an ordinary metric, and treat a rise in it as a regression. - After adding an asynchronous path, pick a recent request, filter by its ID, and check that lines appear from every stage you expect. - Accept and document the entries that legitimately cannot carry one — anything written before the request is identified, such as connection setup or routing failures. The payoff for getting this right is not just faster grepping. A complete, ID-stamped log corpus is the hop that every other correlation workflow lands on: whatever gets you to a specific request, logs are usually where you finish reading.
- Background jobs have no incoming request, so what should go in the correlation field for them?Give the job its own identifier generated at the start of each run, and record what kind of run it was. Faking a request ID is worse than leaving it empty, because it pollutes a real request's result set. If the job was triggered by a request — a queued task, say — carry the originating ID in a separate field so you can see both the job's own unit of work and what caused it.
- Why should a trace ID never become a metric dimension, when it is fine as a log field?A metric stores one time series per distinct combination of dimension values, so a per-request identifier creates a new series for every request and never sees a second sample. That explodes the series count and the memory of the metrics store while producing nothing queryable. Logs and spans are record-oriented — they store one entry per event — so a unique field costs one field, not one series.
- Your logs carry the ID but a request's lines still show gaps. How do you tell a lost context from genuinely idle time?Compare the log record against another signal for the same request. If a trace exists, its spans cover the gap and name the operation, which proves work happened. Failing that, check whether the gap coincides with a known asynchronous boundary, and look for unstamped lines in the same service within the same millisecond range — if lines exist there with no ID, the context was dropped rather than the request being idle.
It is the order number on every piece of paper in a warehouse: without it you sort by date and hope, with it you pull the whole order in one motion — and the slip that lost its number is the one you never learn about.
saying these in an interview costs you the question
- Thinks the tracing library having the ID means the logs automatically have it
- Puts the request identifier on a metric as a dimension
- Generates a fresh ID in each service instead of reading the incoming one
- Assumes a filter on the ID always returns every line the request produced
- Treats missing IDs on background work as a bug rather than an absent request context