Three microservices each write their own log lines to separate log streams while handling one incoming HTTP request. What has to be threaded through that request for an engineer to later pull up every log line from all three services that belongs to that one request, and how does this relate to (but differ from) distributed tracing?
answer
- correlation ID / request ID header
- MDC / thread-local injection
- structured logging + centralized aggregation
- reuse trace ID as correlation ID
- async boundaries lose the ID
basics
~20 sA shared ID gets attached to a request and passed to every service it touches. If each service logs that ID on every line, you can search for the ID and pull all related lines from every service, in order.
solid answer
~40 sLog correlation generates a correlation/request ID (often the same trace ID used for tracing) at the edge, propagates it through every downstream call as a header, and has each service's logging framework inject it into every log line via MDC/thread-local context. Centralized log aggregation (ELK, Loki, Datadog) then lets you filter by that ID across all services to reconstruct the full narrative of one request, ordered by timestamp. It overlaps with tracing - both need context propagated across hops - but logs capture arbitrary unstructured detail (stack traces, business decisions, variable dumps) at any point in code, while traces capture structured span timing/hierarchy; best practice is reusing the trace ID as the correlation ID so you can pivot directly from a slow span in Jaeger to the exact log lines emitted during it.
go deeper
Should know the idea of a shared ID attached to a request that appears in every service's logs so you can find related lines later.
Should describe propagating the ID via a header and injecting it into log lines via the logging framework's context mechanism, plus using centralized log aggregation to search across services.
Should explain reusing the trace ID as the correlation ID to pivot between tracing and logging tools, and know async or background work can lose the correlation context if not handled explicitly.
Should discuss standardizing correlation/log schema across a polyglot fleet, cost and retention trade-offs of centralized log aggregation at scale, and designing propagation so it survives queues, retries, and fan-out without manual per-service wiring.
## What log correlation is **Log correlation** is the practice of tagging every log line that any service emits while handling a given logical request with a shared identifier, so that after the fact, all of those lines - scattered across separate log files or separate services' log streams - can be retrieved together as the complete narrative of that one request. Mechanically, it starts the same way distributed tracing does: an identifier is generated at the entry point of the system (often literally called a **correlation ID** or **request ID**, or reused directly as the trace ID from distributed tracing) and propagated to every downstream call as a header, most commonly something like `X-Request-ID` or, if reusing tracing infrastructure, the trace ID component of the W3C `traceparent` header. ## How the ID reaches every log line The part specific to logging is what happens inside each service once it has that ID: the service's logging framework needs to attach it to every log line emitted while processing that request, not just log it once. - This is typically done via a mechanism like **MDC** (Mapped Diagnostic Context in Java logging frameworks like Logback/Log4j) or an equivalent thread-local/async-local context in other languages - the ID is set into this context as soon as the request is received, and the logging framework is configured to automatically include whatever's in that context as a field on every subsequent log statement on that thread, without every individual log call needing to pass the ID explicitly. - Logs are then shipped to a centralized aggregation system - the **ELK** stack (Elasticsearch/Logstash/Kibana), **Grafana Loki**, or a hosted platform like **Datadog** - which indexes them. - An engineer can then query "show me every log line, from every service, with correlation ID X," sorted by timestamp, to see the full chronological story of exactly what happened to one request as it moved through the system. ## Why it exists The reason this exists is that in a monolith, "look at the log" for a request usually means one file, one process, in order. In microservices, a single user action might touch five services, each with its own log stream, its own log rotation, possibly its own storage backend, running on different hosts with clocks that aren't perfectly synchronized. Without a shared correlation mechanism, reconstructing what happened to one specific request means manually guessing which lines, across several different files, correspond to the same event - based on approximate timestamps and hoping nothing else was happening concurrently on the same instances. Log correlation removes the guesswork entirely by making the request identity an explicit, queryable, indexed field. ## Related to tracing, but not the same thing It's closely related to, but distinct from, distributed tracing, and the two are usually designed to work together rather than as separate systems. | Mechanism | What it captures | |---|---| | **Tracing** | captures structured, hierarchical timing data - spans, their parent-child relationships, durations - purpose-built for answering "where did the time go." | | **Logs** | capture arbitrary, unstructured or semi-structured detail at any point a developer chose to write a log statement - stack traces, business-logic decisions, variable dumps - things tracing spans generally aren't meant to carry in bulk. | Best practice, and increasingly the default in OpenTelemetry-instrumented systems, is to use the same trace ID as the log correlation ID, so an engineer looking at a slow span in a tracing tool like Jaeger can copy its trace ID and immediately pivot to the exact log lines any service emitted during that span in their log aggregator, without maintaining two separate identifier schemes and a mental mapping between them. ## Losing the ID across an async boundary The most common failure mode is losing the correlation ID across an asynchronous boundary. MDC/thread-local propagation works automatically only as long as processing stays on the same logical call stack; the moment a service hands work off asynchronously, the ID doesn't travel with it unless the application code explicitly reads it from the context and writes it into the outgoing message or the new execution context: - publishing a message to a queue for a separate worker process to pick up later, - spawning a detached background task, - or offloading to a different thread pool without explicitly copying the context. The result is a class of orphaned log lines from background/async work that have no correlation ID at all, breaking the reconstructed narrative exactly at the point where the async fan-out happened, which is often precisely where a bug or delay is hiding. ## Why the logs have to be structured A second, more structural requirement is that logs need to be structured (e.g., JSON lines with named fields, one of which is the correlation ID) rather than free-form text with the ID merely appearing somewhere in a message string. **Structured fields** let the aggregation backend index and filter on the correlation ID as a first-class, exact-match query; unstructured text requires fragile substring or regex search across potentially enormous log volumes, with real risk of false positives or missed matches from inconsistent formatting between services.
- Why is it recommended to reuse the distributed tracing trace ID as the log correlation ID rather than generating a separate one?It lets you pivot directly between systems - see a slow span in Jaeger, copy its trace ID, and jump straight to the matching log lines in your log aggregator, or vice versa - without maintaining two separate ID schemes and mapping between them, which is fragile and adds friction during an incident.
- What breaks log correlation if a service spawns an async background job to process part of the request after responding to the caller?If the correlation ID isn't explicitly carried into the async job's context, the background job's logs lose the link to the original request and show up as orphaned entries with no correlation ID, because thread-local/MDC propagation doesn't happen automatically once processing leaves the original call stack.
- How does structured logging make correlation more useful than plain-text log lines that happen to include the ID as a substring?Structured fields let the log aggregator index and query the correlation ID precisely and efficiently as a first-class filter, whereas plain-text requires fragile regex or substring search across potentially huge volumes and risks false matches or missed lines due to inconsistent formatting.
Like a shared case number stamped on every document filed by different departments handling the same customer complaint - pull the case number and you get every department's paperwork in one folder, even though each department kept its own filing cabinet.
saying these in an interview costs you the question
- thinks each service's local logs are enough to debug a distributed request
- doesn't know MDC/thread-local context can be lost across async boundaries
- no mention of propagating the ID via headers
- conflates correlation ID with distributed tracing entirely, unaware they're complementary
- logs in unstructured free text with no consistent field for the ID