In Datadog, what must be true of a service's telemetry for its metrics, traces and logs to link up?
answer
- Two different joins, not one
- Three tags have to agree everywhere
- env, service and version on every signal
- Log records need the trace identifier injected
- Tracer environment variables set the same tags
basics
~20 sEvery signal must carry the same env, service and version tags, set on the Agent, the tracing library and the log records. Log records must also carry the active trace and span identifiers to be joined to one request.
solid answer
~50 sDatadog calls the convention **unified service tagging**: the three tags `env`, `service` and `version` are applied to every signal a workload emits, so that a metric, a span and a log line about the same deployment agree on their identity. In practice you set them once through the tracer's environment variables (`DD_ENV`, `DD_SERVICE`, `DD_VERSION`) and mirror them on the Agent's host or container tags, so infrastructure metrics land under the same `service` as the spans. That gets you *service-level* correlation. Joining one **individual request** to its log lines needs more: the log record has to carry the active trace and span identifiers, which the tracing libraries inject as `dd.trace_id` and `dd.span_id` fields on the log, and the log's `service` attribute must match the APM service name. Miss either half and the trace view's log panel comes back empty.
go deeper
Recall that Datadog links signals through tags, and that env, service and version are the three that matter. Knowing that a log needs to identify its service before it can appear next to a trace is enough at this level.
Explain where each of the three tags is set - tracer environment variables, Agent configuration, log records - and that joining an individual request to its logs additionally needs the trace and span identifiers on the log record.
Demonstrate diagnosis. Given an empty log panel on a trace, say what you inspect first and why: a raw log record for a trace identifier field, then the service name spelling, then whether those logs were indexed at all.
Own the convention as policy across teams: how the three tags are injected at deploy time rather than requested in a wiki, what you do about services that predate the standard, and where you accept that correlation will be partial rather than blocking a delivery.
## Two different joins, often confused Datadog's promise is that you can move from a metric, to the request that caused it, to the log lines that request wrote. That promise rests on **two distinct joins**, and candidates who only know one of them get stuck the first time a pivot comes back empty. The first join is **by identity**: this metric, this span and this log all describe the same deployed thing. The second is **by request**: this log line was written while this specific span was active. They are satisfied by different mechanisms, and a setup can pass one and fail the other. ## The identity join: unified service tagging Datadog's convention is three tags applied everywhere: - **`env`** - which environment the telemetry came from, so production and staging never merge. - **`service`** - the unit an on-call engineer thinks in. This is the tag that ties an APM service to the container metrics of the workload running it. - **`version`** - the build. This is the one teams skip and then regret, because it is what turns a latency graph into a before-and-after of a rollout. The subtlety is that these tags have to be set on **each producer separately**, because the producers are different artefacts: 1. The **tracing library** inside the application, normally through `DD_ENV`, `DD_SERVICE` and `DD_VERSION`, which stamp every span. 2. The **Agent**, so that host and container metrics carry the same values rather than only a hostname. 3. The **log records**, so that a log's `service` attribute holds the same string the tracer uses. When all three agree, the platform can offer the service's infrastructure, its request throughput and its logs as facets of one entity. When they disagree by so much as a hyphen - `ferry-booking` in the tracer, `ferry_booking` in the logs - the platform has no way to know they are the same thing, and the correlation silently degrades into two unrelated services. ## The request join: trace identifiers on log records Matching identity is not enough to find *this request's* logs among a service's millions of lines. For that, the log record must carry the trace context. The Datadog tracing libraries do this with **log injection**: while a span is active, the logger's output gains `dd.trace_id` and `dd.span_id` fields, which the log pipeline maps onto the reserved trace identifier so the trace view can query for exactly those lines. What has to hold for that to work: - Injection must be enabled and the logging framework must be one the tracer can hook. Injection typically works by adding to the logger's diagnostic context, so a logger that formats messages by hand and ignores that context emits nothing to join on. - The log must be **structured**, or parsed server side into structured fields, so the identifiers are fields rather than characters buried in a message string. - The log must actually reach a queryable store. A log that was ingested but excluded from an index is not there to be found, which is why a correlation problem and a cost-control decision can look identical from the trace view. - The identifier has to survive asynchronous work. Logs written from a thread pool or a callback after the span's scope ended carry no trace context. ## Diagnosing an empty pivot | Symptom | Likely cause | |---|---| | Service appears twice with slightly different names | The `service` tag disagrees between tracer and logs | | Infrastructure metrics not shown beside the service | Agent-side tags do not mirror the tracer's tags | | Trace view's log panel empty for every request | Log injection off, or logs unstructured | | Empty only for some requests | Logs written outside the active span's scope | | Empty only for older requests | The logs were ingested but never indexed | | Rollout comparison impossible | `version` never set, so every build looks the same | The diagnostic order that works: check one raw log record first. If it has no trace identifier field, the request join is broken and no amount of tag tidying will fix it. If it has one, compare its `service` value character by character with the APM service name. Only then look at retention and indexing. ## Why this is a senior question The mechanics are simple; the difficulty is that the three tags must be enforced across teams who each deploy independently, and that a single service getting it wrong degrades navigation for everyone investigating a request that passes through it. On a 41-service estate the cost of a missing convention is not one broken dashboard, it is that cross-service debugging becomes guesswork exactly when it matters. That is why interviewers ask it as a fleet question rather than a configuration question.
- What does the version tag actually buy you during a rollout?It splits every signal by build, so error rate, latency and log volume can be compared between the version being rolled out and the one it replaces while both are live. Without it a canary and the stable fleet are one undifferentiated series, and the only way to attribute a regression is to correlate by wall-clock time against a deploy record - which fails as soon as two deploys overlap.
- Your logs carry a service name that differs from the APM service name. What breaks, and how do you fix it?The platform treats them as two unrelated entities: the trace view finds no logs, and the service page shows no log volume. The durable fix is to have the same value feed both, normally by setting the tracer's service environment variable and having the logger read the same source. Where you cannot redeploy immediately, a server-side log pipeline processor can remap the attribute onto the correct service name.
- Why do logs written from a background thread often lose their trace identifiers?Injection works from the trace context that is active on the current execution scope. Work handed to a thread pool, a callback or a detached task runs outside that scope unless the context was explicitly propagated, so the logger has nothing to inject. The logs still arrive and still carry the service tags, which is why the failure looks partial: identity correlation works, request correlation does not.
saying these in an interview costs you the question
- Says correlation works automatically once the Agent is installed
- Sets the service tag on the tracer but never on the log records
- Thinks matching timestamps is enough to attach a log to a trace
- Treats the version tag as optional metadata nobody queries
- Assumes unstructured log text can be joined to a request
- Cannot tell a correlation failure from a log that was never indexed