With a request ID on every log line but no tracing, what can you still not learn?
answer
- membership, not structure
- a set rather than a tree
- no parent link, no call ordering
- timestamps are exposed to clock skew
- no per-request keep decision
basics
~20 sYou can gather every line for one request, but not the call tree, not who called whom, and not where the time actually went. A flat identifier records membership only; parent links and per-operation durations are what tracing adds.
solid answer
~40 sA flat identifier propagated perfectly records exactly one fact: **this line belongs to that request**. What it structurally cannot give you is **structure**. There is no parent link, so you cannot reconstruct which service called which — only that both took part. There are no per-operation durations unless every service logged its own timings, so you cannot say where the time went or separate working from waiting on a dependency. Ordering across hosts falls back to timestamps, which clock skew of a few milliseconds can invert on exactly the fast hops you are judging. Volume control becomes all-or-nothing: with no per-request keep decision, cutting cost means logging less for everything. None of that makes it a bad choice — it is cheap, needs no instrumentation library, and reaches systems a trace cannot.
go deeper
Recall that a shared identifier lets you collect one request's log lines from several services, and that collecting them is not the same as knowing the order or the timing of what happened.
Explain the boundary precisely: membership without structure. No parent link means no call tree, no critical path and no way to separate a service doing work from one waiting on a dependency.
Show you can still operate with only the identifier — reasoning from per-service timings, treating cross-host ordering as approximate — and say when the identifier alone is genuinely the better spend.
Frame the two as complementary and sequence the investment: universal correlated logs first because they reach systems that cannot be instrumented, then biased tracing where call structure actually pays.
Assume the best case for a flat identifier: a request ID is generated at the edge, propagated across every hop including the ones through a message broker, and stamped on every log line in every service. Nothing is lost. This is worth reasoning about precisely, because it is a very common real state — plenty of production estates have exactly this and no tracing — and because the honest boundary between it and tracing is a question interviewers use to see whether a candidate understands what tracing is actually for, rather than treating it as a fancier log field. ## What it buys - **Membership.** One equality filter returns every line every service wrote for that request. That alone solves the dominant day-to-day problem: finding the relevant logs at all. - **Cross-service reach.** The stack trace in one service and the connection error in another become one result set instead of two investigations. - **Cheapness and robustness.** It is one field, needs no instrumentation library, survives a proxy or third-party hop that would strip unknown headers as long as the field is one the hop passes through, and works for any language or runtime that can write a log line. ## What it structurally cannot give you | Question about one slow request | Flat request identifier | Distributed trace | |---|---|---| | Which lines belong to this request? | yes | yes | | Which services participated? | yes | yes | | Which service called which? | **no** | yes, from parent links | | How long did each operation take? | only if each service logged its own timing | yes, per span | | Was a service working or waiting on a dependency? | **no** | yes, from nesting | | Did two calls run in parallel or in sequence? | **no**, timestamps only | yes, from the tree | | Reliable ordering across hosts | timestamps, exposed to clock skew | parent links are skew-immune | | Keep detail for interesting requests only | **no**, volume is all-or-nothing | yes, per-request keep decision | The through-line in the "no" rows is that a flat identifier records **membership without structure**. It is a set, and a trace is a tree. Everything a trace tells you that a set cannot comes from the parent link and the recorded duration on each node. The consequences are specific enough to state. Without parent links you cannot compute a **critical path** — the chain of operations whose durations actually add up to the user-visible latency — so you cannot distinguish a service that is slow from a service that is merely waiting on a slow dependency. Both look identical in a log corpus: lines at the start, lines at the end, a gap in between. Two dependency calls that ran concurrently and two that ran one after another produce the same log pattern, and only one of those shapes can be fixed by making a dependency faster. Clock skew turns the timestamp fallback into a trap rather than merely an approximation. Hosts drift by milliseconds even with disciplined time synchronisation, so reconstructing order from timestamps is trustworthy at the scale of seconds and unreliable at the scale of the fast hops you usually care about. The infuriating symptom is a response line that appears to precede the request line that caused it. The cost profile differs too. Tracing has a per-request keep decision, so an estate can retain complete detail for a small, deliberately biased slice — slow requests, errored requests — and pay almost nothing for the rest. Logs have no equivalent lever that preserves the same property: dropping a percentage of lines destroys the completeness that made the identifier useful in the first place, and dropping by severity keeps the wrong things. ## When a flat identifier is still the right call Presenting this as "correlation IDs are the poor relation" is the wrong reading, and an interviewer will push on it. A flat identifier is the right call when: 1. **The estate is shallow.** Two or three hops, where the call structure is already known from the architecture and a tree tells you little you cannot infer. 2. **Parts of the path cannot be instrumented.** A legacy component, a vendor appliance, or a third-party hop that will never run an instrumentation library can still copy one field forward, so the identifier reaches places a trace cannot. 3. **The budget is genuinely tight.** Consider a vinyl-record marketplace whose telemetry spend was halved and which can afford only a 96-hour horizon for detailed data. Complete, correlated logs across every request often beat a thin uniform trace sample that misses the requests that mattered. The mature position is that the two are complementary rather than alternatives, and that the identifier is the foundation: tracing without correlated logs leaves you knowing where the time went and not what the code said, which is its own kind of half-answer.
- Every service logs its own duration alongside the shared identifier. How close does that get you to a trace?Closer than nothing and short of a trace. You gain per-service timings, so you can rank which service spent the most wall-clock time. You still lack parent links, so you cannot tell whether a service's four seconds were its own work or time blocked on a dependency that also reports four seconds — the durations overlap and you cannot tell by how much. Nesting is what separates self time from waiting.
- Why is ordering by timestamp across services unreliable when it looks perfectly reasonable?Because each host keeps its own clock and they drift by milliseconds even under synchronisation. That is comfortably below the duration of many hops, so lines can be ordered wrongly precisely where the sequence matters. The characteristic symptom is a downstream response appearing to be logged before the upstream request that triggered it, which is a clock artefact rather than a real inversion.
It is a guest list versus a seating plan: the list proves who was at the dinner, but only the plan tells you who was sitting next to whom and who was waiting on whom to pass the salt.
saying these in an interview costs you the question
- Claims a shared identifier gives you the call tree
- Trusts cross-host timestamp ordering at millisecond resolution
- Cannot distinguish a slow service from one waiting on a dependency
- Treats a correlation identifier as a strictly inferior version of a trace
- Thinks sampling log lines gives the same lever as trace sampling