In Jaeger's trace waterfall, how do you tell a genuinely slow service from one merely waiting on a downstream call?
answer
- A parent bar encloses its children
- Subtract the children from the parent
- Client and server spans rarely match
- Mind the gap between two spans
- The dependency graph comes from sampled traces
basics
~20 sCompare a span's duration with the time its children cover. A parent's bar encloses its children, so the slow hop is the one with large self time — duration minus child coverage — not simply the longest bar on the screen.
solid answer
~50 sEach bar in the waterfall is one span, positioned by start time and drawn as long as its duration, indented under its parent. Because a parent's bar **encloses** its children, bar length tells you how long a span was open, not how long that service was busy — so ranking by duration ranks callers above callees. The number that matters is **self time**: the span's duration minus the wall-clock covered by its children. A genuinely slow hop has large self time; a waiting caller has a long bar almost entirely covered by one child. Read the gaps too: before the first child is pre-call work, between children is in-process work, and the gap between a client span and the matching server span is connection-pool wait, queueing at the callee or network — time neither service's own metrics counted. Concurrent children mean you subtract their union, not their sum.
code
text · 5 linesorder-api GET /v1/orders/reserve 318 ms self 11 ms
order-api POST inventory (client) 296 ms self 287 ms <-- the time is here
inventory POST /reserve (server) 9 ms self 9 ms
order-api GET catalogue (client) 11 ms self 2 ms
catalogue GET /variants (server) 9 ms self 9 msgo deeper
Recall what the view shows: one bar per span, indented under its caller, positioned by start time and as long as the span lasted. Knowing that a caller's bar covers its callee's is enough to stop you blaming the top row.
Explain self time and compute it from the picture, including the union rule for concurrent children. An interviewer expects you to read a gap before the first child, a gap between children, and the client-to-server gap correctly.
Show diagnostic judgement on a real trace: attribute a client/server gap to connection-pool wait, queueing or network, know that clock skew is adjusted and marked, and say what the trace cannot tell you without more instrumentation.
Own what has to be true fleet-wide for this reading to work: consistent client and server spans at every boundary, instrumentation on intermediaries, and a sampling policy that keeps enough of the rare paths for the dependency graph to be trusted.
## What the waterfall actually draws Jaeger's trace view is a set of nested bars. Each bar is one span; its horizontal position is the span's start time relative to the trace's start, and its length is the span's duration. Bars are indented under the span they reference as parent, so the picture is a call tree laid over a timeline. The consequence people miss: **a parent's bar covers its children's bars.** A service that spends 296 ms of a 318 ms request blocked on one downstream call has a 296 ms bar, exactly like a service that spent 296 ms computing. The bar length alone tells you how long the span was *open*, not how long that service was *busy*. Ranking spans by duration therefore ranks callers above callees and tells you almost nothing. ## Self time is the whole trick The number you actually want is **self time** (also called exclusive time): the span's duration minus the wall-clock time covered by its children. Read it off the picture by looking for the parts of a parent's bar that no child bar sits under. - **A genuinely slow hop** has large self time. Its own bar is long and its children are short or absent — the work happened here. - **A caller merely waiting** has small self time. Its bar is long but almost entirely covered by one child; the time belongs further down. - **A gap before the first child** is work the service did before making its first outbound call: deserialization, validation, an in-process cache miss, or waiting for a thread from a pool. - **A gap between two children** is in-process work between calls, and is often where an accidental N+1 pattern shows up as a staircase of short bars. - **Overlapping children** mean concurrent calls, so self time is the parent's duration minus the *union* of the children's intervals, not minus their sum. Adding the sum is a classic misreading that makes self time look negative. Take a request against a seed-catalogue ordering service with a 320 ms p99 budget. The root span is 318 ms. One child, the outbound call to inventory, is 296 ms. Under it, inventory's own server-side span is 9 ms. Inventory is not slow. Two hundred and eighty-seven milliseconds went somewhere between the caller starting the call and the callee starting to handle it. ## The client/server gap, and what lives in it That gap is the single most valuable thing the waterfall shows, because it is invisible to both services' own metrics. A client-side span and the matching server-side span measure the same call from opposite ends, and the difference between them contains everything neither service counted: | Where the time went | How it shows in the waterfall | |---|---| | Waiting for a connection from a pool | Long gap before the server span begins, no network activity | | DNS, TCP and TLS setup | Gap on the first call to a host, absent on subsequent ones | | Request queued at the callee before a worker picked it up | Server span starts late but is itself short | | Genuine network latency or retransmission | Gap symmetric on request and response sides | | Response body serialization or client-side deserialization | Gap after the server span ends, before the client span closes | None of those is fixed by optimising the callee, which is why "the trace says inventory is slow" is the wrong conclusion and "the trace says 287 ms elapsed outside inventory" is the right one. One caveat that is genuinely Jaeger-specific: spans in one trace are timestamped by different hosts, whose clocks do not agree. Jaeger applies a **clock-skew adjustment** when it assembles a trace, shifting spans so children fall inside their parents, and marks spans it adjusted. A child that still appears to start before its parent, or a negative-looking gap, is a skew artefact rather than a physical impossibility — and on a badly skewed pair of hosts, small gaps like the ones in the table above stop being trustworthy at all. ## Where the service dependency graph comes from Jaeger's dependency graph is not a service registry, a deployment manifest, or anything derived from network traffic. It is computed **from the traces themselves**: wherever a parent span in one service has a child span in another, that is an edge, and the edges are aggregated with call counts over a time window. Three consequences follow directly: 1. **It only shows what was sampled.** An edge exercised rarely may not appear at all if none of its requests were kept, so the graph is a picture of your sampled traffic and not of your architecture. 2. **It only shows what is instrumented.** A hop through an uninstrumented proxy or a service that drops the incoming trace context disappears, and the graph draws a direct edge that does not exist. 3. **On a scaled deployment it needs a job to build it.** The single-binary distribution can compute the graph on the fly from what it holds in memory; production backends need a separate aggregation job to populate a dependencies store. A newly stood-up Jaeger with an empty dependency graph is almost always this, not a bug.
- A client span reads 296 ms and the matching server span reads 9 ms. What is in the difference?Everything neither service measured: waiting for a connection from a pool, DNS, TCP and TLS setup on a first call, the request sitting in the callee's accept queue before a worker picked it up, network latency or retransmission, and serialization at either end. None of it is fixed by optimising the callee, which is why the correct conclusion is that time went outside it, not that it is slow.
- A child span appears to start slightly before its parent. What is going on?Almost certainly clock skew. Spans in one trace are timestamped by different hosts whose clocks disagree, so Jaeger applies a skew adjustment when assembling the trace and marks the spans it moved. Treat it as an artefact rather than a causality violation — and note that on a badly skewed pair of hosts, small client/server gaps stop being trustworthy evidence at all.
- Where does Jaeger's service dependency graph come from, and why is it sometimes empty?It is computed from the traces themselves: a parent span in one service with a child span in another is an edge, aggregated with call counts over a window. It shows only sampled, instrumented traffic, so rare or uninstrumented hops go missing. On a scaled deployment it needs a separate aggregation job to populate a dependencies store — a freshly deployed Jaeger with an empty graph is usually that job not running.
A waterfall is a set of nested stopwatches. The outer one keeps running while the inner ones do, so a long outer reading means the outer call stayed open that long, not that it was doing anything.
saying these in an interview costs you the question
- Picks the longest bar as the slow service
- Adds concurrent children's durations instead of their union
- Ignores the gap between client and server spans
- Treats a skew-adjusted timestamp as a real anomaly
- Thinks the dependency graph comes from a service registry
- Assumes the graph shows every service, not sampled ones