skip to content

In a distributed trace, a child span starts before its parent and a queued job is nested under the wrong caller. What causes each?

level: seniorimportance: nice to knowfreq 26%

answer

  1. Two separate defects, not one
  2. Each timestamp comes from its own host
  3. Durations survive, comparisons do not
  4. The edge records ambient context
  5. Queues and pools break causality

basics

~20 s

Clock skew explains the first: start times come from each host's own wall clock. The second is that a span's parent records whatever context was active at creation, which stops matching causality once work is queued or pooled.

solid answer

~50 s

Two independent defects make a waterfall lie. **Clocks**: each span's start timestamp is read on the machine that produced it, and hosts disagree by milliseconds routinely and by far more when synchronisation fails, so a child can appear to start before its parent. Durations remain sound — they are measured from a monotonic clock inside one process — but cross-machine comparisons are not, and UIs that shift a remote child to fit inside its caller hide the skew *and* any real one-way latency. **Parentage**: the parent reference records what was active at span creation, which is a perfect proxy for causation only in a synchronous call. A consumer processing 190 queued messages under one poll span collapses 190 causes into one; a pooled worker can inherit the previous task's context; fire-and-forget work outlives its parent. Decoupled handoffs deserve a non-hierarchical reference, not a parent-child edge.

code

pseudocode · 6 lines
pseudocode
parent (host A):  start 09:14:02.118   duration 91ms
child  (host B):  start 09:14:02.075   duration 12ms   // 43ms BEFORE its parent

// host B's wall clock runs 43ms behind host A's.
// durations are measured per host from a monotonic clock and remain correct;
// only the cross-host comparison of the two start times is unsound.

go deeper

for a junior

Recall that a trace's start times come from different machines' clocks, so a slightly impossible ordering across services is usually skew rather than a bug in the application you are looking at.

for a middle

Explain why durations stay trustworthy while cross-host comparisons do not, and why a worker picked out of a pool can end up recording a parent that had nothing to do with the work it is running.

for a senior

Show the diagnostic instinct: separate in-process anomalies from cross-process ones, look for a consistent offset across several traces, and refuse to quote a per-hop network time smaller than the skew you can demonstrate.

for a principal

Decide the estate-wide conventions — where async handoffs are modelled as references rather than parent-child edges, what time synchronisation you require of hosts, and which trace-derived numbers teams are allowed to use in reports.

## Cause one: the timestamps come from different clocks Every span's start time is read from the wall clock of the machine that produced it. Two hosts synchronised by a time daemon still disagree by single-digit milliseconds routinely, by tens of milliseconds across zones or on a busy virtualised host, and by seconds when synchronisation has silently failed on one node. Nothing in a trace reconciles those clocks, so the intuitive invariant "a child cannot start before its parent" holds **only within a single process**. The visible symptoms are all arithmetic artefacts of skew: - A child's bar starts 43 ms to the left of its parent's bar. - A server span appears to finish after its client span ended, or to begin before the client sent anything. - The apparent one-way network time between two services comes out negative in one direction and double in the other. Many tracing UIs paper over this by shifting a remote child's timeline so it fits inside the span that called it, assuming the client span brackets the server span. That makes the picture readable and destroys the evidence: **you cannot read true one-way network latency off a waterfall**, because the offset you would be measuring is indistinguishable from the skew that was corrected away. Durations themselves are safe, because a well-built tracer measures elapsed time from a monotonic clock inside one process; it is only cross-machine *comparisons* that are unreliable. ## Cause two: recorded parentage is not always causality The parent reference on a span records *what was active when this span was created*, which is a proxy for causation that holds beautifully for a synchronous call and can fail completely once work is handed off. - **A batch consumer** pulls 190 messages produced by 190 unrelated requests and processes them under a single poll span. Everything it does is nested under that one parent, so 190 causes collapse into one, and the trace claims a relationship that does not exist. - **A pooled worker** may pick up whatever context the thread last held, attributing the new task to whichever request happened to run there previously. - **Fire-and-forget work** outlives the request that scheduled it, producing a child whose end timestamp falls well outside its parent's interval — genuinely, not because of skew. - **Lazy shared initialisation** — a connection pool warming up, a cache being populated — is attributed entirely to the unlucky request that triggered it, which then looks pathologically slow. The modelling answer is that a parent-child edge should mean "the parent is waiting for this", and a decoupled handoff deserves a non-hierarchical reference between the two spans instead, so that the producer and the consumer stay separate traces that can each point at the other. The instrumentation-side answer is that a task must carry the context of whoever scheduled it rather than inherit whatever the worker happened to be holding. ## What a reader can and cannot trust | Reading taken from a trace | Trustworthy? | Why | |---|---|---| | the duration of one span | yes | measured in one process, from a clock that only moves forward | | the order of two spans in the same process | yes | same clock, same thread of control | | comparing timestamps from two machines | only beyond the skew | independent wall clocks, unreconciled | | one-way network latency from client/server offsets | no | skew and UI correction are indistinguishable from it | | a parent-child edge across a synchronous call | yes | the caller demonstrably waits for the callee | | a parent-child edge across a queue or a pool | often not | the edge records ambient context, not causation | ## How to handle a suspicious waterfall 1. Check whether the anomaly is confined to *cross-process* edges. Impossible orderings inside one service point at an instrumentation bug; impossible orderings only between services point at clocks. 2. Compare several traces through the same pair of services. A consistent offset in the same direction is skew on one host; a random one is not. 3. Treat any per-hop network time smaller than the skew you can demonstrate as unmeasurable, and fall back on durations measured within one process. 4. For an edge that crosses a queue or a pool, verify the semantics before believing the shape: ask whether the parent could actually have been waiting for that child. The overall point is that a waterfall looks like a precise, globally-ordered picture of a distributed execution, and it is not. It is a set of locally-measured intervals stitched together with locally-recorded edges, and both the measurement and the edges have known failure modes worth naming out loud before a conclusion is drawn from them.

  • Can you measure one-way network latency by subtracting a server span's start from its client span's start?
    No. That difference is the true one-way time plus the skew between the two hosts, and the skew is frequently the larger term. Worse, many UIs already shift the child to fit inside the parent, so the number you read has been adjusted by an unknown amount. Network timing needs a measurement that stays on one clock.
  • How would you model a producer and a consumer separated by a queue so the trace does not lie?
    Keep them as separate traces joined by a non-hierarchical reference in each direction, rather than making the consumer's work a child of the producer. Parent-child should mean the parent is waiting for the child; a message that sits in a queue for a while, or is processed in a batch alongside unrelated messages, satisfies neither condition.

It is like reconstructing an incident from security cameras in different buildings: each camera's own footage runs at the right speed, but the wall clocks in the corner of each frame disagree, so the order of events across buildings is a guess.

saying these in an interview costs you the question

  • Trusts cross-host timestamps to the millisecond
  • Reads one-way network latency straight off a waterfall
  • Assumes a recorded parent always means causation
  • Blames the tracing backend for impossible orderings
  • Thinks time synchronisation removes skew entirely