In a New Relic distributed trace, how do you attribute latency when only some services are instrumented?
answer
- Only what arrived can be attributed
- Unattributed time hides inside the caller
- That remainder blends several different causes
- A trace begins at its first instrumented hop
- Aggregate across traces before blaming a service
basics
~20 sNew Relic attributes time only to the spans it received. Work in an uninstrumented hop shows up as unexplained time inside its caller's span, and the platform cannot say whether that was network, queueing, or the missing service's own work.
solid answer
~50 sStart from what arrived. Where New Relic holds both caller and callee it shows the callee's share of the caller's time, and because spans are queryable records, `FROM Span` with a filter and a facet answers which service owns the tail across thousands of requests, not one anecdote. Where a hop is not instrumented you get a gap: a long caller span, short known children, and an unattributed difference. That difference is transit, queueing, the missing service's own work and everything it called, summed. The size of the gap is trustworthy; its composition is not. Two limits are worth stating outright. The trace begins at the first instrumented participant, so anything in front of it is invisible rather than zero. And if a hop does not carry the trace's identity onward, the work behind it appears as unrelated traces - a service's absence is never evidence it was fast.
code
nrql · 5 linesSELECT count(*)
FROM Span
WHERE trace.id = '9a4f2c81b3'
FACET service.name
SINCE 3 hours agogo deeper
Recall that a trace only shows services that reported spans, and that a service missing from the waterfall has not been shown to be fast - it simply did not report anything.
Explain how time is attributed: a parent's duration minus its known children leaves a remainder, and that remainder covers transit, queueing and everything the missing hop did.
Show the method: aggregate span time per service across many traces, bound the gap with client-side timing and the missing hop's own metrics, and say plainly what the data cannot decide.
Own coverage as a budgeted choice. Decide which hops earn instrumentation, publish where latency stays unattributed, and make partial coverage deliberate rather than accidental.
## Start from what actually arrived A distributed trace in New Relic is assembled from the spans the platform received. Everything the trace view can tell you is a statement about those spans, and everything else is inference. That sounds obvious and it is routinely forgotten, because a waterfall diagram is persuasive: it looks like a picture of the request, and it is really a picture of the reporting. What the platform can genuinely reconstruct when it holds both ends of a call: - The wall-clock duration of each span it received and where it sits inside its parent. - The share of a parent's time occupied by children it knows about. - The remainder — the parent's duration minus the time accounted for by known children. - Repetition: the same downstream call made many times inside one parent. - Aggregate behaviour, because spans are queryable records. `FROM Span` with a filter and a facet answers "which service accounts for the tail across thousands of requests", which is a far stronger claim than anything a single waterfall supports. ## The shape of a missing hop An uninstrumented service does not appear as a gap in the diagram. It appears as **unattributed time inside its caller**: a long parent span whose known children are short, with a large remainder nobody claims. That remainder is a sum, not a measurement: - time in transit both ways, - time queued at the callee before work started, - the missing service's own work, - everything the missing service called in turn, - and, if timing sources differ, some measurement error. Nothing in the trace separates those. This is the single most important thing to say out loud: the gap is real, its size is trustworthy, and its *composition* is unknown. | What the trace shows | What it could be | What would settle it | |---|---|---| | Long parent, short known children, big remainder | the missing hop's own work, its downstream calls, queueing, transit | a span from that hop, or its own queue and connection metrics | | A callee's span much shorter than the caller's view of the call | transit, connection setup, or client-side queueing | client-side and server-side spans for the same call | | A trace that begins inside your own services | everything in front: client, edge, gateway | instrumenting the entry point | | Work that appears as separate, unrelated traces | a hop that did not carry the trace's identity onward | the identity surviving that hop | ## Two limits worth naming before someone else does **The trace starts at the first instrumented participant.** If the entry point is not instrumented, the trace's own start time is not the request's start time, and latency in front of it is invisible rather than zero. **A trace you can see may be disconnected from the work you are chasing.** When a hop in the middle does not carry the trace's identity forward, downstream work still produces spans — they simply belong to other traces. The absence of a service from your trace is therefore never evidence that the service was fast. It is evidence that you did not receive spans from it. ## Working the problem Consider a tour-logistics service peaking at 5,400 requests per second whose telemetry budget was just halved, so instrumenting everything is not on the table. A workable method: 1. **Aggregate before accusing.** Query span records for the path, facet by service, and compare per-service time at the percentile you actually care about. The service that dominates the median is frequently not the one that dominates the tail. 2. **Locate the remainders.** Find where the largest unattributed time sits, and on which paths. That is a ranked list of candidates for instrumentation, produced from data rather than from opinion. 3. **Bound the gap cheaply.** You rarely need full instrumentation of the missing hop to make progress. A client-side span around the call, plus that hop's own connection, queue-depth or saturation metrics, is usually enough to split the remainder into transit versus work. 4. **Spend the budget where the uncertainty is expensive.** Instrument the hop sitting inside the largest remainder on the highest-traffic path first. Uncertainty about a hop that contributes two milliseconds is uncertainty you can afford. 5. **Write down what stays unattributed.** A coverage note — these hops report spans, these do not — turns a recurring argument into a known limitation. ## What separates a strong answer Weak answers read the waterfall as truth and blame whichever service is visible at the bottom. Strong answers do three things: they say precisely what the platform received, they name the remainder and refuse to interpret it as one cause, and they move from a single trace to an aggregate before assigning responsibility. Adding the budget dimension — that full coverage is a choice with a price, and that partial coverage should be deliberate rather than accidental — is what makes the answer sound like someone who has run this system rather than read one trace.
- The gap sits between a caller and an uninstrumented proxy. What is the cheapest way to shrink your uncertainty?You rarely need full instrumentation of the missing hop. A client-side span around the call gives you the caller's view of its duration, and the proxy's own connection counts, queue depth and saturation metrics bound how much of the remainder is queueing versus transit versus work. That is usually enough to settle which side of the call to investigate, at a fraction of the ingest cost.
- Why is one slow trace a weak basis for naming the service at fault?It is a single sample of a distribution. The hop that dominates the median is often not the hop that dominates the tail, and one trace cannot distinguish a systemic problem from a coincidence. Aggregate span time per service across many traces on the same path, compare at the percentile you actually care about, and only then assign responsibility.
- Your telemetry budget was just halved. How do you choose which uninstrumented hops to instrument first?Rank by the unattributed time they sit inside, weighted by traffic on that path. Instrument where the uncertainty is expensive - the hop inside the largest remainder on your busiest route - and accept remainders on low-traffic or low-latency paths as a deliberate blind spot. Then write the coverage down, so the gap is a known limitation rather than a recurring argument.
saying these in an interview costs you the question
- Blames the last service that appears in the waterfall
- Reads unattributed time as pure network latency
- Assumes a missing span means the hop was fast
- Draws a conclusion about a service from one trace
- Thinks the trace shows the whole request path by definition