skip to content

Traces & Spans for LLM Calls

You learn how to model an LLM interaction as a trace: a span per model call, nested inside agent, tool, and retrieval spans, tied together by a session or run id. Once the shape is right, a bad answer becomes a specific span you can open.

on this pageshow

questions

5

How would you model one turn of a tool-using LLM agent as a trace of spans?

level: middleimportance: must knowfreq 58%

answer

  1. one turn, one trace
  2. parent is what caused it
  3. run over agent over model call
  4. tools and retrieval are leaves
  5. flat siblings tell you nothing

basics

~20 s

Make the turn one trace: a root run span, agent spans for each reasoning cycle, and child spans for every model call, tool call and retrieval. Each span's parent is whatever caused it and outlives it.

solid answer

~50 s

One user turn is one trace. The root is a **run span** covering the whole invocation. Under it sit **agent spans** — one per reasoning cycle or delegated subagent — and under those the leaves: a **model-call span** per request to the model, a **tool span** per tool execution, a **retrieval span** per search. A travel-booking turn might give a run span over three tool spans (flight search, seat map, hotel availability), a retrieval span, and four model-call spans, ordered as they happened. The rule for parents is causality plus lifetime: a span's parent is the operation that triggered it and that is still open while it runs. Get the shape right and a bad answer becomes a specific leaf you can open. Get it wrong — every span a flat sibling of the root, or one span per turn — and all you learn is that the turn was slow, not where.

code

json · 17 lines
json
{
  "span": "run: plan_trip_turn",
  "duration_ms": 4130,
  "attributes": { "session.id": "s-8842", "turn": 3, "app.version": "2026.06.1" },
  "children": [
    { "span": "agent: iteration 1", "children": [
        { "span": "model_call", "duration_ms": 610 },
        { "span": "tool: flight_search", "duration_ms": 900 } ] },
    { "span": "agent: iteration 2", "children": [
        { "span": "model_call", "duration_ms": 540 },
        { "span": "tool: seat_map", "duration_ms": 320 },
        { "span": "tool: hotel_availability", "duration_ms": 410 },
        { "span": "retrieval: hotel_notes", "duration_ms": 88 } ] },
    { "span": "agent: iteration 3", "children": [
        { "span": "model_call", "duration_ms": 1180 } ] }
  ]
}

go deeper

for a junior

Know the vocabulary and be able to name the span types in an LLM turn — run, agent, model call, tool, retrieval — and say that a trace is the tree they form for one user turn.

for a middle

Be ready to draw the tree for a concrete agent turn and justify each parent choice by causality and lifetime, including where parallel tool calls and retries go.

for a senior

Show that you design the span shape so failures localize: consistent taxonomy, spans that survive streaming and retries, and enough attributes on the root to find bad traces before you open them.

for a principal

Own the span convention as a cross-team contract. Argue what you standardize versus leave free, how you keep instrumentation cost proportionate, and how you avoid a fleet where every team's traces need bespoke interpretation.

## What a span is here A span is a timed record of one operation: a name, start and end timestamps, a status, key/value attributes, and a pointer to its parent. A trace is the tree those parent pointers form. None of that is LLM-specific. What *is* specific to LLM work is which operations you choose to make into spans, and how you nest them — because an LLM system fails by producing a plausible wrong answer far more often than by throwing, and the only cheap way to localize a plausible wrong answer is to see every hop it passed through. ## The span types worth having - **Run (or invocation) span** — the root of the turn. It covers everything the harness did in response to one user input: all reasoning cycles, all tools, all model calls. Its attributes are the ones you slice on later: session and turn ids, the app or prompt version, the entry point, and a verdict if you have one. - **Agent span** — one reasoning cycle, or one delegated subagent's whole life. It groups the model call that decided what to do with the tool calls that decision produced. In a single-loop agent you get one agent span per iteration; with subagents, a subagent's span is a child of the orchestrator's span and its own subtree hangs beneath it. - **Model-call span** — exactly one request to the model. Not the loop, not the retry sequence: one request, one response. Its attributes carry the model, the operation, sampling settings, the stop/finish reason, and (subject to your capture policy) the assembled input and the output. - **Tool span** — one tool execution, with the arguments the model produced and the result the tool returned, plus a status that distinguishes a tool that failed from a tool that returned an unhelpful answer. - **Retrieval span** — one search. The interesting attribute is the *shape* of the result: how many candidates came back, and how many survived filtering or reranking into the prompt. ## Nesting rules Parent = the operation that caused this one and is still open while it runs. Concretely: - Tool and retrieval spans are children of the agent span whose model call requested them, not children of the model-call span — the model call has already ended by the time the tool runs. - Parallel tool calls are **siblings with overlapping time ranges** under the same agent span. Never chain them parent-to-child to record their order; the timestamps already carry order, and a false chain hides the concurrency. - A retried model call is **two model-call spans** under the same parent, each with its own status, not one span with a retry counter. You want to see the first failure's latency separately. - A streamed model call's span ends when the stream ends (or is cancelled), not at the first token. Record first-token timing as an attribute or event on the span instead; ending the span early makes every latency number wrong. - If a tool runs in another service, its span belongs in this trace — which means the trace context has to reach that service. If it does not, you get an orphan trace and lose the join exactly when you need it. ## Worked example A trip-planning assistant is asked to hold a flight and suggest a hotel. The trace: `run` (4.1 s) → `agent iteration 1` → `model call` (decides to search flights) → `tool: flight_search` (0.9 s); `agent iteration 2` → `model call` → `tool: seat_map` and `tool: hotel_availability` in parallel, plus a `retrieval` span over the hotel notes corpus; `agent iteration 3` → `model call` (writes the answer). Eleven spans, four of them model calls. Reading top-down you can see how many reasoning cycles the turn cost; reading the leaves you can see which single operation was slow or wrong. ## Why the shape is the whole game Two failure modes of bad span design dominate in practice. The first is the **single span per turn**, with the prompt and the answer as attributes: it gives you turn latency and nothing else, so every investigation falls back to log grepping. The second is the **flat trace**, where each model call, tool and retrieval is a direct child of the root: you can see the parts, but not which model call caused which tool call, which is precisely the relationship that explains a wrong action. Both are cheap to instrument and expensive to live with. Deciding the taxonomy once, and enforcing it as a convention across services, is what makes traces from different teams comparable — and what lets a single dashboard answer 'how many reasoning steps does this feature take?' without bespoke parsing.

  • Where do three tool calls issued in parallel from one reasoning step belong in the tree?
    As three sibling spans under the same agent span, with overlapping start and end times. The timestamps already record order and concurrency, so chaining them parent-to-child is wrong: it invents a causal relationship, makes the critical path look serial, and inflates each child's apparent latency. Merging them into one tool span with a count attribute loses the per-tool status you need when only one of them failed.
  • Should a streamed model call's span end at the first token or the last?
    At the last token, or when the stream is cancelled — the span measures the operation, and the operation is not over at first token. Record the time to first token as an attribute or a span event so you keep both numbers. Ending the span at first token systematically understates model latency and hides slow or stalled tails, which is usually the thing you were investigating.
  • How do you represent a model call that failed and was retried twice?
    Three model-call spans under the same agent span, in order, the first two with error status and the third with its own outcome. That preserves each attempt's latency and failure reason and makes the retry cost visible in the tree. Collapsing them into one span with a retry-count attribute hides which attempt was slow and makes the turn look like a single expensive call.

saying these in an interview costs you the question

  • Emitting one span per turn and calling that a trace
  • Making every model call and tool a direct child of the root
  • Wrapping the whole agent loop inside one model-call span
  • Ending a streamed call's span at the first token
  • Hiding retries inside a single span with a counter attribute

context

open as a page

A trip-planning agent recommended a fully booked hotel — how do you localize the fault in its trace?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Open the turn's trace and split the question in two: did the availability fact ever reach the model? The retrieval and tool spans answer that; the model-call span's recorded input answers whether the fact was present and ignored.

open as a page

How do session and run ids relate to trace ids across a multi-turn LLM conversation?

level: middleimportance: should knowfreq 44%

basics

~20 s

Each turn gets its own trace and trace id. A session id, recorded on the spans, groups the turns of one conversation; a run id names one agent invocation, which can cover more than one trace when the invocation is retried.

open as a page

Why follow the OpenTelemetry GenAI semantic conventions when tracing LLM calls?

level: seniorimportance: should knowfreq 37%

basics

~20 s

They fix the attribute names on a model-call span — provider, requested model, token usage, finish reason — so any backend can render and query your LLM spans without per-vendor mapping. As of mid-2026 they are still pre-stable, so pin a version.

open as a page

How do you set a prompt and completion payload capture policy for LLM traces?

level: principalimportance: should knowfreq 29%

basics

~20 s

Capture full prompt and completion payloads on a small sampled share of traces, plus always on errors and flagged turns. Keep attribute-only spans everywhere else, cap payload size, redact on the way out, and retain payloads for less time.

open as a page