skip to content

What must a multi-agent trace record for failure analysis to be possible later?

level: middleimportance: should knowfreq 45%

answer

  1. you cannot re-run your way to the cause
  2. one correlation id across the whole task
  3. spans nested to preserve who called whom
  4. verbatim payloads, not summaries of them
  5. plan sampling, retention and redaction

basics

~20 s

One correlation id spanning the whole task, one span per agent turn nested under the orchestrator, and verbatim inputs, outputs, handoff payloads, tool calls, model identity, token counts and latency. Runs are nondeterministic, so the trace is the only faithful record.

solid answer

~50 s

Treat the trace as a flight recorder: after the fact it is the only thing that survives, because re-running a nondeterministic agentic task does not reproduce the same run. The minimum useful record is a single correlation id for the task, a span per agent turn nested under the orchestrator's span, and inside each span the **verbatim** input context, the output, every tool call with its arguments and result, the handoff payload passed to the next agent, the model identity and version, token counts and latency. The verbatim payloads are the part teams skip and the part attribution needs — without them you can see that an agent ran but not what it decided. Use a vendor-neutral schema rather than a bespoke one: OpenTelemetry's GenAI semantic conventions define agent, workflow, tool and model spans with token and latency metrics, so traces stay portable across frameworks. Then budget for the consequences: traces are large and carry user data, so plan sampling, retention and redaction up front.

code

json · 15 lines
json
{
  "trace_id": "4f2c9e01",
  "span_id": "a17b",
  "parent_span_id": "orchestrator-root",
  "agent": "eligibility",
  "model": "provider-model-v3",
  "input_context": "<verbatim prompt and retrieved documents>",
  "output": "<verbatim agent output>",
  "handoff_payload": {"verdict": "conditionally_eligible", "reasons": ["missing_tax_id"]},
  "tool_calls": [{"name": "lookup_registry", "args": {"id": "G-4471"}, "status": "ok", "ms": 240}],
  "input_tokens": 8140,
  "output_tokens": 612,
  "latency_ms": 3120,
  "status": "completed"
}

go deeper

for a junior

Know that every agent turn should be recorded with its inputs, outputs and tool calls under one id for the whole task, because a rerun will not reproduce the same run.

for a middle

Explain the span hierarchy that preserves which agent invoked which, name the fields attribution needs — especially the verbatim handoff payload — and know that OpenTelemetry's GenAI semantic conventions give a vendor-neutral schema.

for a senior

Demonstrate using traces operationally: replaying a run to find where state diverged, substituting a step's output to test a hypothesis, and mining failed production runs into new evaluation cases.

for a principal

Own tracing as policy: instrument at the orchestration layer so coverage cannot drift, set sampling and retention against real storage cost, and treat the trace store as a regulated data surface because it aggregates more user content than the application does.

## Why tracing is the substrate, not a nice-to-have In a single-model call, the request and response are the whole story. In a multi-agent system, the interesting behaviour lives *between* the agents: what the orchestrator decided to delegate, what context each subagent was given, what it returned, and what the next agent made of it. None of that is visible from the final output. And because the run is nondeterministic — sampling, retrieval results, tool responses and timing all vary — you cannot recover it by executing the task again. Every downstream evaluation activity depends on the trace existing: per-agent scoring on real inputs, first-divergence analysis, counterfactual replay, cost accounting, and mining production runs into eval cases. ## The structure: one task, nested spans The shape that works is hierarchical. A single **trace id** identifies the whole task, from the user's request to the final answer. Under it sits a span for the orchestrator's run, and under that a span per agent turn, with tool calls as child spans of the turn that issued them. Nesting matters because it preserves *who invoked whom* — a flat log of agent turns loses the delegation structure, which is precisely what you need when attributing a failure across a graph of agents. Parallel subagents produce sibling spans under the same parent, each with its own start and end, which is how you later distinguish a genuinely concurrent branch from a sequential one. ## What each span must carry - **Verbatim input context.** The actual prompt and retrieved material the agent saw, not a summary of it. A summary of the input cannot answer "did this agent have the information it needed?" - **Verbatim output.** Including the reasoning surface if the provider exposes one, and the final artifact. - **The handoff payload.** The structured object or message passed to the next agent, recorded at the boundary. This is the single most valuable field for multi-agent attribution and the one most often omitted, because teams log each agent's output but not what was actually forwarded after any post-processing. - **Tool calls.** Name, arguments, result or error, and duration. - **Model identity and version, plus the sampling settings.** Without it you cannot tell whether last week's regression came from a model change or a prompt change. - **Token counts and latency**, per span, so cost and time decompose by agent rather than arriving as one bill. - **Status and termination reason** — completed, budget exhausted, error, cancelled. ## Standardise the schema Bespoke trace formats rot: they are tied to one framework, they break when you change orchestrators, and every tool you want to point at them needs an adapter. OpenTelemetry's GenAI semantic conventions give a vendor-neutral schema covering agent, workflow, tool and model spans with token and latency metrics, and by 2026 it is the common denominator across observability vendors and agent frameworks. Emitting to it means your traces outlive your framework choice and can be read by generic tooling. ## Replay is the payoff A complete trace lets you replay a run like a flight recorder: step through the agent turns in order and watch where the state went wrong. It also enables the two techniques that make attribution rigorous. **Substitution replay** re-executes downstream agents with one step's output replaced by a corrected value, confirming or refuting an attribution hypothesis. **Input-faithful scoring** grades each agent on the input it actually received, which is the only honest per-agent number. Neither is possible without verbatim payloads. Traces are also the best source of new evaluation cases: the failures you find in production are, by construction, the distribution your fixtures were missing. ## The costs you must plan for **Volume.** A multi-agent run carries far more text than a chat turn, and full-fidelity traces of every production run get expensive quickly. The usual policy is to keep everything for failing or flagged runs and sample successful ones, with a short hot-retention window and longer cold storage for the sampled set. **Sensitive data.** Verbatim context contains whatever the user and your retrieval layer put into it — personal data, documents, credentials pasted by mistake. Redact at write time rather than at read time, and treat trace stores as a data-protection surface with real access control, since they aggregate more user content in one place than the application database does. **Instrumentation drift.** If each agent is instrumented by hand, coverage decays as the system changes and you discover the gap exactly when you need the trace. Instrument at the orchestration layer, so every agent turn is captured by construction rather than by a developer remembering. ## The failure mode to recognise The common story is a team with dashboards full of aggregate metrics — success rate, average latency, tokens per run — and no way to answer "what happened in run 4f2c?". Aggregates tell you something is wrong; only a trace tells you what. If the first question in your incident review cannot be answered from the record, the tracing is inadequate no matter how good the dashboards look.

  • Why record the handoff payload separately when you already log each agent's output?
    Because the two are often not the same object. Orchestration layers routinely post-process an agent's output — truncating, summarising, mapping to a schema, or attaching extra context — before handing it on. Attribution needs what the receiving agent actually saw. When only the producer's raw output is stored, a defect introduced in the plumbing between agents is invisible and gets misattributed to one of the models.
  • How do you keep trace volume manageable in production without losing diagnostic value?
    Retain full fidelity for runs that failed, were flagged by a user, or hit a budget or safety guard, and sample successful runs at a low rate for baseline coverage. Keep a short hot window for interactive debugging with cheaper cold storage behind it, and redact sensitive fields at write time so retention policy is not also a data-protection decision made later.
  • What breaks in your evaluation pipeline if agents are instrumented individually rather than at the orchestration layer?
    Coverage becomes uneven and decays over time: a newly added agent, a retry path, or an error branch ships without instrumentation, and the gap surfaces only when a failure lands there. Instrumenting where delegation happens means every agent turn is captured by construction, spans nest correctly, and adding an agent does not require anyone to remember to trace it.

saying these in an interview costs you the question

  • Assuming a failed run can be reproduced by executing it again
  • Logging summaries of agent inputs instead of verbatim context
  • Flat per-agent logs with no correlation id or nesting
  • Storing traces without redaction or a retention policy
  • Relying on aggregate dashboards to explain an individual failed run

context