skip to content

A trip-planning agent recommended a fully booked hotel — how do you localize the fault in its trace?

level: seniorimportance: must knowfreq 52%

answer

  1. green status, wrong answer
  2. did the fact reach the model?
  3. count in versus count out
  4. check tool arguments, not just results
  5. finish reason explains cut-off answers

basics

~20 s

Open the turn's trace and split the question in two: did the availability fact ever reach the model? The retrieval and tool spans answer that; the model-call span's recorded input answers whether the fact was present and ignored.

solid answer

~50 s

Do not start from status — a wrong answer is almost always a trace of successful calls, so the root span will be green. Start from the leaves and binary-search the hop where the truth disappeared. Open the retrieval span: if it returned 8 candidate hotels and only 2 reached the model call, a filter or rerank dropped the available one and this is a retrieval bug, not a model bug. Open the tool span next: if `hotel_availability` was called with the wrong dates or the wrong city, the tool answered a different question correctly. Only if the correct fact is visible in the model call's recorded input is this a reasoning or prompt problem — and then check `finish_reasons` for truncation and where in a large context the fact landed. The discipline is that every hop either preserved the fact or destroyed it, and the trace tells you which one in a single read.

code

json · 15 lines
json
{
  "span": "retrieval: hotel_availability_notes",
  "status": "OK",
  "attributes": {
    "query": "boutique hotel lisbon 12-15 sep 2 adults",
    "candidates.returned": 8,
    "candidates.after_rerank": 2,
    "rerank.score_cutoff": 0.62
  },
  "next_sibling": {
    "span": "model_call",
    "status": "OK",
    "attributes": { "input.documents": 2, "finish_reason": "stop" }
  }
}

go deeper

for a junior

Know that you start from the trace for the failing turn and look at what each step returned, rather than guessing at the prompt.

for a middle

Be able to walk the tree in order and say what each span type tells you: retrieval counts, tool arguments, model-call input and finish reason.

for a senior

Show the hypothesis structure — did the fact reach the model or not — and that you can tell a retrieval or argument bug from a reasoning bug without rewriting prompts speculatively.

for a principal

Own the instrumentation that makes this walk possible at all: which shapes and attributes are mandatory, what the team is allowed not to capture, and how a repeat investigation becomes a saved query rather than a heroic read.

## Why status is useless here The defining property of LLM failures is that they are successful. No exception was thrown, every HTTP call returned 200, the tool answered, the model responded. So the first instinct from classical debugging — filter for errors — finds nothing, and people conclude the trace is not helpful. It is; you just have to read it differently. Instead of asking 'what failed', you ask 'where did the correct information stop being present'. ## The two-way split Every wrong answer of this kind reduces to one of two cases: 1. **The fact never reached the model.** It was not retrieved, it was retrieved and then filtered out, the tool was called with wrong arguments, the tool returned a stale cached answer, or the result was truncated before it was serialised into context. 2. **The fact reached the model and was not used.** It was in the assembled input and the model still recommended the booked hotel — an instruction-following, context-position or truncation problem. These have completely different fixes, and confusing them is the most expensive mistake in LLM debugging: teams rewrite prompts for weeks over what was a reranker cutoff. The trace decides the split in one look, provided the spans record shapes and inputs. ## The walk, span by span **Root run span.** Read the attributes, not the status: session id, turn index, prompt and app version, entry point. If the failure started last Tuesday, the version attribute is what tells you it coincided with a release. **Retrieval span.** The single most informative attribute is the count pair: candidates returned versus candidates that survived into the prompt. Eight hotels in, two out is a complete explanation on its own if the available hotel is among the six dropped. Also check the query text — an agent that searched for the wrong city was never going to be saved by a better reranker. **Tool span.** Look at the arguments the model produced, not just the result. Availability checked for the wrong date range, the wrong occupancy, or a stale cached response are all common, and all look like a healthy span. Check timings too: a suspiciously fast tool call often means a cache hit on data that has since changed. **Model-call span.** If you captured the assembled input for this trace, this is where the question is settled. Was the availability information in the prompt? If yes, and the model contradicted it, you are in case 2. Check `finish_reasons` — output stopped for length means the answer was cut mid-thought. Check how large the input was and where the fact sat within it; a fact buried in the middle of a very large context is materially less likely to be used, and that is a context-construction problem rather than a model defect. **Missing spans.** An absent tool span has two very different causes, and separating them is the first move: either the model never emitted the tool call at all — visible in the preceding model-call span's output, which will contain no tool-use request — or the tool ran somewhere the trace did not reach and its span was never attached. The first is a model/prompt problem; the second is an instrumentation gap that will keep costing you. ## Making the walk fast On a trace with 300 spans, browsing is not a strategy. Three things make it quick. Put outcome attributes on the root — the feature, the verdict if you have one, the number of agent iterations — so you can find suspect traces before opening them. Record the count-in/count-out shape on every retrieval and tool span, since shapes are cheap to store and often sufficient without payloads. And keep the span taxonomy consistent, so you can filter a trace to just its retrieval spans, or just its model calls, instead of reading the tree linearly. ## What you say in an interview The weak answer is 'I would look at the trace'. The strong answer names the hypothesis structure — did the fact reach the model, yes or no — says which span answers it, and admits what the trace cannot tell you: if you did not capture the assembled input for this trace, you cannot resolve the split, and your next move is to reproduce the turn with capture enabled or to look at a sampled sibling trace on the same path. That admission is the difference between someone who has debugged a production LLM system and someone who has read about it.

  • The retrieval span shows the available hotel in its results, and the model still recommended the booked one. What next?
    Move to the model-call span's recorded input and read what was actually assembled: was the availability line present, where did it sit, and did anything else in context contradict it. Large inputs with the fact buried mid-context, or an instruction earlier in the prompt that outranks it, explain most of these. That makes it a context-construction fix — ordering, summarising, or promoting the fact — rather than a model swap.
  • How do you keep this walk fast on a trace with 300 spans?
    Do not read linearly. Filter by span type to see only retrieval or only tool spans, sort leaves by duration and by status, and collapse subagent subtrees until you need them. Most importantly, put outcome-shaped attributes on the root span — feature, iteration count, verdict — so you locate the suspect traces before opening any of them, and so a saved query can pull the whole class.
  • What if the tool span you expect is missing from the trace entirely?
    Separate two causes before anything else. If the preceding model-call span's output contains no tool-use request, the model never called the tool — a prompt or schema problem. If the model did request it but no span exists, the tool executed outside the trace and its span was never attached, which is an instrumentation gap. Treating the second as the first sends you rewriting prompts for a wiring bug.

saying these in an interview costs you the question

  • Assuming a green root span means nothing went wrong
  • Blaming the model before checking what retrieval actually returned
  • Reading application logs when the trace already shows the hop
  • Treating the tool's raw result as identical to what the model saw
  • Ignoring the finish reason when the answer was cut off

context