How do you trace a LlamaIndex query end-to-end with Arize Phoenix or LlamaTrace?
answer
- one line, installed before construction
- spans mirror the pipeline stages
- the retrieve span is the first place to look
- rendered prompt, not the template
- traces carry your document text
basics
~20 sInstall the Phoenix callback integration and call set_global_handler("arize_phoenix") before building anything; for hosted LlamaTrace pass its endpoint and an API key header. Every query then emits a span tree covering retrieval and synthesis, with retrieved node text, scores and the exact prompts.
solid answer
~50 sLlamaIndex exposes one-line observability: `from llama_index.core import set_global_handler` then `set_global_handler("arize_phoenix")`, with the `llama-index-callbacks-arize-phoenix` package installed. For the hosted LlamaTrace service you pass the endpoint and supply the API key through the OTLP headers environment variable. The handler must be installed before you construct indexes and query engines, because it is picked up at component construction. What you get is a span tree per query: the query engine span, a retrieve span carrying every returned node's text and similarity score, and synthesis spans carrying the fully rendered prompt and the model's reply, plus latency and token counts. That structure is what makes debugging mechanical — open a bad answer, look at the retrieve span first, and if the supporting fact is not in the retrieved nodes it is a retrieval problem, not a prompt problem. For custom telemetry, `llama_index.core.instrumentation` lets you register your own event and span handlers instead.
code
python · 13 linesimport os
from llama_index.core import Document, VectorStoreIndex, set_global_handler
# Hosted LlamaTrace: key travels in the OTLP headers env var.
os.environ["OTEL_EXPORTER_OTLP_HEADERS"] = f"api_key={os.environ['PHOENIX_API_KEY']}"
set_global_handler("arize_phoenix", endpoint="https://llamatrace.com/v1/traces")
# Everything built AFTER the handler is installed gets traced.
index = VectorStoreIndex.from_documents(
[Document(text="Refunds are accepted within 30 days of purchase.")]
)
index.as_query_engine().query("What is the refund window?")go deeper
Know that one call to set_global_handler with the Phoenix integration turns on tracing, and that it must run before you build the index or you will see empty traces.
Describe what the span tree contains — retrieved nodes with scores, the rendered prompt, the completion, latency and tokens — and how those spans map onto the query pipeline.
Demonstrate the debugging algorithm: read the retrieve span first to split retrieval failures from synthesis failures, then use span latencies to find where the wall-clock time went.
Own the observability policy — sampling rates, retention, who may read traces containing corpus text, and whether tracing lands in a vendor UI or your own OpenTelemetry pipeline via custom instrumentation handlers.
## Turning it on Observability in LlamaIndex is deliberately a one-liner. `set_global_handler("arize_phoenix")`, imported from `llama_index.core`, installs a callback handler globally; the string names an integration that must be installed as a separate package. Arize Phoenix runs locally — you launch it and open a local UI — while LlamaTrace is its hosted counterpart, reached by passing an endpoint and authenticating through the OTLP headers environment variable that the OpenTelemetry exporter reads. The ordering rule catches people out: install the handler **before** you build indexes, retrievers and query engines. Components capture the callback manager when they are constructed, so a handler installed afterwards produces a suspiciously empty trace view and a wasted afternoon. ## What a trace contains One query becomes a tree of spans, and the tree shape mirrors the pipeline: - The top-level query-engine span, with total latency. - An embedding span for encoding the query. - A retrieve span holding the returned nodes — their text, their metadata and their similarity scores. This is the single most valuable payload in the trace. - A rerank span, if a node postprocessor ran, showing what order changed. - One or more synthesis spans, each with the fully rendered prompt actually sent to the model — after template substitution and context stuffing — and the raw completion, with token counts and latency attached. Seeing the rendered prompt matters more than it sounds. Templates, system prompts, chat history and retrieved context are assembled by layers you did not write, and the difference between what you think you sent and what you sent is where a surprising number of bugs live. ## The debugging algorithm Tracing turns "the answer was wrong" into a decision procedure with two branches: 1. Open the retrieve span and ask whether the fact needed to answer correctly is present in any returned node. If it is **not**, the generator never had a chance — the problem is upstream: chunk size cutting the fact in half, too small a top-k, an embedding model that does not separate your domain vocabulary, or a metadata filter excluding the right document. No amount of prompt work fixes it. 2. If the fact **is** present and the answer still missed it, the problem is downstream: the synthesizer's strategy, prompt wording, context ordering, or the model itself. Look at the rendered prompt — often the fact is buried in position seven of ten chunks, or a refine chain overwrote a correct intermediate answer. This two-branch split is the whole reason to run tracing rather than print statements. It stops the team from tuning prompts to fix a retrieval bug. ## Beyond single-query debugging Spans carry latency, so a trace view also answers the operational question — where did the eight seconds go? Typically it is either a slow rerank model, a synthesizer making several sequential LLM calls, or a cold vector store. The span tree shows the sequential structure directly, including which parts could have run concurrently and did not. Phoenix and LlamaTrace also let you attach evaluation results back onto traced spans, which closes the loop: sample production traffic, score it with the faithfulness and relevancy evaluators, and view failing scores alongside the exact retrieval that produced them. That is materially better than a spreadsheet of scores with no context. ## Alternatives and custom handlers For local, service-free debugging, `LlamaDebugHandler` from `llama_index.core.callbacks` records events in-process; `get_llm_inputs_outputs()` returns the prompt/response pairs, which is enough to answer "what did we actually send" without standing anything up. For production telemetry that must land in your own stack, the `llama_index.core.instrumentation` module is the modern extension point: get the dispatcher, and register a `BaseEventHandler` or `BaseSpanHandler` subclass that forwards events wherever you need. That is how you emit into an existing OpenTelemetry pipeline or your own metrics store rather than a vendor UI. ## Operational cautions Traces contain the retrieved document text and the full prompt, so they contain whatever is in your corpus — treat the trace store as data of the same sensitivity as the source documents, and think about retention and access before pointing production at a hosted endpoint. Full-fidelity tracing of every request is also volume you pay for in storage; sampling, with a rule that always keeps traces for low-scoring or errored requests, is the usual compromise.
- You installed the handler but the trace view is empty. What is the first thing you check?Ordering. Components capture the callback manager when they are constructed, so an index or query engine built before `set_global_handler` ran never emits anything. Move the call to the top of your module, before any index or engine is created, and rebuild them. The second check is that the integration package is actually installed — the handler name is resolved by string, so a missing package fails at the wiring layer rather than at the query.
- The trace shows the right chunk was retrieved but the answer still missed the fact. Where do you look next?At the synthesis spans, specifically the rendered prompt. Common causes are position — the fact sat deep in a long stuffed context — and multi-call synthesis, where a refine chain overwrote a correct intermediate answer on a later chunk. Both are visible in the span tree: count the synthesis calls, and read what each one was given. The fixes are reordering or reranking, a shorter context, or a different synthesis strategy.
- What is the risk of sending production traces to a hosted tracing service?Traces carry the retrieved document text and the fully rendered prompt, so they hold whatever your corpus holds — customer records, contracts, internal policy. The trace store therefore inherits the corpus's sensitivity classification, including access control and retention. If that is unacceptable, run the tracing backend yourself, or register a custom span handler through the instrumentation module that redacts payloads before export and keeps only structure, scores and latency.
saying these in an interview costs you the question
- Installing the handler after building the index and seeing nothing
- Debugging a wrong answer by editing prompts before checking retrieval
- Thinking the trace shows the template rather than the rendered prompt
- Tracing every production request with no sampling or retention policy
- Assuming traces are safe to ship anywhere regardless of corpus sensitivity