skip to content

LlamaIndex

You will learn the data framework for RAG: ingestion and loaders, node parsing and chunking, indexes and vector stores, retrievers and query engines, agentic workflows, and evaluation. Interviewers ask because LlamaIndex is organised around the data pipeline, which forces you to hold opinions on chunking, index choice, and how retrieval quality gets measured.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

In LlamaIndex, what does QueryEngineTool.from_defaults do and why does its description matter?

level: juniorimportance: must knowfreq 66%

answer

  1. Wraps a query engine as a callable tool
  2. The model only reads metadata
  3. Description is a routing prompt
  4. name plus description plus arg schema
  5. One tool per corpus enables routing

basics

~10 s

QueryEngineTool.from_defaults wraps an existing query engine in LlamaIndex's tool interface with a name and description. That description is the only thing the model reads when deciding whether to route a question to that corpus.

solid answer

~50 s

`QueryEngineTool.from_defaults(query_engine=..., name=..., description=...)` turns a query engine into something an agent can call. The tool exposes a single natural-language input; invoking it runs the underlying query engine and returns a `ToolOutput` whose `content` is the synthesized answer and whose `raw_output` still carries the source nodes. What the LLM actually receives is only the metadata — the name, the description, and the argument schema. The implementation is invisible to it. So the description is a routing prompt, not documentation: it should state which corpus this covers, what kinds of question it answers, and any scope limits like time range or product line. Wrapping several query engines as several tools is the cleanest way to turn a fixed RAG pipeline into an agent that decides *whether* and *where* to search. Vague descriptions such as "useful for questions about documents" are the usual cause of an agent that searches the wrong index or does not search at all.

code

python · 18 lines
python
from llama_index.core import Document, VectorStoreIndex
from llama_index.core.tools import QueryEngineTool
from llama_index.core.agent.workflow import FunctionAgent

index = VectorStoreIndex.from_documents(
    [Document(text="Employees accrue 25 days of paid leave per year.")]
)

hr_tool = QueryEngineTool.from_defaults(
    query_engine=index.as_query_engine(similarity_top_k=3),
    name="hr_policy_search",
    description=(
        "Answers questions about the 2025 employee handbook: paid leave, "
        "expenses and remote-work policy. Input is a full question."
    ),
)

agent = FunctionAgent(tools=[hr_tool], llm="openai/gpt-4o-mini")

go deeper

for a junior

Know the one-liner: QueryEngineTool.from_defaults(query_engine=..., name=..., description=...), and be able to say that the model picks tools purely from name and description.

for a middle

Explain what actually crosses the wire — metadata and argument schema, never the implementation — and show how two well-bounded descriptions turn a fixed pipeline into a routing agent.

for a senior

Be ready to debug bad routing in production: separate retrieval failures from selection failures, keep raw_output for citations, and cap tool count before selection accuracy degrades.

for a principal

Own the tradeoff of agentic retrieval versus a fixed pipeline: each tool call is a full retrieval plus synthesis, so argue the latency and cost budget before letting a model decide when to search.

## The tool contract An agent in LlamaIndex is an LLM plus a list of tools. Every tool carries metadata: a `name`, a `description`, and a schema for its arguments. That metadata is what gets serialized into the model's tool-calling payload (or, for a text-driven ReAct agent, rendered into the prompt). The Python implementation behind the tool is never shown to the model. Everything the model knows about your retrieval stack is the sentence you wrote in `description`. ## Wrapping a query engine A query engine is LlamaIndex's "ask a question over an index, get a synthesized answer" object. `QueryEngineTool.from_defaults(query_engine=qe, name="...", description="...")` adapts it to the tool interface. The generated tool takes one string input, calls the query engine with it, and returns a `ToolOutput`. `ToolOutput.content` is the string the agent sees; `raw_output` holds the underlying response object, so source nodes and scores are still reachable for citations even though the model only reads the text. The fuller form is `QueryEngineTool(query_engine=qe, metadata=ToolMetadata(name=..., description=...))`; `from_defaults` is the convenience constructor. `return_direct=True` makes the agent stop and return the tool's output verbatim instead of feeding it back through another LLM turn — useful when the query engine already produces the final answer and a second synthesis pass would only add cost and drift. ## Naming and describing for routing Treat `name` as an identifier: short, snake_case, meaningful (`hr_policy_search`, not `tool1`). Treat `description` as a prompt fragment. A good one answers three questions: - **What is in this corpus?** ("the 2025 employee handbook") - **What kinds of question does it answer?** ("leave, expenses, remote-work policy") - **What should the input look like?** ("a full natural-language question, not keywords") When two tools cover adjacent corpora, the descriptions must draw the boundary explicitly, because the model has nothing else to disambiguate with. "Product documentation" and "internal engineering wiki" will be confused; "customer-facing setup and troubleshooting docs" versus "internal design decisions and postmortems" will not. ## From fixed pipeline to agent This is the bridge the leaf is named for. A plain RAG pipeline always retrieves. An agent holding two or three `QueryEngineTool`s decides *whether* to retrieve at all, *which* corpus to hit, and *how many times* — it can ask a follow-up query after reading the first answer. The cost is that each tool call is a full retrieval plus synthesis, so an agent that calls three tools costs roughly three RAG queries plus the agent's own reasoning turns. Budget for that before replacing a pipeline with an agent. ## Other tool sources `FunctionTool.from_defaults(fn=my_function)` wraps an ordinary Python function; the docstring and type hints become the description and schema, which is why a bare undocumented function makes a bad tool. A `ToolSpec` is a bundle: you subclass `BaseToolSpec`, list the exposed method names in `spec_functions`, and call `.to_tool_list()` to get a list of tools you can splice into the agent's tool list. Integration packages ship ready-made specs for services like mail or search, which is how you add a whole family of related tools in one line. ## Failure modes worth naming - Descriptions written for humans ("searches the vector index") instead of for routing. - Too many tools: past roughly a dozen, selection accuracy degrades and the tool block eats context. Consolidate, or put a router in front. - The agent forwarding the user's raw question unchanged into a corpus that needs a narrower query — usually fixed by telling it in the description what a good input looks like. - Exceptions inside the query engine surface to the agent as an errored tool result rather than crashing the run, so a broken retriever can look like "the agent gave a vague answer" until you inspect the tool events. - Losing citations because you only kept `ToolOutput.content`; keep `raw_output` if the product needs sources. ## Testing it Call `tool.call("a representative question")` directly, outside any agent, to confirm the retrieval half works. Then test selection separately: give the agent several questions that should route differently and check which tool fired. Those are two different bugs and they are fixed in two different places — the index, or the description.

  • What does return_direct=True change about how the agent handles that tool's output?
    With `return_direct=True`, the agent stops after the tool call and returns the tool's output as the final answer instead of feeding it back to the LLM for another turn. It saves a synthesis round trip and prevents the model from rewriting an already-good answer, but it also means no post-processing, no combining with other tool results, and no chance for the agent to notice the answer was inadequate and query again.
  • How would you keep source citations when a query engine is called through a tool?
    The tool returns a `ToolOutput` whose `content` is the synthesized string but whose `raw_output` still holds the query engine's response object, including `source_nodes` with text and scores. Capture the tool-result events from the run and read `raw_output` there, rather than trying to parse citations out of the agent's final prose.
  • When would you use a ToolSpec instead of building QueryEngineTools by hand?
    A `ToolSpec` bundles several related operations behind one object: you subclass `BaseToolSpec`, list the exposed methods in `spec_functions`, and `.to_tool_list()` yields the tools. It is the right shape when a single service has many actions — read, search, create — that share auth and setup. For a handful of independent corpora, separate `QueryEngineTool`s with hand-written descriptions give you better routing control.

The description is the label on a filing cabinet drawer. Whoever is looking for a document never opens the cabinet to check — they read the label and decide.

saying these in an interview costs you the question

  • Thinks the agent inspects the index or retriever code
  • Writes descriptions for developers instead of for routing
  • Assumes tool selection improves as you add more tools
  • Believes a tool exception crashes the whole agent run
  • Confuses QueryEngineTool with FunctionTool's plain-function wrapper

context

open as a page

In LlamaIndex's SentenceSplitter, what do chunk_size and chunk_overlap measure?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Both are token counts, not characters. chunk_size (default 1024) caps how many tokens a node may hold, measured with the tokenizer LlamaIndex is configured with; chunk_overlap (default 200) repeats that many trailing tokens at the start of the next node.

open as a page

In LlamaIndex, what does VectorStoreIndex.from_documents() actually do?

level: juniorimportance: must knowfreq 78%

basics

~20 s

VectorStoreIndex.from_documents() splits each document into nodes, calls the embedding model once per node, and writes those vectors plus the node text into a vector store — by default an in-memory SimpleVectorStore that vanishes when the process exits.

open as a page

What does SimpleDirectoryReader do in LlamaIndex, and how do you control which files it loads?

level: juniorimportance: must knowfreq 68%

basics

~10 s

SimpleDirectoryReader walks a folder, picks a reader based on each file's extension, and returns Document objects with file metadata attached. You narrow what it touches with input_files, required_exts, exclude, exclude_hidden, recursive and num_files_limit.

open as a page

In LlamaIndex, what is the difference between index.as_retriever() and index.as_query_engine()?

level: juniorimportance: must knowfreq 76%

basics

~20 s

In LlamaIndex, as_retriever() returns an object whose retrieve() call gives back scored nodes and makes no LLM call. as_query_engine() wraps a retriever with a response synthesizer, so query() also calls the LLM and returns an answer plus its source nodes.

open as a page

In LlamaIndex, when do you pick ReActAgent over FunctionAgent, and what does it cost?

level: middleimportance: must knowfreq 72%

basics

~20 s

FunctionAgent uses the model's native tool-calling API, so it needs a model that supports it. ReActAgent prompts the model to write reasoning and actions as text and parses them, which works with any model but costs more tokens and can fail to parse.

open as a page

How does document metadata change what LlamaIndex's node parser emits?

level: middleimportance: must knowfreq 50%

basics

~20 s

Metadata is copied onto every node, prepended to the node's text for embedding and for the LLM, and counted against chunk_size — so long metadata shrinks the real text budget and can even make it non-positive. Use excluded_embed_metadata_keys and excluded_llm_metadata_keys to control what leaks where.

open as a page

In LlamaIndex, how do FaithfulnessEvaluator, RelevancyEvaluator and CorrectnessEvaluator differ?

level: middleimportance: must knowfreq 58%

basics

~20 s

FaithfulnessEvaluator checks whether the answer is supported by the retrieved context. RelevancyEvaluator checks whether the retrieved context and answer address the query. CorrectnessEvaluator scores the answer against a labelled reference answer, so only it needs ground truth.

open as a page

When would you choose SummaryIndex or TreeIndex over VectorStoreIndex in LlamaIndex?

level: middleimportance: must knowfreq 70%

basics

~20 s

Choose by the shape of the question. VectorStoreIndex suits targeted lookups over a large corpus. SummaryIndex keeps every node in order and is for whole-document synthesis over a small set. TreeIndex spends LLM calls at build to create a summary hierarchy for large-scope questions.

open as a page

How do you persist a LlamaIndex index and reload it without re-embedding?

level: middleimportance: must knowfreq 72%

basics

~10 s

Call index.storage_context.persist(persist_dir="./storage") to write the docstore, index store and default vector store to disk, then rebuild with StorageContext.from_defaults(persist_dir="./storage") and load_index_from_storage(storage_context). Re-calling from_documents() would re-embed everything.

open as a page

In LlamaIndex, what is the difference between a Document and a Node?

level: middleimportance: must knowfreq 75%

basics

~20 s

A Document is one whole source item as loaded — a file, a page, an API record. Nodes are the pieces produced from it, each with its own id, metadata, embedding, and a stored reference back to the source Document.

open as a page

In LlamaIndex, how do the compact, refine and tree_summarize response modes differ?

level: middleimportance: must knowfreq 72%

basics

~20 s

refine calls the LLM once per node, each call improving the running answer. compact, the default, packs nodes into as few context-sized prompts as possible before refining, so it makes far fewer calls. tree_summarize summarizes batches and then summarizes the summaries, bottom-up.

open as a page

What does similarity_top_k control in a LlamaIndex retriever, and how do you tune it?

level: middleimportance: must knowfreq 70%

basics

~20 s

similarity_top_k is how many nodes a LlamaIndex retriever asks the vector store for, defaulting to 2 on VectorIndexRetriever. Raising it buys recall at the cost of prompt tokens, latency and distraction, so the usual pattern is a wide top_k narrowed by a reranker.

open as a page

How does a LlamaIndex Workflow decide which @step runs next?

level: middleimportance: should knowfreq 58%

basics

~20 s

Nothing declares the order. Each @step method's parameter type says which event it consumes and its return type says which events it emits, so LlamaIndex wires steps by matching event types. A run starts with StartEvent and ends when a step returns StopEvent.

open as a page

What are NodeRelationship.PREVIOUS and NEXT for on LlamaIndex nodes?

level: middleimportance: should knowfreq 42%

basics

~20 s

They are pointers stored in each node's relationships dict that link it to the chunks immediately before and after it in the original document. They let you expand a retrieved chunk back into its surrounding context instead of relying only on chunk_overlap.

open as a page

When would you pick TokenTextSplitter over SentenceSplitter in LlamaIndex?

level: middleimportance: should knowfreq 52%

basics

~20 s

Pick TokenTextSplitter when the text has no reliable sentence structure — logs, code, transcripts, scraped markup — and you want tight, uniform token windows. SentenceSplitter is the better default for prose because it refuses to cut mid-sentence.

open as a page

How does LlamaIndex's BatchEvalRunner run evaluators, and where does it bite at scale?

level: middleimportance: should knowfreq 42%

basics

~20 s

BatchEvalRunner takes a dict of named evaluators and runs them concurrently over many queries with asyncio, capped by its workers setting. It returns a dict keyed by evaluator name holding one EvaluationResult per row. The cost is judge LLM calls: queries times evaluators.

open as a page

How do you write a custom LlamaIndex loader by subclassing BaseReader?

level: middleimportance: should knowfreq 38%

basics

~20 s

Subclass BaseReader and implement load_data(), returning a list of Document objects with text, a metadata dict, and a stable id_ derived from the source system. Use lazy_load_data() to stream large sources instead of materialising everything.

open as a page

How do you resume a long-running LlamaIndex Workflow after the process restarts?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Serialize the run's Context with ctx.to_dict() and store it, then rebuild it with Context.from_dict(workflow, data) and pass it to the next run. WorkflowCheckpointer captures per-step snapshots, but they live in memory unless you persist them yourself.

open as a page

In LlamaIndex AgentWorkflow, how does one agent hand off to another?

level: seniorimportance: should knowfreq 42%

basics

~20 s

AgentWorkflow gives each agent a handoff tool listing the agents named in its can_handoff_to. Calling it transfers control to that agent, which continues with the same shared Context and conversation history rather than starting fresh.

open as a page

What does LlamaIndex's HierarchicalNodeParser produce, and what must you store?

level: seniorimportance: should knowfreq 34%

basics

~20 s

It chunks each document several times at decreasing sizes and links the levels with PARENT and CHILD relationships. You embed only the leaf nodes returned by get_leaf_nodes, but every node of every level must go into the docstore or the parent links dangle.

open as a page

What does LlamaIndex's SemanticSplitterNodeParser cost you at ingest time?

level: seniorimportance: should knowfreq 36%

basics

~20 s

It requires an embed_model and embeds sentence groups across the whole corpus before any node exists, so ingestion becomes an extra full embedding pass with the latency, spend and rate limits that implies. It also gives no hard token ceiling on the chunks it emits.

open as a page

How do you attribute token cost in a LlamaIndex app using TokenCountingHandler?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Attach a TokenCountingHandler to a CallbackManager (usually via Settings.callback_manager), give it a tokenizer matching your model, then read prompt_llm_token_count, completion_llm_token_count and total_embedding_token_count. Counters accumulate until you call reset_counts(), so scope or reset them per request.

open as a page

How do you trace a LlamaIndex query end-to-end with Arize Phoenix or LlamaTrace?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Install the Phoenix callback integration and call set_global_handler("arize_phoenix") before building anything; for hosted LlamaTrace pass its endpoint and an API key header. Every query then emits a span tree covering retrieval and synthesis, with retrieved node text, scores and the exact prompts.

open as a page

What does building a PropertyGraphIndex in LlamaIndex cost versus a VectorStoreIndex?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A VectorStoreIndex build spends one embedding call per node. A PropertyGraphIndex additionally runs LLM extractors over every chunk to pull out entities and relations, so the build is an LLM bill, is far slower, and produces non-deterministic structure that changes between runs.

open as a page

How do you keep a LlamaIndex index in sync when source documents change or are deleted?

level: seniorimportance: should knowfreq 52%

basics

~10 s

Give every document a stable id, then call index.refresh_ref_docs(documents): LlamaIndex compares each document's hash against the docstore and re-inserts only what changed. Removals need an explicit index.delete_ref_doc(doc_id, delete_from_docstore=True).

open as a page

How do you keep LlamaIndex metadata out of embeddings while the LLM still sees it?

level: seniorimportance: should knowfreq 42%

basics

~10 s

Set excluded_embed_metadata_keys on the Document or node to hide keys from the embedded text, and excluded_llm_metadata_keys to hide keys from the prompt. Verify with get_content(metadata_mode=MetadataMode.EMBED) or MetadataMode.LLM.

open as a page

How does LlamaIndex's IngestionPipeline avoid re-embedding unchanged documents?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Two separate mechanisms: an IngestionCache that skips a transformation when the same input hash and transformation were seen before, and an attached docstore that compares each document's id and hash to decide insert, skip or upsert. Both need stable document ids.

open as a page

How do you build hybrid BM25-plus-vector retrieval in LlamaIndex?

level: seniorimportance: should knowfreq 55%

basics

~10 s

Either use a vector store that supports native hybrid search through as_retriever(vector_store_query_mode="hybrid"), or run a BM25Retriever and a vector retriever side by side and merge them with QueryFusionRetriever using mode="reciprocal_rerank".

open as a page

How does a reranker like CohereRerank fit into a LlamaIndex query engine?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Rerankers are node postprocessors: pass CohereRerank or SentenceTransformerRerank in node_postprocessors, and they rescore the retrieved nodes with a cross-encoder and keep only top_n. They run in list order between retrieval and synthesis, and they replace the embedding scores with their own.

open as a page

showing 1–30 of 35