skip to content

Agents & Workflows

You will learn the agentic layer: ReActAgent and FunctionCallingAgent, exposing query engines as tools through QueryEngineTool and ToolSpec, and the event-driven Workflow API with context passing, checkpointing, and multi-agent orchestration. Interviewers ask because wrapping a query engine as a tool is the cleanest bridge from a fixed RAG pipeline to an agent that decides when to search.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

questions

6

In LlamaIndex, what does QueryEngineTool.from_defaults do and why does its description matter?

level: juniorimportance: must knowfreq 66%

answer

  1. Wraps a query engine as a callable tool
  2. The model only reads metadata
  3. Description is a routing prompt
  4. name plus description plus arg schema
  5. One tool per corpus enables routing

basics

~10 s

QueryEngineTool.from_defaults wraps an existing query engine in LlamaIndex's tool interface with a name and description. That description is the only thing the model reads when deciding whether to route a question to that corpus.

solid answer

~50 s

`QueryEngineTool.from_defaults(query_engine=..., name=..., description=...)` turns a query engine into something an agent can call. The tool exposes a single natural-language input; invoking it runs the underlying query engine and returns a `ToolOutput` whose `content` is the synthesized answer and whose `raw_output` still carries the source nodes. What the LLM actually receives is only the metadata — the name, the description, and the argument schema. The implementation is invisible to it. So the description is a routing prompt, not documentation: it should state which corpus this covers, what kinds of question it answers, and any scope limits like time range or product line. Wrapping several query engines as several tools is the cleanest way to turn a fixed RAG pipeline into an agent that decides *whether* and *where* to search. Vague descriptions such as "useful for questions about documents" are the usual cause of an agent that searches the wrong index or does not search at all.

code

python · 18 lines
python
from llama_index.core import Document, VectorStoreIndex
from llama_index.core.tools import QueryEngineTool
from llama_index.core.agent.workflow import FunctionAgent

index = VectorStoreIndex.from_documents(
    [Document(text="Employees accrue 25 days of paid leave per year.")]
)

hr_tool = QueryEngineTool.from_defaults(
    query_engine=index.as_query_engine(similarity_top_k=3),
    name="hr_policy_search",
    description=(
        "Answers questions about the 2025 employee handbook: paid leave, "
        "expenses and remote-work policy. Input is a full question."
    ),
)

agent = FunctionAgent(tools=[hr_tool], llm="openai/gpt-4o-mini")

go deeper

for a junior

Know the one-liner: QueryEngineTool.from_defaults(query_engine=..., name=..., description=...), and be able to say that the model picks tools purely from name and description.

for a middle

Explain what actually crosses the wire — metadata and argument schema, never the implementation — and show how two well-bounded descriptions turn a fixed pipeline into a routing agent.

for a senior

Be ready to debug bad routing in production: separate retrieval failures from selection failures, keep raw_output for citations, and cap tool count before selection accuracy degrades.

for a principal

Own the tradeoff of agentic retrieval versus a fixed pipeline: each tool call is a full retrieval plus synthesis, so argue the latency and cost budget before letting a model decide when to search.

## The tool contract An agent in LlamaIndex is an LLM plus a list of tools. Every tool carries metadata: a `name`, a `description`, and a schema for its arguments. That metadata is what gets serialized into the model's tool-calling payload (or, for a text-driven ReAct agent, rendered into the prompt). The Python implementation behind the tool is never shown to the model. Everything the model knows about your retrieval stack is the sentence you wrote in `description`. ## Wrapping a query engine A query engine is LlamaIndex's "ask a question over an index, get a synthesized answer" object. `QueryEngineTool.from_defaults(query_engine=qe, name="...", description="...")` adapts it to the tool interface. The generated tool takes one string input, calls the query engine with it, and returns a `ToolOutput`. `ToolOutput.content` is the string the agent sees; `raw_output` holds the underlying response object, so source nodes and scores are still reachable for citations even though the model only reads the text. The fuller form is `QueryEngineTool(query_engine=qe, metadata=ToolMetadata(name=..., description=...))`; `from_defaults` is the convenience constructor. `return_direct=True` makes the agent stop and return the tool's output verbatim instead of feeding it back through another LLM turn — useful when the query engine already produces the final answer and a second synthesis pass would only add cost and drift. ## Naming and describing for routing Treat `name` as an identifier: short, snake_case, meaningful (`hr_policy_search`, not `tool1`). Treat `description` as a prompt fragment. A good one answers three questions: - **What is in this corpus?** ("the 2025 employee handbook") - **What kinds of question does it answer?** ("leave, expenses, remote-work policy") - **What should the input look like?** ("a full natural-language question, not keywords") When two tools cover adjacent corpora, the descriptions must draw the boundary explicitly, because the model has nothing else to disambiguate with. "Product documentation" and "internal engineering wiki" will be confused; "customer-facing setup and troubleshooting docs" versus "internal design decisions and postmortems" will not. ## From fixed pipeline to agent This is the bridge the leaf is named for. A plain RAG pipeline always retrieves. An agent holding two or three `QueryEngineTool`s decides *whether* to retrieve at all, *which* corpus to hit, and *how many times* — it can ask a follow-up query after reading the first answer. The cost is that each tool call is a full retrieval plus synthesis, so an agent that calls three tools costs roughly three RAG queries plus the agent's own reasoning turns. Budget for that before replacing a pipeline with an agent. ## Other tool sources `FunctionTool.from_defaults(fn=my_function)` wraps an ordinary Python function; the docstring and type hints become the description and schema, which is why a bare undocumented function makes a bad tool. A `ToolSpec` is a bundle: you subclass `BaseToolSpec`, list the exposed method names in `spec_functions`, and call `.to_tool_list()` to get a list of tools you can splice into the agent's tool list. Integration packages ship ready-made specs for services like mail or search, which is how you add a whole family of related tools in one line. ## Failure modes worth naming - Descriptions written for humans ("searches the vector index") instead of for routing. - Too many tools: past roughly a dozen, selection accuracy degrades and the tool block eats context. Consolidate, or put a router in front. - The agent forwarding the user's raw question unchanged into a corpus that needs a narrower query — usually fixed by telling it in the description what a good input looks like. - Exceptions inside the query engine surface to the agent as an errored tool result rather than crashing the run, so a broken retriever can look like "the agent gave a vague answer" until you inspect the tool events. - Losing citations because you only kept `ToolOutput.content`; keep `raw_output` if the product needs sources. ## Testing it Call `tool.call("a representative question")` directly, outside any agent, to confirm the retrieval half works. Then test selection separately: give the agent several questions that should route differently and check which tool fired. Those are two different bugs and they are fixed in two different places — the index, or the description.

  • What does return_direct=True change about how the agent handles that tool's output?
    With `return_direct=True`, the agent stops after the tool call and returns the tool's output as the final answer instead of feeding it back to the LLM for another turn. It saves a synthesis round trip and prevents the model from rewriting an already-good answer, but it also means no post-processing, no combining with other tool results, and no chance for the agent to notice the answer was inadequate and query again.
  • How would you keep source citations when a query engine is called through a tool?
    The tool returns a `ToolOutput` whose `content` is the synthesized string but whose `raw_output` still holds the query engine's response object, including `source_nodes` with text and scores. Capture the tool-result events from the run and read `raw_output` there, rather than trying to parse citations out of the agent's final prose.
  • When would you use a ToolSpec instead of building QueryEngineTools by hand?
    A `ToolSpec` bundles several related operations behind one object: you subclass `BaseToolSpec`, list the exposed methods in `spec_functions`, and `.to_tool_list()` yields the tools. It is the right shape when a single service has many actions — read, search, create — that share auth and setup. For a handful of independent corpora, separate `QueryEngineTool`s with hand-written descriptions give you better routing control.

The description is the label on a filing cabinet drawer. Whoever is looking for a document never opens the cabinet to check — they read the label and decide.

saying these in an interview costs you the question

  • Thinks the agent inspects the index or retriever code
  • Writes descriptions for developers instead of for routing
  • Assumes tool selection improves as you add more tools
  • Believes a tool exception crashes the whole agent run
  • Confuses QueryEngineTool with FunctionTool's plain-function wrapper

context

open as a page

In LlamaIndex, when do you pick ReActAgent over FunctionAgent, and what does it cost?

level: middleimportance: must knowfreq 72%

basics

~20 s

FunctionAgent uses the model's native tool-calling API, so it needs a model that supports it. ReActAgent prompts the model to write reasoning and actions as text and parses them, which works with any model but costs more tokens and can fail to parse.

open as a page

How does a LlamaIndex Workflow decide which @step runs next?

level: middleimportance: should knowfreq 58%

basics

~20 s

Nothing declares the order. Each @step method's parameter type says which event it consumes and its return type says which events it emits, so LlamaIndex wires steps by matching event types. A run starts with StartEvent and ends when a step returns StopEvent.

open as a page

How do you resume a long-running LlamaIndex Workflow after the process restarts?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Serialize the run's Context with ctx.to_dict() and store it, then rebuild it with Context.from_dict(workflow, data) and pass it to the next run. WorkflowCheckpointer captures per-step snapshots, but they live in memory unless you persist them yourself.

open as a page

In LlamaIndex AgentWorkflow, how does one agent hand off to another?

level: seniorimportance: should knowfreq 42%

basics

~20 s

AgentWorkflow gives each agent a handoff tool listing the agents named in its can_handoff_to. Calling it transfers control to that agent, which continues with the same shared Context and conversation history rather than starting fresh.

open as a page

When would you replace a LlamaIndex FunctionAgent with a hand-written Workflow?

level: principalimportance: should knowfreq 36%

basics

~20 s

When the control flow is known in advance. A prebuilt agent pays an LLM turn to decide every step; a hand-written Workflow encodes the sequence in typed events and steps, making it cheaper, deterministic and testable — at the price of owning the loop yourself.

open as a page