skip to content

In LlamaIndex, when do you pick ReActAgent over FunctionAgent, and what does it cost?

level: middleimportance: must knowfreq 72%

answer

  1. Both live under agent.workflow, same run interface
  2. Native tool API versus prompted text
  3. Thought / Action / Observation gets parsed
  4. Parse failures cost turns
  5. Any model versus tool-calling model

basics

~20 s

FunctionAgent uses the model's native tool-calling API, so it needs a model that supports it. ReActAgent prompts the model to write reasoning and actions as text and parses them, which works with any model but costs more tokens and can fail to parse.

solid answer

~50 s

In llama-index-core 0.14 both classes live in `llama_index.core.agent.workflow` and share one interface — `agent.run(...)`, streamed events, an optional `Context` — so switching is a one-line change. The difference is how a tool call is expressed. `FunctionAgent` sends tool schemas through the provider's function-calling API and receives structured tool calls back, so arguments arrive already typed and the provider may emit several calls in one turn. It requires an LLM that advertises function calling. `ReActAgent` puts the tool descriptions in the prompt and asks the model to emit Thought / Action / Action Input / Observation text, which the agent parses. That works with any model, including local ones with no tool API, and gives you a readable reasoning trace. The cost is real: more tokens per step, weaker guarantees on argument shape, and parse failures that turn into wasted retry turns. Default to `FunctionAgent`; reach for `ReActAgent` when the model cannot do tool calling or you specifically want visible reasoning.

code

python · 22 lines
python
import asyncio
from llama_index.core.agent.workflow import FunctionAgent, ReActAgent
from llama_index.core.tools import FunctionTool


def add(a: int, b: int) -> int:
    """Add two integers and return the sum."""
    return a + b


tools = [FunctionTool.from_defaults(fn=add)]

structured = FunctionAgent(tools=tools, llm="openai/gpt-4o-mini")
prompted = ReActAgent(tools=tools, llm="ollama/llama3.1")


async def main() -> None:
    print(await structured.run("What is 21 plus 21?"))
    print(await prompted.run("What is 21 plus 21?"))


asyncio.run(main())

go deeper

for a junior

Be able to name both classes and say the one-line difference: native tool-calling API versus reasoning emitted as parsed text.

for a middle

Explain the mechanics on both sides — schema-constrained arguments and possible parallel calls versus a prompted Thought/Action format — and justify defaulting to FunctionAgent.

for a senior

Show production judgment: check the model's function-calling capability at wiring time, recognise parse-failure loops as a format problem rather than a reasoning one, and account for scratchpad token growth.

for a principal

Own the model-and-framework coupling: choosing ReAct to stay portable across self-hosted models is a real strategy with a measurable token and reliability tax, and that tradeoff should be decided deliberately, not by default.

## The two agent classes As of llama-index-core 0.14, the current agent surface is `llama_index.core.agent.workflow`, which exports `FunctionAgent`, `ReActAgent` and `AgentWorkflow`. Both agents are themselves Workflow objects underneath, so they expose the same lifecycle: `handler = agent.run("question")`, `await handler` for the final `AgentOutput`, and `handler.stream_events()` for intermediate events. Older material shows `ReActAgent.from_tools(...)` and worker/runner splits; that is the pre-workflow API and should not be presented as current. ## FunctionAgent: structured tool calling `FunctionAgent` serializes each tool's name, description and JSON argument schema into the provider's tool-calling request. The model responds with a structured tool call — a tool name plus a JSON argument object — which the agent executes and feeds back as a tool message. Consequences: - **Arguments are typed.** The provider constrains output to the schema, so multi-argument tools work reliably. - **Parallel calls are possible** when the provider supports them: one turn can request several tools, and the agent runs them before the next LLM call. - **Fewer tokens.** No reasoning scaffold is spent in the prompt or the completion. - **It requires capability.** The LLM must support function calling; LlamaIndex surfaces this on the model's metadata as `is_function_calling_model`. Point `FunctionAgent` at a model without it and tool use will not work. ## ReActAgent: reasoning as text `ReActAgent` renders the tool list into the system prompt along with a format contract and asks the model to alternate Thought, Action, Action Input and Observation lines. The agent parses each block, executes the named tool, appends the observation, and loops until the model emits a final answer. Consequences: - **Model-agnostic.** Any instruction-following model can drive it — a local Llama or Mistral deployment with no tool API included. - **Visible reasoning.** The Thought lines are actual model output you can stream, log and inspect, which is helpful for debugging and for products that want to show work. - **Token cost.** The format instructions, the reasoning text and the re-sent scratchpad grow every step. On a five-step task that is a large multiplier over the structured path. - **Parse fragility.** A model that drifts from the format — fenced code, an extra prose paragraph, malformed JSON in Action Input — produces a parse error that the agent must feed back as a correction, burning a turn. Smaller models drift more. - **Weaker argument fidelity.** Nothing constrains Action Input to your schema; complex nested arguments are where this bites. ## How to choose Default to `FunctionAgent` when your provider supports tool calling. It is cheaper, faster and more reliable, and the interface is identical, so nothing is locked in. Choose `ReActAgent` when: - The model has no tool-calling API (self-hosted, older, or an OSS checkpoint behind a plain completion endpoint). - You need the chain of thought as a first-class artifact for audit or UI. - You are debugging why an agent picks the wrong tool: reading its Thought lines is often faster than inferring intent from structured calls. Do not choose ReAct because "ReAct is the standard agent pattern." The loop is the same either way — plan, call a tool, read the result, repeat. Only the encoding of the call differs. ## Operational notes Both agents are stateless across `run()` calls unless you pass a `Context`: create one and reuse it (`ctx = Context(agent)`, then `await agent.run("...", ctx=ctx)`) to keep chat history and any tool-written state between turns. Both stream the same event types — token deltas from the LLM, tool-call events, tool-result events — so observability code written against one works against the other. Tool count matters more for ReAct, because every tool's description sits in the prompt on every step. Ten tools that were fine for `FunctionAgent` can noticeably inflate a ReAct run's cost. ## A common trap Engineers benchmark a ReAct agent on a weak model, see it loop or mis-parse, and conclude "agents do not work." Half the time the fix is not prompt engineering: it is switching to a tool-calling model plus `FunctionAgent`, which removes the entire class of format-adherence failures. Conversely, forcing `FunctionAgent` onto a model without the capability produces an agent that answers from parametric memory and never calls a tool — check the model's capability flag before blaming the tools.

  • How would you tell, before running anything, whether a model can back a FunctionAgent?
    LlamaIndex exposes capability on the LLM's metadata — `llm.metadata.is_function_calling_model` is the flag `FunctionAgent` depends on. If it is false, tool schemas never reach a tool API and the agent will answer from parametric memory instead of calling anything. Check it at wiring time and fall back to `ReActAgent` explicitly rather than discovering it as a silent behavioural failure in production.
  • A ReActAgent keeps failing to parse its own Action Input. What do you try first?
    First confirm the model is the problem by running the same tools under a tool-calling model and `FunctionAgent`; if that works, the loop logic is fine and the failure is format adherence. Then reduce the surface: fewer tools, flatter argument schemas, single-string inputs instead of nested JSON, and shorter tool descriptions. Prompt tweaking helps least; simplifying what the model must emit helps most.
  • Does switching between the two classes change how you stream or persist a run?
    No. Both are Workflow subclasses in llama-index-core 0.14 with the same surface: `agent.run(...)` returns a handler, `handler.stream_events()` yields the same LLM-delta, tool-call and tool-result event types, and passing the same `Context` object across calls preserves chat history and tool-written state. Observability and persistence code written against one class works unchanged against the other.

saying these in an interview costs you the question

  • Claims ReAct is more capable rather than differently encoded
  • Thinks FunctionAgent works on any model regardless of tool support
  • Ignores the token cost of the re-sent ReAct scratchpad
  • Says the two classes need different orchestration code
  • Presents ReActAgent.from_tools as the current API

context