skip to content

When would you replace a LlamaIndex FunctionAgent with a hand-written Workflow?

level: principalimportance: should knowfreq 36%

answer

  1. Same runtime, different decision owner
  2. Known flow versus discovered flow
  3. Model turns are priced decisions
  4. Steps are unit-testable, agents are not
  5. The hybrid: agents inside a graph

basics

~20 s

When the control flow is known in advance. A prebuilt agent pays an LLM turn to decide every step; a hand-written Workflow encodes the sequence in typed events and steps, making it cheaper, deterministic and testable — at the price of owning the loop yourself.

solid answer

~50 s

The prebuilt agents in llama-index-core 0.14 are themselves Workflows, so this is not a rewrite, it is choosing where the decisions live. `FunctionAgent` delegates control flow to the model: every step costs a reasoning turn, and the path varies run to run. That is worth paying for when what happens next genuinely depends on what was found. When the sequence is known — classify, retrieve, verify, answer — encoding it as `@step` methods and typed events removes the variance. You get deterministic branching, real parallel fan-out, per-step testing, and a graph a reviewer can read. You can still embed an agent inside a single step where judgment is required, which is usually the right hybrid. What you give up is agility and machinery: you own retries, memory, event streaming, and prompt assembly, and adding a capability means an event class and a step rather than appending to a tool list. Decide on control-flow variance, latency and cost budgets, and auditability — not on which construct is newer.

go deeper

for a junior

Know that both are built on the same workflow machinery, and that the difference is whether the model or your code chooses the next step.

for a middle

Explain the concrete gains of explicit steps — deterministic branching, fan-out with joins, unit-testable steps — and the machinery you take on in exchange.

for a senior

Argue from operational evidence: latency variance from model-decided step counts, spend per reasoning turn, and irreversible side effects that need explicit gates rather than an optional tool call.

for a principal

Own the migration strategy: prototype with an agent, instrument which decisions are effectively constant, promote those into steps, and defend the hybrid shape against both dogmas rather than picking a construct.

## The framing This question is not "agent versus workflow" as rival technologies. In llama-index-core 0.14 the agent classes are implemented as Workflows, and `AgentWorkflow` is one too. The real question is: **who decides what happens next — the model, or your code?** A prebuilt agent answers "the model, every time." A hand-written `Workflow` answers "my code, except where I deliberately delegate." ## What the prebuilt agent buys - **Adaptivity.** Unknown or branching task shapes get handled without you enumerating them. - **Speed of iteration.** New capability equals a new tool in the list. No graph edits. - **Less code.** The loop, tool dispatch, memory and streaming events already exist and are tested. The cost is variance in three dimensions that matter to a lead: **latency** (step count is not bounded by design), **spend** (each decision is a priced turn, and the transcript grows), and **behaviour** (two runs on the same input can take different paths, which makes regression testing statistical rather than exact). ## What a hand-written Workflow buys - **Determinism where you want it.** A branch decided by a Python condition on retrieved scores is free, instant and reproducible; the same branch decided by the model is none of those. - **Real parallelism.** Fan out with `ctx.send_event(...)`, run steps concurrently with `@step(num_workers=n)`, and join with `ctx.collect_events(...)`. An agent loop is sequential by construction unless the provider happens to emit parallel tool calls. - **Testability.** A step is an async function from an event to an event. You can unit-test it with a fabricated input event and no model at all. Agent behaviour can only be evaluated end to end. - **Auditability.** The event graph is the specification. In a regulated setting, "the model decided" is a weaker answer than "step three routes to review whenever the confidence field is below the threshold." - **Cost control.** You place the LLM calls. Cheap deterministic steps do the plumbing; expensive calls happen where they earn their keep. The cost is ownership. Retries, timeouts beyond the run-level budget, memory shape, which events you stream to the UI, error compensation — all yours. And the thing teams underestimate: a rigid graph resists product churn. Every new case is a new event type and a new step, reviewed and tested, where the agent version was a one-line tool addition. ## The decision criteria I would actually apply 1. **Is the control flow known?** If you can draw the flowchart and it does not change per input, encode it. If drawing it requires "it depends what the search returns," keep the agent. 2. **What is the latency budget?** A hard p95 is hard to hold when step count is model-decided. 3. **What does a wrong path cost?** Irreversible side effects — payments, emails, writes to systems of record — argue for explicit steps with explicit gates, not a tool the model may call twice. 4. **Who must be able to review it?** If a non-author has to sign off on behaviour, a typed graph reviews far better than a prompt plus a tool list. 5. **How fast is the surface changing?** Early exploration favours the agent; a stabilised flow favours the graph. ## The hybrid is usually the answer The mature shape is a workflow skeleton with agents inside it. Deterministic steps handle intake, routing on structured signals, retrieval, validation and persistence; one or two steps instantiate an agent for the genuinely open-ended part — "figure out which of these six sources answers the question and how." You keep the graph's testability and cost profile while retaining model judgment exactly where the problem is open-ended. Migration usually runs that direction too: prototype with `FunctionAgent`, watch which decisions the model makes identically every time, and promote those into steps. ## What a weak answer sounds like "Workflows are more advanced, so use them for production." That skips the tradeoff. It is equally wrong to say "agents are the future, hand-written control flow is legacy" — control flow that does not need a model to decide it should not pay a model to decide it. State the criteria, name the hybrid, and describe how you would measure the migration: instrument the agent's chosen paths, and promote a decision to code once the distribution collapses to one branch.

  • How do you decide which agent decisions to promote into explicit steps?
    Instrument the running agent: log the tool chosen at each step against the input class. Where the distribution collapses — the same tool is chosen more than nine times in ten for a recognisable input shape — the model is not exercising judgment, it is paying for a lookup. Promote those into a deterministic step and leave the genuinely branching decisions to the agent. That gives you a measured migration rather than a rewrite on instinct.
  • What does the workflow route make harder that the agent route made easy?
    Change. Adding a capability to an agent is appending a tool; adding one to a graph is a new event class, a new step, wiring by annotation, and tests. You also inherit machinery the agent gave you free: retry policy, memory shape, which events you stream to a UI, and compensation when a mid-graph step fails after earlier steps produced side effects. Budget for that before committing a fast-moving surface to a fixed graph.
  • Where does human approval sit more naturally, and why?
    In the workflow. An explicit gate is a step that emits an input-required event and waits for a response, so the approval point is visible in the graph, testable, and impossible for a model to route around. In an agent, approval is a tool the model may or may not call, and its placement depends on prompt adherence — acceptable for advisory checks, not for irreversible actions.

saying these in an interview costs you the question

  • Calls workflows the advanced option and agents the beginner one
  • Treats the choice as a rewrite rather than relocating decisions
  • Ignores that model-decided step counts break latency budgets
  • Overlooks that steps are unit-testable in isolation
  • Never considers agents embedded inside workflow steps

context