When would you replace a LlamaIndex FunctionAgent with a hand-written Workflow?
answer
- Same runtime, different decision owner
- Known flow versus discovered flow
- Model turns are priced decisions
- Steps are unit-testable, agents are not
- The hybrid: agents inside a graph
basics
~20 sWhen the control flow is known in advance. A prebuilt agent pays an LLM turn to decide every step; a hand-written Workflow encodes the sequence in typed events and steps, making it cheaper, deterministic and testable — at the price of owning the loop yourself.
solid answer
~50 sThe prebuilt agents in llama-index-core 0.14 are themselves Workflows, so this is not a rewrite, it is choosing where the decisions live. `FunctionAgent` delegates control flow to the model: every step costs a reasoning turn, and the path varies run to run. That is worth paying for when what happens next genuinely depends on what was found. When the sequence is known — classify, retrieve, verify, answer — encoding it as `@step` methods and typed events removes the variance. You get deterministic branching, real parallel fan-out, per-step testing, and a graph a reviewer can read. You can still embed an agent inside a single step where judgment is required, which is usually the right hybrid. What you give up is agility and machinery: you own retries, memory, event streaming, and prompt assembly, and adding a capability means an event class and a step rather than appending to a tool list. Decide on control-flow variance, latency and cost budgets, and auditability — not on which construct is newer.
go deeper
Know that both are built on the same workflow machinery, and that the difference is whether the model or your code chooses the next step.
Explain the concrete gains of explicit steps — deterministic branching, fan-out with joins, unit-testable steps — and the machinery you take on in exchange.
Argue from operational evidence: latency variance from model-decided step counts, spend per reasoning turn, and irreversible side effects that need explicit gates rather than an optional tool call.
Own the migration strategy: prototype with an agent, instrument which decisions are effectively constant, promote those into steps, and defend the hybrid shape against both dogmas rather than picking a construct.
## The framing This question is not "agent versus workflow" as rival technologies. In llama-index-core 0.14 the agent classes are implemented as Workflows, and `AgentWorkflow` is one too. The real question is: **who decides what happens next — the model, or your code?** A prebuilt agent answers "the model, every time." A hand-written `Workflow` answers "my code, except where I deliberately delegate." ## What the prebuilt agent buys - **Adaptivity.** Unknown or branching task shapes get handled without you enumerating them. - **Speed of iteration.** New capability equals a new tool in the list. No graph edits. - **Less code.** The loop, tool dispatch, memory and streaming events already exist and are tested. The cost is variance in three dimensions that matter to a lead: **latency** (step count is not bounded by design), **spend** (each decision is a priced turn, and the transcript grows), and **behaviour** (two runs on the same input can take different paths, which makes regression testing statistical rather than exact). ## What a hand-written Workflow buys - **Determinism where you want it.** A branch decided by a Python condition on retrieved scores is free, instant and reproducible; the same branch decided by the model is none of those. - **Real parallelism.** Fan out with `ctx.send_event(...)`, run steps concurrently with `@step(num_workers=n)`, and join with `ctx.collect_events(...)`. An agent loop is sequential by construction unless the provider happens to emit parallel tool calls. - **Testability.** A step is an async function from an event to an event. You can unit-test it with a fabricated input event and no model at all. Agent behaviour can only be evaluated end to end. - **Auditability.** The event graph is the specification. In a regulated setting, "the model decided" is a weaker answer than "step three routes to review whenever the confidence field is below the threshold." - **Cost control.** You place the LLM calls. Cheap deterministic steps do the plumbing; expensive calls happen where they earn their keep. The cost is ownership. Retries, timeouts beyond the run-level budget, memory shape, which events you stream to the UI, error compensation — all yours. And the thing teams underestimate: a rigid graph resists product churn. Every new case is a new event type and a new step, reviewed and tested, where the agent version was a one-line tool addition. ## The decision criteria I would actually apply 1. **Is the control flow known?** If you can draw the flowchart and it does not change per input, encode it. If drawing it requires "it depends what the search returns," keep the agent. 2. **What is the latency budget?** A hard p95 is hard to hold when step count is model-decided. 3. **What does a wrong path cost?** Irreversible side effects — payments, emails, writes to systems of record — argue for explicit steps with explicit gates, not a tool the model may call twice. 4. **Who must be able to review it?** If a non-author has to sign off on behaviour, a typed graph reviews far better than a prompt plus a tool list. 5. **How fast is the surface changing?** Early exploration favours the agent; a stabilised flow favours the graph. ## The hybrid is usually the answer The mature shape is a workflow skeleton with agents inside it. Deterministic steps handle intake, routing on structured signals, retrieval, validation and persistence; one or two steps instantiate an agent for the genuinely open-ended part — "figure out which of these six sources answers the question and how." You keep the graph's testability and cost profile while retaining model judgment exactly where the problem is open-ended. Migration usually runs that direction too: prototype with `FunctionAgent`, watch which decisions the model makes identically every time, and promote those into steps. ## What a weak answer sounds like "Workflows are more advanced, so use them for production." That skips the tradeoff. It is equally wrong to say "agents are the future, hand-written control flow is legacy" — control flow that does not need a model to decide it should not pay a model to decide it. State the criteria, name the hybrid, and describe how you would measure the migration: instrument the agent's chosen paths, and promote a decision to code once the distribution collapses to one branch.
- How do you decide which agent decisions to promote into explicit steps?Instrument the running agent: log the tool chosen at each step against the input class. Where the distribution collapses — the same tool is chosen more than nine times in ten for a recognisable input shape — the model is not exercising judgment, it is paying for a lookup. Promote those into a deterministic step and leave the genuinely branching decisions to the agent. That gives you a measured migration rather than a rewrite on instinct.
- What does the workflow route make harder that the agent route made easy?Change. Adding a capability to an agent is appending a tool; adding one to a graph is a new event class, a new step, wiring by annotation, and tests. You also inherit machinery the agent gave you free: retry policy, memory shape, which events you stream to a UI, and compensation when a mid-graph step fails after earlier steps produced side effects. Budget for that before committing a fast-moving surface to a fixed graph.
- Where does human approval sit more naturally, and why?In the workflow. An explicit gate is a step that emits an input-required event and waits for a response, so the approval point is visible in the graph, testable, and impossible for a model to route around. In an agent, approval is a tool the model may or may not call, and its placement depends on prompt adherence — acceptable for advisory checks, not for irreversible actions.
saying these in an interview costs you the question
- Calls workflows the advanced option and agents the beginner one
- Treats the choice as a rewrite rather than relocating decisions
- Ignores that model-decided step counts break latency budgets
- Overlooks that steps are unit-testable in isolation
- Never considers agents embedded inside workflow steps