skip to content

When would you ship a fixed LangChain pipeline instead of create_agent?

level: principalimportance: should knowfreq 45%

answer

  1. is the control flow known in advance?
  2. paying model calls for a decision you already made
  3. latency you cannot put an SLA on
  4. tools the model can reach are the blast radius
  5. bounded agent inside a deterministic skeleton

basics

~20 s

Use a fixed pipeline whenever the sequence of steps is known before the request arrives. An agent spends model calls deciding what you already know, adds latency variance and non-determinism, and makes evaluation and incident response markedly harder.

solid answer

~50 s

The deciding question is whether control flow depends on data only visible at runtime. If the answer is no — retrieve then answer, classify then route, extract then validate — a composed pipeline wins on every axis that matters in production: one model call instead of several, bounded and predictable latency, deterministic structure you can unit-test, and a failure that points at a specific step. `create_agent` earns its cost when the number and order of tool calls genuinely cannot be predicted: open-ended investigation, tasks where each result changes the next question, or a long tail of requests no fixed graph covers. Even then, bound it — cap the step budget, keep the tool catalogue small and well described, and gate destructive tools behind approval. The common middle ground is a fixed skeleton with one agentic step inside it, so the unpredictable part is isolated and everything around it stays deterministic. In v1 that composes naturally, because the agent is itself a runnable you can embed.

go deeper

for a junior

Be able to say that an agent decides its own steps while a fixed chain runs the steps you wrote, and that the fixed version is cheaper and more predictable.

for a middle

Explain the concrete costs of the loop — extra model calls, a transcript that is re-sent every turn, variable latency — and give an example of a task whose steps are known in advance.

for a senior

Argue the production case: bounded p99, deterministic failure attribution, a blast radius limited to wired tools, and the hybrid pattern of a bounded agentic step inside a deterministic skeleton.

for a principal

Own the decision and its organisational cost. Justify it with trace data on real path distributions, state what agent adoption obliges the team to fund (evals, tracing, budgets, tool governance), and be explicit that improving models shrink the reliability argument but not the latency, cost or blast-radius ones.

## The one question that decides it Does the control flow depend on information that only exists at runtime? That is the whole test. An agent is a mechanism for buying runtime control flow with model calls. If you already know the sequence, you are paying for a decision you do not need. Concretely, these do not need an agent: retrieve-then-answer RAG; classify-then-route; extract-then-validate-then-persist; summarise-each-then-merge. These do: "investigate why this customer's invoice is wrong" where each lookup determines the next; troubleshooting flows over a large tool surface; multi-hop research with no fixed hop count. ## What the agent actually costs **Model calls.** Every turn is a call, and each re-sends the whole accumulated conversation including previous tool outputs. A five-turn agent run is not five times a single call — it is superlinear in tokens because the transcript grows. A fixed pipeline with one call is often an order of magnitude cheaper for the same outcome. **Latency variance.** A pipeline has bounded latency; an agent has a distribution with a long tail. If the surface is user-facing with an SLA, a p99 you cannot state is a product problem, and "we set the step limit high enough" is not an answer. **Non-determinism.** The same input can take different paths on different days, and a model upgrade changes the distribution of paths. That makes regression testing statistical rather than assertive: you need an eval set and a tolerance, not a unit test. **Debuggability.** When a pipeline fails you know which step. When an agent fails you have a transcript to read and a question about why the model chose what it chose. On-call cost is real and rarely priced in. **Blast radius.** An agent chooses which tools to invoke. Any tool that writes, spends or deletes is now reachable by a path an attacker can influence through the prompt. A fixed pipeline reaches only the tools you wired. ## What the agent buys Coverage of the long tail. A fixed graph handles the cases you enumerated; an agent degrades gracefully on the ones you did not. For an internal support assistant fielding thousands of distinct question shapes, enumerating them is not feasible and the agent's flexibility is the product. It also collapses combinatorial routing: ten tools with data-dependent ordering is a graph nobody wants to hand-draw. ## The pattern that usually wins A deterministic skeleton with a bounded agentic core. Validate and classify deterministically. Retrieve deterministically. Then, only for the branch that genuinely needs it, run an agent with a small tool set and a tight step budget. Post-process its output deterministically — validate against a schema, check citations, enforce policy. Because a v1 agent is itself a runnable, embedding it inside a larger composition is straightforward, and you keep the property that most requests never enter the unbounded region at all. ## Levers that shrink the gap If you keep the agent, several LangChain-level controls make it behave more like a pipeline: - **Constrain tool choice.** `bind_tools(tools, tool_choice=...)` can force a specific tool or require that some tool is called, removing the model's freedom on the first hop. - **Shrink the catalogue.** Selection accuracy falls as the tool count rises and descriptions start overlapping. Splitting one twenty-tool agent into narrower agents usually beats prompt-tuning. - **Structured finish.** Use `response_format` so the run ends with a validated object rather than prose, which lets downstream code trust the output. - **Bound the budget.** A step or token budget enforced in middleware, with `recursion_limit` as the runaway guard, keeps worst-case cost stateable. - **Gate the dangerous tools.** Human approval before writes turns an unbounded blast radius into a reviewable one. ## Evaluating the decision honestly Build the fixed version first when you can, then measure what it fails to cover. If the gap is a small enumerable set, extend the pipeline. If the gap is a long tail with no shape, that is the evidence that justifies the agent — and you now have a baseline to compare cost, latency and quality against. Teams that start with an agent because it demos well rarely acquire that baseline and end up unable to say whether the agent is earning anything. ## Organisational angle The choice sets what your team must be good at. Pipelines demand prompt and retrieval quality. Agents demand eval infrastructure, tracing, cost controls, tool-catalogue governance and an on-call story for non-deterministic failures. Choosing agents without funding that second list is the most common way these systems become unowned.

  • How would you justify replacing a working agent with a fixed pipeline to a sceptical team?
    With the trace data. Sample real runs and measure the distribution of tool sequences: if most requests take the same two or three paths, those paths are enumerable and can be wired deterministically, with the agent retained only for the residual tail. Then compare cost per request, p95 latency and eval score between the two on the same inputs. The argument is a measurement, not a preference.
  • What must be funded before an organisation adopts agents broadly?
    An eval set with a quality bar, tracing that captures full transcripts, per-run cost and step budgets, a governed tool catalogue with reviewed descriptions and permissions, and an on-call runbook for non-deterministic failures. Without those, the system cannot be diagnosed or improved, and quality drifts silently whenever a model or a tool description changes.
  • Where does the agent-versus-pipeline line move as models improve?
    It moves in the agent's favour for selection accuracy and multi-step coherence, which shrinks the reliability argument. It does not move for latency variance, cost superlinearity from a growing transcript, or blast radius, since those are properties of the loop rather than the model. So the durable reasons to prefer a fixed pipeline are the operational ones, and design decisions should rest on those rather than on today's tool-choice error rate.

saying these in an interview costs you the question

  • Reaching for an agent because it demos better than a pipeline
  • Ignoring that each turn re-sends the whole growing transcript
  • Quoting an average latency for a run with an unbounded tail
  • Giving one agent twenty tools and tuning the prompt instead of splitting
  • Treating agent quality as unit-testable rather than eval-driven

context