How do you decide which tasks get a fixed workflow, plan-and-execute, or ReAct?
answer
- how much of the sequence is knowable upfront
- the no-agent baseline is the honest comparison
- irreversible actions argue for structure
- escalate on measured failure, not fashion
- promote converged loops back down to plans
basics
~20 sMatch the strategy to how much of the step sequence you can know in advance. Fully known steps get hard-coded code, a knowable shape with variable content gets an upfront plan, and genuinely open-ended work gets a ReAct loop.
solid answer
~50 sAsk one question per task class: *before the first tool runs, how much of the sequence do I already know?* If you know all of it, do not build an agent. A shipment-status feature that always does lookup, format, reply is three lines of code with at most one model call for wording — turning it into an agent buys nondeterminism, cost variance and a worse audit trail in exchange for nothing. If you know the shape but not the content — the same seven reconciliation steps with different identifiers each month — plan upfront and execute; you get determinism, parallelism and a plan you can review before side effects fire. If you genuinely cannot enumerate the steps, because each depends on what the last returned, use ReAct and accept the token and latency bill. Decide this per task class, not once for the product, and start with the cheapest strategy that passes your evals — escalate only on measured failure.
go deeper
Know that not every LLM feature is an agent: when the steps are always the same, ordinary code with one model call is the right build, and it is cheaper and more predictable.
Be able to sort a task by how much of its sequence is knowable upfront, and explain what agents cost you — variable latency, variable spend, harder debugging — that a workflow does not.
Bring the measurements: baseline scores for the non-agent version, plan completion rate, trace convergence, and reversibility of the actions involved. Argue the choice from those, not from architecture preference.
Own the policy — written escalation criteria per task class, an enforced no-agent baseline, and the discipline to demote converged loops back to plans or code as evidence arrives. Be candid that the boundaries are contested and defend your rule, not a consensus.
## The question that orders everything Before the first tool call, how much of the step sequence do you already know? That single question sorts nearly every task into one of three buckets, and it is the framing a principal is expected to bring — because the failure mode at this level is not choosing the wrong loop, it is choosing an agent at all when a function would do. ## Bucket one: you know all the steps A shipment-status lookup does the same three things every time: resolve the tracking ID, fetch the status, render a reply. There is no decision for a model to make about sequence. Implement it as code. If you want natural-sounding phrasing, make one model call inside the last step — that is an LLM feature, not an agent. What you avoid is substantial: a fixed workflow has a predictable cost per request, a predictable latency, a stack trace when it fails, and no possibility of the system deciding to call a tool you did not intend. Every one of those properties is expensive to recover once a model owns the control flow. The honest baseline for any agent proposal is "what does the non-agent version cost and score?", and teams that skip it routinely ship agents that are slower, dearer and less reliable than the `if` statement they replaced. ## Bucket two: you know the shape, not the content A monthly reconciliation across four systems is the same seven steps with different identifiers; a customer onboarding is the same checks in the same order with different documents. Here the sequence is derivable upfront but the arguments are not, so a planning call that emits a step list with substituted values gives you determinism, concurrency across independent steps, and — often the deciding factor — a plan that a human or a policy check can inspect before anything with side effects executes. In regulated domains that inspectable artifact is worth more than the token saving. ## Bucket three: you cannot enumerate the steps Incident diagnosis, exploratory research, open-ended code changes. What you query second is determined by what the first result said. Any upfront plan here is a guess in a structured costume, and its determinism is worse than useless because it is deterministic about the wrong steps. Run ReAct, budget for the round trips, and invest in observability instead of structure. ## Second-order considerations that move the line **Reversibility.** The more irreversible the actions — money moves, customer emails, production writes — the more you want a reviewable plan or plain code, and the less you want a model choosing freely at each step. Adaptivity is cheap to grant on read-only work and expensive on write paths. **Variance tolerance.** Agents produce a distribution of behaviours, not one behaviour. If your product promises a consistent experience or your contract promises a response time, that distribution is a liability, and a strategy that narrows it earns its constraints. **Operational surface.** Every strategy you run is a distinct failure mode, a distinct eval harness and a distinct on-call story. A portfolio with all three is legitimate for a large product and pure overhead for a small one. Consolidating on two is usually the right organizational answer even when three would be marginally better per task. **Migration in both directions.** The interesting move is not just workflow-to-agent. When traces show a ReAct loop converging on the same path for 95% of inputs, that path has become knowable and should be promoted to a plan or to code, with the loop kept only as the fallback for the tail. Teams reliably do the first migration and almost never do the second, and that is where a lot of avoidable cost lives. ## How to actually run the decision Start at the cheapest strategy that could work and escalate only on measured failure against a real eval set. Escalation criteria should be written down before you build: "if the fixed workflow cannot handle more than 80% of inputs without a human, we plan; if generated plans complete without revision less than 70% of the time, we loop." Without those numbers the decision defaults to whatever is fashionable, which in 2026 means everything becomes an agent. And be honest that this is contested ground. There is no settled industry answer on where the boundaries sit; what is defensible is having a stated rule, a baseline you measured against, and a willingness to move a task back down the ladder when the evidence says its uncertainty was smaller than you assumed.
- How do you tell whether a task genuinely needs an agent?Try to write its step list in advance. If you can cover the large majority of real inputs with an enumerable sequence, it is a workflow with an LLM inside a step, not an agent. Build that version first and measure it — it becomes the baseline any agent proposal has to beat on quality, cost and latency, and it frequently wins.
- What would make you move a task from ReAct back to a fixed workflow?Traces converging: if the loop takes essentially the same path for the overwhelming majority of inputs, that path is knowable and should be promoted to code or a plan, with the loop retained only for the tail. This migration is rarely done and is where a lot of recoverable cost and latency sits.
- How does action reversibility change the choice?It shifts you toward structure. Read-only exploration is a cheap place to grant a model free choice at each step; write paths that move money, contact customers or change production are not. There the reviewable upfront plan — or plain code — is worth its rigidity, because the cost of a wrong step is not tokens but a real-world effect you cannot take back.
- What is the cost of running all three strategies across one product?Three failure modes, three eval harnesses, three sets of runbooks and three things a new engineer must learn before touching anything. That portfolio is defensible at scale where task classes genuinely differ, and pure overhead on a small product. Consolidating on two is often the right organizational call even when three would score marginally better per task.
saying these in an interview costs you the question
- Defaults every task to an agent because agents are the current architecture
- Never builds or measures the non-agent baseline before proposing a loop
- Assumes a fixed workflow cannot contain an LLM call inside a step
- Ignores that nondeterminism carries operational, contractual and audit cost
- Only ever migrates tasks up the ladder, never back down when traces converge