When should an orchestrator drive the ReAct loop instead of letting the model continue it?
answer
- who decides that another pass happens
- guarantee versus strong suggestion
- checkpoint per pass, resume at twelve
- states encode paths you thought of
- outer skeleton, free loop inside
basics
~20 sTake loop control into a state machine when the run must be auditable, resumable, or constrained — regulated or irreversible operations, long-running work that must survive a restart, and tasks whose legal step order is known. Leave continuation to the model when the path genuinely varies per task.
solid answer
~50 sThe question is who decides that another pass happens and what may occur in it. In model-driven control, the agent keeps emitting actions until it answers, and the harness is a thin executor — maximally adaptive, and the right default when the path differs per task. In orchestrator-driven control, a state machine owns the transitions: it decides which tools are permitted in the current state, validates each proposed action before execution, persists a checkpoint per pass so an interrupted run can resume, and can require an approval before a state that does something irreversible. You pay for that in flexibility, since a path nobody anticipated is a path the machine will not allow, and in the engineering cost of maintaining the states. The common resolution is a hybrid: a fixed outer skeleton over the phases you know, with a free ReAct loop inside the one or two steps that genuinely need to improvise.
go deeper
Understand the basic split: either the model keeps choosing actions until it answers, or surrounding code decides what may happen at each step.
Explain what the harness gains by owning transitions — tool permissions per state, validation before execution, a checkpoint per pass — and that it pays for this in flexibility when a task needs an unanticipated path.
Argue the choice from real constraints: irreversibility, resume across restarts, what an incident review must reconstruct. Be able to describe a hybrid with a fixed outer skeleton and a free loop inside the improvising step.
Own the criteria rather than the preference — which failures are unacceptable, what the audit record must support, whether restarts may lose work — and the sequencing argument that states should be derived from observed traces rather than designed ahead of them.
## The axis in question Every agent loop answers one question each pass: *does another pass happen, and what is allowed in it?* Where that decision lives is an architectural choice, and it is independent of whether the reasoning is interleaved. **Model-driven continuation.** The model's output decides. It emits an action and the loop continues; it emits an answer and the loop stops. The harness executes tools and appends observations, and otherwise stays out of the way. **Orchestrator-driven continuation.** A state machine in the harness owns the transitions. Each state defines which tools are legal, what must be true to advance, and what happens on failure. The model is consulted inside a state — it still reasons and proposes actions — but it does not decide the run's shape. ## What orchestrator control actually buys **Enforceable constraints.** "This agent may not write before it has read" or "the refund tool is unreachable outside the verified state" become properties of the machine rather than requests in a prompt. A prompt instruction is a strong suggestion; a state transition is a guarantee. **Durability and resume.** With a checkpoint persisted per pass, a run that dies at pass twelve of a forty-minute job resumes at pass twelve rather than restarting. For long-horizon work this is often the decisive argument, and it is very hard to bolt on afterwards. **Auditability with structure.** A raw transcript is auditable in the sense that you can read it. A state machine gives you something better: which state each action occurred in, why the run advanced, and where it stopped — the difference between a log and a record, which is what regulated review actually wants. **A place to put a gate.** When one step is irreversible, orchestrator control gives you a named transition to interrupt on. Without it, the pause has to be a tool the model chooses to call, which means it is optional exactly when it matters. ## What it costs **Adaptivity.** The states encode the paths you thought of. Real tasks produce paths you did not, and the machine's response is to refuse or to fall out of its states. This is the same tradeoff as any workflow engine: predictability bought with coverage. **Engineering and drift.** States, transitions, guards and their tests are real code that must track a changing task. Teams routinely underestimate this and end up with a machine whose states no longer describe what the agent does. **Wasted capability.** If the model would have found a better route and the machine forbade it, you paid for a capable model and constrained it to a script. Sometimes that is exactly right; it should be a decision, not an accident. ## How to choose Favour **orchestrator control** when: steps are irreversible or externally consequential (money moves, a crew is dispatched, production changes); the run must survive a process restart; a regulator, auditor or incident review will read the record; the legal step order is genuinely known and stable; or several agents and humans share one workflow and need a common notion of where it is. Favour **model-driven control** when: the path varies per task in ways you cannot enumerate; the tools are read-only or cheaply reversible; the work is exploratory — investigation, research, debugging; and you are still learning what the task's real shape is, because premature states freeze in a shape you have not yet understood. **The hybrid is the common landing point** and the answer most senior interviewers are listening for: an outer skeleton over the phases you are confident about — gather, act, verify, report — with an unconstrained ReAct loop inside whichever phase genuinely needs to improvise. You get durability and gating on the boundaries that matter and keep adaptivity where it earns its keep. ## A useful ordering Start model-driven and instrument heavily. Watch real traces and let the recurring shape reveal itself; then harden the parts that have stabilised into states, leaving the volatile parts free. States derived from observed traces describe the task. States designed before any traces exist describe a guess, and you will spend the following months amending them. ## What a principal is expected to own Not a preference for one architecture, but the criteria: which failures are unacceptable, what the record must support after an incident, whether a restart may lose work, and which steps must never proceed unsupervised. Candidates who declare one style universally correct are answering a different, easier question.
- Why not just put the constraint in the prompt instead of building a state machine?Because a prompt instruction is a suggestion the model may fail to follow under pressure, ambiguity or adversarial input, and the failure is silent. A transition guard is enforced outside the model, so violating it is impossible rather than unlikely. Reserve the machinery for constraints whose violation is genuinely unacceptable — for the rest, the prompt is cheaper and more flexible.
- How do you decide which phases deserve to be states in a hybrid design?Take them from observed traces, not from a design session. Phases that appear in nearly every successful run in the same order have stabilised and are safe to harden. Phases whose shape differs run to run should stay inside a free loop. Adding a state that reality then contradicts costs more than leaving it unconstrained.
- What breaks first when a team over-constrains an agent into states?Coverage. Tasks that need a route nobody enumerated either fail or get forced down a wrong path, and the fix arrives as an ever-growing set of special-case states. Watch for a rising rate of runs that terminate inside a state with no legal transition — that is the machine telling you its model of the task is too narrow.
saying these in an interview costs you the question
- Declares one control style universally correct
- Thinks a prompt instruction enforces a constraint
- Designs the state machine before observing any traces
- Ignores resume and durability as a deciding factor
- Assumes constraining the loop costs nothing in coverage