When does fine-tuning on ReAct traces beat few-shot prompting the loop?
answer
- scaffolding in the prompt or in the weights
- narrow, high-volume, stable — all three
- what happens when a tool is added
- filtered trajectories, not raw traces
- hybrid: format in weights, policy in prompt
basics
~20 sWhen the workflow is narrow, high-volume and stable, exemplars have grown long without fixing format or tool-selection errors, and thousands of verified trajectories exist. Fine-tuning moves the scaffolding into the weights, buying a shorter prompt at the price of a retrain whenever the tools change.
solid answer
~50 sThe question is where the loop's scaffolding lives: in the prompt, re-sent on every step, or in the weights. Prompting keeps everything editable — change policy by editing text, ship in minutes, add a tool without retraining. Fine-tuning on verified trajectories pays when a narrow, high-volume workflow has already exhausted prompt work: exemplars are long, format and tool-selection errors persist, and prefix tokens are a real share of cost and latency at your volume. It also lets a smaller model run a loop a larger one was needed for. The costs are structural, not just upfront: the tool catalogue and action grammar are baked in, so adding or renaming a tool means a new dataset and a new run, and you lose the ability to change behaviour by editing a sentence. This is genuinely contested ground in 2026 — a strong instruction-tuned model with good scaffolding often matches a fine-tune at far lower operational cost, and each model generation erodes the margin further.
go deeper
Know the basic distinction: the loop's format and examples can either sit in the prompt on every call, or be trained into the model so it produces them without being shown.
Explain what fine-tuning buys — shorter prompts, more reliable format — and what it costs, chiefly that the tool set and grammar are baked in and change now means retraining.
Show the decision process: establish a measured prompted baseline, exhaust cheap prompt work, quantify what the fine-tune must buy, and insist on outcome-filtered trajectories rather than raw traces.
Own the strategic call. Weigh a training pipeline coupled to your roadmap against prompt agility, argue the hybrid split of format-in-weights and policy-in-prompt, and state plainly that the margin narrows with each model generation.
## The real question A ReAct agent needs the model to do three things: emit a well-formed step, pick the right tool, and know the domain's conventions. Prompting supplies all three as text that rides along on every single call. Fine-tuning moves some of them into the weights. The design question is not "is fine-tuning good" — it is **which of those three things should live where**, and what you give up by moving them. ## A concrete frame Take a hotel booking-modification workflow: change dates, change room type, apply a rate rule, refund or charge the difference, notify the guest. Six tools, a fixed policy, hundreds of thousands of runs a month, and a task shape that barely changes quarter to quarter. Prompted, each run carries a system block, six exemplar trajectories and the format description on every step of a five-to-eight step loop. Fine-tuned on validated trajectories from that same workflow, the exemplars disappear, the prompt shrinks to the tools and the case, and the model emits the step format natively because that is what it was trained to produce. That is the shape where fine-tuning pays. Notice what makes it work: **narrow, high-volume, stable**. ## The signals that favour fine-tuning - **Prompt work has plateaued.** You have iterated on exemplars and format description, and a stubborn residue of malformed steps or wrong-tool choices remains. Weights can encode what prose could not. - **Volume makes prefix tokens material.** At high request rates, exemplar tokens multiplied by steps per run become a genuine line item, and shorter prompts mean lower latency per step. (Prompt caching blunts this considerably, which is why the argument is weaker than it was.) - **You want to run a smaller model.** Trajectory fine-tuning is often used to distil a capable model's successful runs into a cheaper one that could not follow the loop from a prompt alone. - **You have enough verified data.** Typically thousands of trajectories, filtered by outcome — runs that actually succeeded, checked by a verifier rather than by vibes. Fine-tuning on unfiltered traces teaches the model your existing failures. - **The task is stable.** If the toolset and policy have not moved in six months, baking them in is a bet you can win. ## The costs, which are ongoing **Change becomes a retrain.** This is the big one. With a prompt, adding a seventh tool or reversing a refund policy is an edit. With a fine-tune, the model has learned a catalogue and a grammar; new tools are out of distribution, and the model may keep reaching for the shapes it was trained on. You now have a training pipeline coupled to your product roadmap. **You lose fast rollback.** A prompt regression is reverted in minutes. A weights regression means restoring a prior checkpoint and re-validating, which is a different class of operation. **Data curation is the real work.** Trace hygiene — deduplication, removing runs that succeeded by accident, stripping anything sensitive, keeping the distribution representative — is where most of the effort goes, and it is recurring, because the distribution shifts. **Brittleness off-distribution.** A trajectory-tuned model is excellent inside the shape it saw and can be worse than the prompted baseline outside it. If your "narrow" workflow turns out to have a long tail, you inherit that. **Evaluation obligations.** You now need an eval suite good enough to compare two models, not just two prompts, and to gate every retrain. ## The hybrid, which is usually the right answer Split the three things. Fine-tune the parts that are stable and mechanical — the step format, the alternation discipline, the domain's register. Keep the parts that change in the prompt — the tool catalogue, the policy, the escalation rules. You get shorter prompts and reliable format compliance while retaining the ability to change behaviour by editing text. It also softens the tool-change problem: the model has learned *how to emit an action*, not *which actions exist*. ## Be honest that this is contested There is no settled consensus in 2026, and a candidate who states one confidently is signalling less than one who characterises the disagreement. The reasons it stays open: - Each model generation follows scaffolded prompts better than the last, so the gap a fine-tune closes keeps shrinking, and yesterday's clear win becomes today's marginal one. - Prompt caching removed much of the token-cost argument. - Meanwhile, structured tool-calling interfaces removed much of the *format* argument, since format compliance now comes from the API rather than from your exemplars. What is left for fine-tuning is mostly judgement — which tool, when to stop, how to handle this domain's edge cases — which is the harder thing to teach and the more valuable thing to own. ## How to decide The defensible process: establish a prompted baseline with proper measurement, exhaust cheap prompt work, quantify what a fine-tune would need to buy in accuracy, latency or unit cost, and estimate the ongoing cost of the retrain pipeline against your expected rate of tool and policy change. If the workflow will be re-scoped next quarter, do not bake it into weights. If it has been stable for a year and runs at volume, it is a reasonable bet — and the hybrid is usually a better one.
- Where do the training trajectories usually come from?Mostly from a stronger model's runs on the same workflow, filtered by outcome — kept only when a verifier confirmed the run actually succeeded — plus curated human-corrected traces. The filtering is the point: unfiltered production traces teach the model your existing failure modes, including confident wrong answers. Expect the curation, deduplication and sensitive-data stripping to be the bulk of the effort.
- What is the hybrid you would usually reach for instead?Fine-tune the stable, mechanical parts — step format, alternation discipline, domain register — and keep the volatile parts in the prompt: tool catalogue, policy, escalation rules. You get short prompts and reliable formatting without coupling your product roadmap to a training pipeline, because the model learned how to emit an action rather than which actions exist.
- How has the rise of structured tool calling changed this decision?It removed much of the format argument. When the API enforces the call shape, exemplars teaching syntax are largely redundant, so the remaining case for fine-tuning is judgement — which tool suits which subgoal, when to stop, how this domain's edge cases resolve. That is harder to teach and slower to obsolete, but it also means the easy win from fine-tuning the format is gone.
- What would make you refuse to fine-tune a workflow that otherwise fits?An unstable toolset. If tools are added or renamed monthly, or policy is still being negotiated, the retrain cadence will dominate the benefit and the model will keep drifting toward grammar it learned rather than what now exists. Volume and plateaued prompt work are not enough on their own — stability is the third condition, and it is the one teams skip.
saying these in an interview costs you the question
- Treats fine-tuning as a general fix for weak prompting
- Ignores that adding a tool now requires a retrain
- Trains on raw production traces without outcome filtering
- Assumes fine-tuning removes the need for an eval suite
- Claims the prompted-versus-fine-tuned question is settled