When should an agent call tools from sandbox code instead of one call per turn?
answer
- Two paths, one tool implementation
- Cost scales with calls, or with turns
- Loops belong in code, judgment in turns
- Small n does not repay a sandbox
- A gated action must stay gateable
basics
~20 sWhen the work is a wide, mechanical fan-out with known control flow. Writing a loop that invokes the tool inside a sandbox costs one model turn instead of hundreds, while conversational calls remain right for short, exploratory work where each result changes the next decision.
solid answer
~50 sThere are two invocation paths in 2026 practice. The conversational one — a tool-call block, a result turn, another model inference — costs a full round trip per call, and its cost scales with the number of calls. The second path is programmatic: the model writes code in a sandbox, and tools are invoked from that code, so loops, conditionals and filtering run without a model turn each. Checking stock for 200 SKUs is a `for` loop and one turn, not 200 turns. **Reach for it when the control flow is expressible in advance and the fan-out is wide**; stay conversational when each result genuinely changes the next step, when there are only a handful of calls, or when every call needs human oversight. The cost is real: you now operate a sandbox, model-written code fails in ways a schema-validated call cannot, and a step executed inside code is a step no reviewer saw.
code
python · 9 linesdef restock_report(skus, store, check_stock, threshold=5):
"""Runs inside the sandbox: one model turn, len(skus) tool calls."""
low = []
for sku in skus:
units = check_stock(store=store, sku=sku)["units"]
if units < threshold:
low.append((sku, threshold - units))
low.sort(key=lambda row: row[1], reverse=True)
return low[:10]go deeper
Know that there are two ways an agent reaches a tool: one call per model turn, or code in a sandbox that calls the tool many times within a single turn.
Explain why the conversational path costs a model inference per call and how a loop in code collapses a wide fan-out into one turn, keeping intermediate results out of the context window entirely.
Argue the boundary with examples: exploratory work stays conversational, mechanical fan-out goes to code, and gated destructive actions stay where the harness can approve them one at a time. Name the new failure modes you inherit.
Own the routing policy and its costs — sandbox isolation and egress limits, which tools are programmatically callable at all, how audit parity is preserved, and when a repeatedly generated script should simply become a deterministic workflow you version.
## Two paths to the same tool The conversational path is the one everyone learns first: the model emits a tool-call block, your runtime executes it, you append a result turn, and the model runs again. Its defining property is that **every call costs a model inference**, in latency, in tokens, and in money. Ten calls, ten inferences. The programmatic path inverts this. The model writes code — typically Python — which runs in a sandbox, and the tools are callable from inside that code. Providers expose this by marking a tool as invocable by the code-execution environment rather than only by the conversation. One turn produces a script; the script may call a tool once or five hundred times; the model sees only what the script returns. Both paths reach the same tool implementation. What differs is who drives the loop: the model, one decision at a time, or code the model wrote once. ## Where programmatic invocation wins **Wide fan-out.** A restaurant chain checking stock for 200 SKUs across stores is the canonical shape. Conversationally that is hundreds of round trips even with parallel calls batching some of them. As a loop it is one turn, and the wall-clock difference is the difference between minutes and seconds. **Known control flow.** If you can describe the procedure — for each SKU, call the tool, keep the ones under the reorder threshold, sort by shortfall — then the model does not need to be in the loop to execute it. It needs to be in the loop to *write* it. **Data reduction before the answer.** Intermediate results never enter the conversation. Only what the script returns does. This is usually a bigger win than the turn count, because raw tool output is the fastest-growing part of an agent's context. **Composition.** Joining two tools' outputs, retrying with backoff, or applying arithmetic across results is natural in code and awkward as a sequence of turns where the model must hold intermediate values in its head. ## Where conversational calls remain right **Genuine exploration.** When the next action depends on what the last result said — the shape of most debugging, research and support work — the model has to see each result. Writing a script that guesses the branch is worse than taking the turns. **Small n.** Three calls do not justify a sandbox. The fixed cost of generating code, starting an environment and interpreting a traceback exceeds what you save. **Supervised or destructive actions.** If a human approves each side effect, a loop that fires fifty of them inside a sandbox has bypassed the control you built. Actions with real consequences belong where the harness can gate them individually. **Auditability requirements.** In a regulated flow, "the agent decided to do X, here is the call" is a record. "The agent ran a script that did 40 things" is a weaker one unless you instrument the sandbox to emit an equivalent trail. ## What the second path costs you This is not a free upgrade, and a principal-level answer says so. - **You are now running untrusted code.** Sandbox isolation, network egress limits, filesystem scoping, CPU and memory caps, and time limits all become your problem. Model-written code plus a tool with credentials is exactly the shape prompt-injection attacks aim at. - **New failure modes.** A schema-validated tool call fails in bounded ways. A script fails with a traceback, an infinite loop, a swallowed exception, or a subtly wrong filter that silently returns the wrong half of the data. - **Debuggability shifts.** The trace stops being a readable sequence of calls and becomes a program plus its output. You need logging inside the sandbox to get back what the conversational path gave you for free. - **Model variance.** Code generation quality varies more across models and prompts than structured argument filling does. A weaker model that calls tools competently may write brittle scripts. ## How mature systems actually decide They run both, and route. Short interactive work stays conversational. Batch and analytical work goes to code. A frequent hybrid is to let the model act conversationally while it explores, then switch to a script once the procedure is clear — the model has effectively written a workflow that it now executes. And the honest end of the spectrum: when the procedure is stable enough to write as code every time, the strongest option is often to stop generating it per run and ship it as a deterministic workflow with the model reduced to the parts that need judgment. ## Reflecting the state of practice As of mid-2026 the code path has moved from experiment to a standard tool in the agent-engineering kit, with reported context reductions in the high nineties of a percent on fan-out-heavy tasks. It is not a replacement for the conversational protocol; it is the second gear, and knowing which gear a workload is in is the judgment the question is testing.
- What does moving invocation into a sandbox do to your security posture?It widens the blast radius considerably. You are executing model-written code that can reach tools holding credentials, so untrusted content anywhere in the context becomes a path to arbitrary tool use. The mitigations are structural: strict sandbox isolation, network egress allowlists, scoped filesystem access, resource and time caps, and keeping destructive or irreversible tools off the programmatically callable list entirely so they stay behind an approval gate.
- How does the debugging story differ between the two paths?Conversationally, the trace is a readable sequence of calls and results, and a wrong step is visible as a bad call. In code mode the trace is a script plus whatever it chose to return, so a wrong filter or a swallowed exception can hide a hundred calls behind a plausible summary. You get parity back only by instrumenting the sandbox to log each tool invocation and its arguments as spans, which is work the conversational path gave you free.
- If the model keeps writing the same script every run, what does that tell you?That the procedure has stopped being a judgment call and become a workflow. Regenerating it each run pays for nondeterminism you no longer want and adds a failure mode — a slightly different script — for no benefit. Promote it to deterministic code you own and version, and keep the model for the parts that genuinely vary: interpreting the request, handling exceptions the code cannot classify, and writing the final answer.
saying these in an interview costs you the question
- Treating code mode as a strict upgrade over tool calls
- Ignoring that model-written code is untrusted code
- Using a sandbox for three calls
- Letting gated destructive actions run inside a loop
- Assuming the trace stays as auditable as before