skip to content

Explain internal vs user-controlled tool execution and the returnDirect option. When would you disable internal execution?

level: seniorimportance: should knowfreq 30%

answer

  1. internal = auto loop (default)
  2. internalToolExecutionEnabled(false) = you run it
  3. human-in-the-loop / different process / audit
  4. returnDirect=true -> result straight to caller, skip model
  5. ToolCallingManager.executeToolCalls

basics

~20 s

By default Spring runs tools internally: it executes the requested tool and re-calls the model automatically until a final answer. You can disable this (internalToolExecutionEnabled=false) to get the raw tool-call request and run it yourself. returnDirect=true returns the tool's result straight to the caller instead of sending it back to the model.

solid answer

~50 s

Spring AI has two tool-execution modes. In the default internal mode, ToolCallingManager runs the requested tool(s) and loops back to the model automatically, so .content() yields final text. In user-controlled mode you set ToolCallingChatOptions.internalToolExecutionEnabled(false): the ChatResponse comes back containing the tool-call request(s) without executing them, and you decide whether/how to run them — useful for human-in-the-loop approval, executing tools in another process/service, custom auditing, or when tool execution must be async or transactional under your control. You then run the tools (optionally via ToolCallingManager.executeToolCalls) and re-submit. Separately, @Tool(returnDirect=true) short-circuits the loop: the tool's result is returned directly to the caller and NOT sent back to the model for further reasoning — good when the tool output IS the answer (e.g. a booking confirmation) and you want to save a round-trip and avoid the model paraphrasing it.

code

java · 16 lines
java
// returnDirect: the booking confirmation IS the answer
class BookingTools {
    @Tool(description = "Book a table and return the confirmation", returnDirect = true)
    Confirmation book(String restaurant, String time) {
        return bookingService.book(restaurant, time); // returned as-is to caller
    }
}

// User-controlled execution for human approval
ChatOptions opts = ToolCallingChatOptions.builder()
        .toolCallbacks(ToolCallbacks.from(new BookingTools()))
        .internalToolExecutionEnabled(false)
        .build();

ChatResponse resp = chatModel.call(new Prompt("Book Nobu at 8pm", opts));
// resp now holds the tool-call request(s) -> show to a human, then execute + resubmit

go deeper

for a junior

Know the default is automatic execution.

for a middle

Can describe returnDirect and that internal execution can be turned off.

for a senior

Chooses user-controlled execution for approval/audit/distributed cases and manages the resubmit loop correctly.

for a principal

Designs the execution model around safety (approval gates for destructive tools), latency/cost (returnDirect), and where tool side effects must run.

## Two execution modes ### Internal (default) With default options, when the model requests a tool, Spring's `ToolCallingManager` immediately executes it and re-invokes the model with the result, repeating until the model stops requesting tools. From the caller's view it's one `.call().content()` that returns final text. This is what you want most of the time. ### User-controlled Set on the options: ```java ChatOptions opts = ToolCallingChatOptions.builder() .toolCallbacks(myCallbacks) .internalToolExecutionEnabled(false) .build(); ChatResponse resp = chatModel.call(new Prompt("...", opts)); ``` Now Spring does **not** execute the tools. The `ChatResponse` contains the model's **tool-call requests** (name + arguments). You inspect them and decide what to do: - run them yourself, then build the tool-response messages and call the model again; - or delegate to `ToolCallingManager.executeToolCalls(prompt, response)` which runs them and returns updated conversation state you resubmit. **When to disable internal execution:** - **Human-in-the-loop / approval** — a person must approve a destructive action (refund, delete) before it runs. - **Distributed execution** — the tool must run in a different service, process, or security context. - **Custom orchestration** — you need explicit transaction boundaries, retries, rate-limiting, or auditing around each tool call. - **Async / deferred** — the action can't complete synchronously within the chat round-trip. ## returnDirect `@Tool(returnDirect = true)` (or `returnDirect(true)` on a programmatic callback) changes what happens **after** a tool runs: instead of feeding the result back to the model to continue reasoning, Spring returns the tool's result **directly to the caller** and ends the loop. Use it when: - The tool output *is* the final answer (e.g. a generated report, a booking id, structured data the client will render) and you don't want the model to paraphrase or hallucinate around it. - You want to **save a model round-trip** (cost/latency) since re-calling the model is unnecessary. - You need the **exact, unaltered** tool output preserved. Contrast: with the default `returnDirect=false`, the result re-enters the conversation and the model produces the user-facing text. Note: if the model requests several tools in one turn and any has `returnDirect=true`, the loop terminates and those direct results are returned — design accordingly. ## How they combine - Internal + `returnDirect=false` (default): fully automatic, model writes the final answer. - Internal + `returnDirect=true`: Spring runs the tool automatically but returns its raw result, skipping the final model call. - User-controlled: you run tools yourself regardless of `returnDirect`. ## Gotchas - Disabling internal execution means **you** own the loop — forgetting to resubmit leaves the conversation stuck at the tool-request stage. - `returnDirect=true` bypasses the model's ability to format/summarize; the caller must handle raw tool output. - With user-controlled mode you must faithfully attach tool results to the right tool-call IDs, or the follow-up model call will be inconsistent. - These are per-options/per-tool settings, not global model behavior — set them where they apply.

  • What is the practical effect of returnDirect=true on latency and cost?
    It skips the extra model round-trip that would normally format the tool result, cutting one LLM call (lower latency and token cost). The trade-off is the caller receives the raw tool output rather than a model-written natural-language answer.
  • After disabling internal execution, how do you complete the interaction?
    You read the tool-call requests from the ChatResponse, execute the tools yourself (or via ToolCallingManager.executeToolCalls), attach each result to its tool-call id as a tool-response message, and re-submit the prompt so the model can continue to a final answer.

saying these in an interview costs you the question

  • Believing tools are always executed by Spring automatically with no way to intercept
  • Thinking returnDirect just changes the return type rather than short-circuiting the model loop
  • Assuming user-controlled mode still auto-resubmits results

context