skip to content

Tool Use and Function Calling

How a model reaches outside itself: you describe tools as JSON schemas, the model emits a call, your code runs it, and the result goes back into the conversation. Interviewers focus here because the schema, the error path, and the permissions you grant a tool decide whether an agent is useful or dangerous.

part ofAI agentsoverview, primer and where to startread it →
on this pageshow

questions

20

In LLM function calling, what does the model emit and what turn must you append?

level: juniorimportance: must knowfreq 82%

answer

  1. The model proposes, your code runs it
  2. Two turns, not one
  3. A call id joins request to answer
  4. Echo the assistant turn verbatim
  5. Loop until a turn has no call

basics

~20 s

Instead of prose the model returns a structured tool-call block naming the tool, its arguments and a call id. Your code executes the tool, appends a tool-result turn carrying that same id, and calls the model again with the extended history.

solid answer

~50 s

Giving a model tools does not give it network access — it extends what the model may emit. Alongside text it can produce a **tool-call block** (`tool_use` / `tool_calls`, depending on the provider) holding a tool name, an arguments object, and a provider-generated call id. Nothing has run yet; your program is the runtime. You execute the tool, then append two things to the conversation: the assistant turn exactly as returned, and a following turn containing a **tool-result block that echoes the call id**. Then you call the model again with the whole history. It reads the result as ordinary context and either answers or requests more tools, and you loop until a turn arrives with no tool call. The model is stateless across these calls — the growing message list is the entire state.

code

json · 11 lines
json
[
  { "role": "user", "content": "How many umbrellas are left in the Brighton store?" },
  { "role": "assistant", "content": [
      { "type": "tool_use", "id": "call_a1", "name": "check_stock",
        "input": { "store": "brighton", "sku": "UMB-01" } }
  ] },
  { "role": "user", "content": [
      { "type": "tool_result", "tool_use_id": "call_a1", "content": "{\"units\": 14}" }
  ] },
  { "role": "assistant", "content": "Brighton has 14 umbrellas in stock." }
]

go deeper

for a junior

Be able to narrate the round trip in order: model emits a tool-call block, your code runs the tool, you append a result turn with the same id, you call the model again. Say plainly that the model never executes anything itself.

for a middle

Explain why the assistant turn is replayed verbatim and what the call id joins together, and describe the loop's exit condition. Expect a follow-up on where argument validation belongs in the flow.

for a senior

Show that you treat the message array as the agent's whole state: tool results dominate context growth, order is memory, and refusing a call still requires a result turn. Talk about caps, validation and audit at the execution boundary.

for a principal

Own the boundary as a policy surface. Decide what the harness enforces between proposal and execution — authorization, approval gates, spend caps, redaction — and what history rewriting is permitted, since every edit changes what the model believes happened.

## What a tool call really is When you attach tools to a model request, you are not opening a network path for the model. You are widening its output vocabulary. Alongside ordinary text tokens, the model may now emit a **structured tool-call block**. That block carries three things: the tool's name, an arguments object shaped by the tool's declared schema, and a call id the provider generates for this specific invocation. At the moment the response comes back, nothing has executed. The model has written a request, and your program is the runtime that decides whether and how to honour it. This is the single most useful thing to internalize about function calling: the model proposes, your code disposes. Every guardrail you want — argument validation, permission checks, rate limits, approval prompts — lives in your code between receiving the block and running the tool. ## The round trip, step by step 1. **Request.** You send the conversation history plus the tool definitions. 2. **Assistant turn.** The model replies. The turn may contain text, one or more tool-call blocks, or both. A stop reason (commonly `tool_use` or `tool_calls`) signals that the model is waiting on you rather than finished. 3. **Execute.** Your code parses the arguments, validates them, and runs the real function, HTTP request or query. 4. **Append two turns.** First the assistant turn **exactly as the model returned it**, then a following turn containing a tool-result block whose id field (`tool_use_id` / `tool_call_id`) repeats the id from the call. 5. **Call again.** Send the extended history. The model now sees its own request and the answer, and either produces a final text answer or requests more tools. You loop steps 2–5 until an assistant turn arrives with no tool call in it. That terminal turn is the answer to give the user. ## Why the id exists The id is the join key between a request and its answer. With a single call it looks like ceremony, but the moment a turn contains several calls — or you execute them concurrently and they finish out of order — the id is the only thing that tells the model which answer belongs to which request. Providers reject a result whose id does not match an outstanding call, which is a feature: it catches pairing bugs at the API boundary rather than as quietly wrong reasoning three turns later. ## Why the assistant turn goes back verbatim A common early mistake is to drop the assistant tool-call turn and send only the result, or to paraphrase the call as prose ("I looked up the stock and found 14"). Both break the contract. The model's next reasoning step is conditioned on seeing its own request next to the answer; without the request, the result is a floating fact with no provenance, and providers generally reject a result block with no preceding call to match. Keep the raw block and replay it. ## Conversation state across turns The model holds no state between HTTP calls. Everything it knows on turn seven is in the message array you send on turn seven. That has three practical consequences. Tool results accumulate in the window and are the fastest-growing part of an agent's context. The order of turns is the agent's memory, so appending out of order corrupts its view of what happened. And any turn you edit or drop — for compaction, redaction or cost — changes what the model believes, which is a design decision, not a housekeeping detail. ## Multiple calls in one turn A turn may contain more than one tool-call block when the requests are independent. The contract is that **all** of them must be answered before the model runs again: you send one following turn holding one result block per call, each tagged with its own id. ## What usually goes wrong - Sending the result without echoing the assistant turn — the provider errors, or the model loses the thread. - Answering only some of the calls in a multi-call turn. - Treating raw arguments as trusted. They are model-generated text; validate types, ranges and permissions before executing anything with side effects. - Forgetting that the model chose to call the tool — if it calls the wrong one, the fix is in the tool's description or your loop, not in the transport. - Assuming a tool call means the tool ran. Until your code runs it, the call is only a proposal, and returning a result for a tool you deliberately refused to run is a legitimate, and often correct, move.

  • If your code refuses to run a tool the model asked for, what do you send back?
    You still send a result turn for that call id — the contract requires every outstanding call to be answered before the model runs again. Put a short factual explanation in the result content, such as that the operation needs approval or the arguments failed validation. The model reads it as context and can re-plan, ask the user, or choose another tool, which is far better behaviour than a dropped call or an aborted loop.
  • What actually ends the tool-calling loop?
    An assistant turn that contains no tool-call block — the stop reason flips from a tool-use signal to an ordinary end-of-turn, and the text in that turn is the final answer. Because a model can keep requesting tools indefinitely, production loops also carry hard caps on iterations, wall-clock time and spend, and treat hitting a cap as an error path with a partial answer rather than as normal completion.
  • Where do you validate the model's arguments in this flow?
    Between receiving the tool-call block and executing anything. The arguments are model-generated text that merely claims to match your schema, so you parse and validate types, ranges, enums and identifiers, then apply authorization for the current user before any side effect. Failures come back as a normal result turn describing the problem, which lets the model correct itself instead of your loop crashing.

saying these in an interview costs you the question

  • Thinking the model executes the tool itself
  • Sending only the result without the assistant call turn
  • Believing the model remembers past turns without the history
  • Treating model-supplied arguments as validated input
  • Answering just one call in a multi-call turn

context

open as a page

In a function-calling tool definition, which parts does the model actually see?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The model sees only three things: the tool's name, its natural-language description, and the JSON Schema describing its parameters. Implementation code, docstrings and internal comments never reach it, so everything it needs must live in those three fields.

open as a page

In LLM tool calling, how are parallel tool results returned and paired to their calls?

level: middleimportance: must knowfreq 66%

basics

~20 s

One assistant turn can carry several independent tool-call blocks. All of them are answered in a single following turn holding one result block per call, each matched by the call id it echoes — never by array position or completion order.

open as a page

When a tool call fails, why return an is_error tool result rather than letting the exception propagate?

level: middleimportance: must knowfreq 68%

basics

~20 s

An exception ends the agent's turn; a tool result flagged as an error keeps the loop alive and hands the model a fact it can act on. The model can then fix an argument, pick another tool, or report honestly that the step failed.

open as a page

Why does an agent's tool-selection accuracy fall as its catalog grows from 20 to 400 tools?

level: middleimportance: must knowfreq 72%

basics

~20 s

Every tool definition is loaded into the prompt at once, so hundreds of them cost tens of thousands of tokens and force the model to discriminate between many near-identical options in a single pass. Crowding plus semantic overlap pushes selection accuracy down.

open as a page

Why is a tool's description field a prompt rather than documentation?

level: middleimportance: must knowfreq 66%

basics

~20 s

The description is the model's only instruction about a capability, so it must be written to drive a decision: what the tool does, what must be true before calling it, where argument values come from, and which nearby requests it must not be used for.

open as a page

An incident agent's log-search tool returns 40 MB of JSON. Where do you reduce it?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Reduce it before it becomes context. Aggregate inside the tool at write time, or let the agent run the query as code in a sandbox and return only the derived findings. Trimming history afterwards is damage control, not the fix.

open as a page

How should an agent serialize a tool's return value before feeding it back to the model?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Serialize it into a short, self-describing string the model can read on its own: name the fields, keep units and status explicit, include the identifiers the next call will need, and leave out everything the model cannot act on.

open as a page

What do the tool_choice modes auto, required and none control in LLM tool calling?

level: middleimportance: should knowfreq 54%

basics

~20 s

They constrain what the model may emit on that one request: auto lets it choose between text and a tool call, required (called any by some providers) forces some tool call, a named tool forces that specific one, and none forbids tool calls entirely.

open as a page

Why must an agent treat a fetched web page returned by a tool as data rather than instructions?

level: middleimportance: should knowfreq 45%

basics

~20 s

Because nobody you trust wrote it. A tool result is content authored by an external system, and a model reading it has no intrinsic way to tell narration from a command, so any imperative text inside it must carry no authority in your agent.

open as a page

How does progressive disclosure shrink the context cost of a 400-tool agent catalog?

level: middleimportance: should knowfreq 46%

basics

~20 s

Load one cheap line per tool — name plus a one-sentence purpose — and keep the full parameter documentation, examples and usage rules out of context until the agent activates that tool. Most turns then pay for metadata only, not for the whole catalog.

open as a page

When should a tool parameter be required versus optional in its JSON Schema?

level: middleimportance: should knowfreq 54%

basics

~20 s

Mark a parameter required when the tool cannot do anything useful without it and the value must come from the conversation. Every optional field is an invitation for the model to invent a plausible value, so keep optionals few, describe them tightly, and default the rest in your handler.

open as a page

What does strict mode on a tool schema guarantee, and what does it cost?

level: middleimportance: should knowfreq 44%

basics

~20 s

Strict or structured-output modes constrain decoding so emitted arguments are guaranteed to match the declared schema — right keys, right types, no invented extras. They guarantee shape only, not truth, and they buy that with a restricted JSON Schema subset and some first-use latency.

open as a page

Why can't you act on a streamed tool call's arguments before the stream ends?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Arguments arrive as incremental fragments of a JSON string, so the buffer is syntactically invalid until the final chunk lands. Nothing can be parsed, validated or executed from a partial call — only the tool name, which arrives first, is usable early.

open as a page

Which tool failures should your wrapper retry in code, and which should the model see?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Retry mechanical, transient failures in code — rate limits, connection resets, brief timeouts — with bounded backoff, and only when the call is safe to repeat. Surface anything the model could decide differently about: bad arguments, missing records, denied permissions, exhausted retries.

open as a page

How do namespaces and toolsets stop overlapping agent tools from being selected wrongly?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Namespaces prefix each tool with its owning domain, so two similar capabilities read as different things and the model can see who owns which intent. Toolsets go further and keep the colliding pair out of the same prompt by loading only the tools eligible for the current role or request class.

open as a page

Your agent's tool schemas cost 2,100 tokens per request — how do you cut that?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Tool definitions are re-sent on every request, so their cost is per-turn, not one-off. Cut by removing parameters the handler can supply, flattening deep nesting, compressing per-field prose to decision-relevant text, and splitting overloaded tools — then verify accuracy did not drop.

open as a page

When should an agent call tools from sandbox code instead of one call per turn?

level: principalimportance: should knowfreq 34%

basics

~20 s

When the work is a wide, mechanical fan-out with known control flow. Writing a loop that invokes the tool inside a sandbox costs one model turn instead of hundreds, while conversational calls remain right for short, exploratory work where each result changes the next decision.

open as a page

How would you measure tool-selection accuracy as teams keep adding tools to a shared agent?

level: principalimportance: should knowfreq 38%

basics

~20 s

Score selection on its own: a labelled set of single user turns, a fixed catalog, one tool call each, compared against the gold tool — including turns whose correct answer is to call nothing. Re-run it at several catalog sizes and gate new tools on it.

open as a page