skip to content

Tool & Function Calling

How you let the model call your code: you describe functions with JSON Schema, the model answers with tool_calls instead of prose, you execute them and feed the results back. The interview angle is the loop — who runs the tool, who validates the arguments, and how you stop it iterating forever.

on this pageshow

questions

6

How do you define a tool in OpenAI's Chat Completions API and return its result?

level: middleimportance: must knowfreq 78%

answer

  1. The model requests, your code executes
  2. Two message shapes go back
  3. Arguments arrive as text, not objects
  4. Every id needs an answer
  5. role 'tool' plus tool_call_id

basics

~20 s

You send a tools array whose entries describe a function with a JSON Schema. The model replies with message.tool_calls instead of prose; your code runs each call, then appends a role 'tool' message carrying the matching tool_call_id.

solid answer

~40 s

Each entry of `tools` is `{"type": "function", "function": {"name", "description", "parameters"}}`, where `parameters` is a JSON Schema object. When the model decides to use one, the response comes back with `finish_reason: "tool_calls"`, `message.content` usually null, and `message.tool_calls` — a list of `{id, type, function: {name, arguments}}`. **`arguments` is a JSON-encoded string, not a dict**, so you parse it yourself. Your application executes the function; OpenAI never runs any code for you. You then append the assistant message verbatim (tool_calls intact) and, for every call, a `{"role": "tool", "tool_call_id": <id>, "content": <string>}` message, and call the API again with the same tools. The model reads the results and produces the user-facing answer.

code

python · 36 lines
python
import json, os
from openai import OpenAI

client = OpenAI()
MODEL = os.environ["OPENAI_MODEL"]

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Current temperature for a city, in Celsius.",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string", "description": "City name"}},
            "required": ["city"],
            "additionalProperties": False,
        },
    },
}]

messages = [{"role": "user", "content": "What is the weather in Oslo?"}]
first = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)
msg = first.choices[0].message
messages.append(msg)

for call in msg.tool_calls or []:
    args = json.loads(call.function.arguments)   # arguments is a JSON string
    result = {"city": args["city"], "temp_c": 7}
    messages.append({
        "role": "tool",
        "tool_call_id": call.id,
        "content": json.dumps(result),
    })

second = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)
print(second.choices[0].message.content)

go deeper

for a junior

Be able to say plainly that the model only asks for a function to be called and your code runs it, and that the result goes back as a message with role 'tool'.

for a middle

Walk the full round trip from memory: tools array with JSON Schema, finish_reason tool_calls, parsing the arguments string, appending the assistant message plus one tool message per tool_call_id.

for a senior

Show how you keep the transcript valid while trimming history, how you validate model-supplied arguments before they reach a system of record, and how you surface tool errors as content so the model can recover.

for a principal

Own the contract: who defines tool schemas, how they are versioned alongside the services behind them, and what the blast radius is when a model calls a mutating tool with plausible but wrong arguments.

## What tool calling actually is Tool calling (historically "function calling") does not let the model execute anything. It is a structured-output protocol: you describe the functions your application can run, and instead of prose the model emits a machine-readable request naming one of them plus arguments. Every side effect — the HTTP call, the SQL query, the file write — happens in your process, under your auth and your validation. That split is the point of most interview questions on this topic. ## Step 1 — describing the tools On a Chat Completions request you pass `tools`, a list where each element is: `{"type": "function", "function": {"name": ..., "description": ..., "parameters": {JSON Schema}}}` `name` is the identifier you will dispatch on. `description` is prompt text the model reads — it is the single strongest signal for *when* to pick this tool, so write it for a reader, not as an afterthought. `parameters` is a JSON Schema object (`type: "object"` with `properties` and `required`); per-property `description` strings matter for the same reason. A function that takes no arguments still needs an object schema with empty `properties`. An optional `strict` flag on the function definition asks the API to constrain generated arguments to the schema rather than merely suggesting it. The whole `tools` block is serialized into the prompt, so it costs input tokens on every request that carries it. ## Step 2 — reading the response When the model wants a tool, the choice comes back with `finish_reason: "tool_calls"` and `message.tool_calls` populated; `message.content` is typically null. Each element looks like `{"id": "call_abc", "type": "function", "function": {"name": "get_weather", "arguments": "{\"city\":\"Oslo\"}"}}`. Two details trip people up. First, `arguments` is a **string** containing JSON, not a parsed object — `json.loads` it. Second, that string is model-generated: it can be malformed, can carry values outside your enum, and can name entities that do not exist. Parse defensively and validate against your own schema or model class before touching a database. Dispatch by looking `function.name` up in an explicit map of allowed handlers. Never resolve the name dynamically against your module namespace. ## Step 3 — feeding results back The follow-up request is the same conversation plus two kinds of new messages: 1. The assistant message you just received, appended **verbatim** with its `tool_calls` array intact. If you drop it, the tool messages have nothing to attach to. 2. One message per call: `{"role": "tool", "tool_call_id": call.id, "content": "..."}`. `content` must be a string, so serialize structured results with `json.dumps`. There is no `name` requirement — the `tool_call_id` is what binds the result to the request. Send that with the same `tools` list (the model may want another round), and the next response is usually ordinary prose that incorporates the tool output. ## What the API enforces Every `tool_call_id` in an assistant message must be answered by a tool message before the next completion, and the tool messages must follow that assistant message in the array. Omit one and the request fails with a 400 rather than degrading gracefully. Likewise a tool message with no preceding assistant `tool_calls` is rejected. This makes conversation history a small state machine you have to keep consistent — a common source of bugs once you start trimming old turns to control context growth, because trimming an assistant tool_calls message without its tool replies (or vice versa) corrupts the transcript. ## Errors and results are both just content If your handler throws, the idiomatic recovery is to put a short error description into the tool message content — `{"error": "city not found"}` — rather than raising out of the loop. The model can then apologise, ask the user, or retry with different arguments. Return the raw exception text only if it carries no secrets. ## Common mistakes - Treating `arguments` as a dict and getting an attribute error at runtime. - Forgetting to append the assistant message, so the tool replies are orphaned. - Passing a dict as tool `content` instead of a JSON string. - Returning the tool result straight to the user instead of letting the model phrase it — you lose the synthesis step the loop exists for. - Assuming the model called the tool because it was there: with the default `tool_choice: "auto"` it may just answer in prose, so always branch on whether `tool_calls` is present.

  • What does setting strict to true on a function definition buy you?
    It asks the API to constrain the generated arguments to your JSON Schema rather than treating it as a hint, so you stop seeing missing required fields or invented property names. It comes with schema restrictions — notably that every property is listed in required and objects disallow additional properties — so schemas must be written for it deliberately. It does not validate semantics: a well-formed city name can still be a city that does not exist.
  • The model returns a tool call whose arguments fail to parse as JSON. What do you do?
    Do not crash the loop. Append a tool message for that tool_call_id whose content says the arguments were unparseable, and let the model retry — it usually reissues the call correctly. Cap the retries per call so a persistently broken tool cannot spin. If it repeats, that is a signal the schema is ambiguous or too permissive; tighten it or enable strict argument generation.
  • Do you have to resend the tools array on the follow-up request?
    Yes, if you want the model to be able to call anything again. The API is stateless per request: whatever tools are not in this request do not exist for this turn. Dropping them on the second call is a legitimate technique to force a prose answer, but it is a deliberate choice, not the default — and it changes the prompt prefix, which costs you any caching of the stable tool block.

saying these in an interview costs you the question

  • Claiming OpenAI executes your function on its servers
  • Treating tool_calls[].function.arguments as an already-parsed dict
  • Forgetting to append the assistant message before the tool replies
  • Trusting model-supplied arguments straight into a query or shell
  • Assuming a tool is always called just because tools were provided

context

open as a page

What do OpenAI's tool_choice values auto, required and none do?

level: middleimportance: must knowfreq 62%

basics

~20 s

tool_choice auto lets the model decide and is the default when tools are present; required forces at least one tool call this turn; none forbids calling while still sending the definitions; and a named function object pins the call to that one function.

open as a page

How do parallel tool calls work in OpenAI's API, and when do you disable them?

level: middleimportance: should knowfreq 50%

basics

~20 s

One assistant message can carry several entries in tool_calls; run them concurrently, then append one tool message per tool_call_id before the next request. Set parallel_tool_calls to false when calls must be ordered or have side effects.

open as a page

How do you keep an OpenAI tool-calling agent loop from running forever?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Loop until a response has no tool_calls and finish_reason is stop, and bound it: a maximum step count, a token or cost budget, per-tool timeouts, and duplicate-call detection. Force a prose answer with tool_choice none on the final step.

open as a page

How do you reconstruct tool_calls from an OpenAI streaming chat completion?

level: seniorimportance: should knowfreq 40%

basics

~20 s

With stream enabled, tool calls arrive as fragments in delta.tool_calls. Each fragment has an index; the first for that index carries the id and function name, later ones carry pieces of the arguments string you concatenate before parsing.

open as a page

How many tools should one OpenAI request expose, and what breaks with too many?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Tool definitions are serialized into the prompt and billed as input tokens on every request, and selection accuracy falls as similar tools multiply. Keep the per-request set small and distinct; route or split across sub-agents rather than growing one surface.

open as a page