How do you reconstruct tool_calls from an OpenAI streaming chat completion?
answer
- The array arrives in pieces
- There is a join key on every fragment
- Names come early, arguments come slowly
- Parse only once, at the end
- Truncation here is corruption, not brevity
basics
~20 sWith stream enabled, tool calls arrive as fragments in delta.tool_calls. Each fragment has an index; the first for that index carries the id and function name, later ones carry pieces of the arguments string you concatenate before parsing.
solid answer
~40 sStreaming does not give you a finished `tool_calls` array — it gives you deltas. Each chunk's `choices[0].delta.tool_calls` is a list of partial entries, and the **`index`** field is what identifies which call a fragment belongs to; ids do not repeat on every chunk. The first fragment at a given index carries `id` and `function.name`; subsequent fragments carry `function.arguments` as string pieces that you append in arrival order. So you accumulate into a map keyed by index — `{id, name, args_buffer}` — and only once the stream ends do you have JSON complete enough to parse. Attempting `json.loads` on a fragment fails, since a fragment can end mid-key. The final chunk carries `finish_reason: "tool_calls"`, which is your signal to stop accumulating, parse each buffer and dispatch.
code
python · 41 linesimport json, os
from openai import OpenAI
client = OpenAI()
MODEL = os.environ["OPENAI_MODEL"]
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current temperature for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
acc, reason = {}, None
stream = client.chat.completions.create(
model=MODEL,
messages=[{"role": "user", "content": "Weather in Oslo and Lisbon?"}],
tools=tools,
stream=True,
)
for chunk in stream:
choice = chunk.choices[0]
reason = choice.finish_reason or reason
for frag in choice.delta.tool_calls or []:
slot = acc.setdefault(frag.index, {"id": None, "name": None, "args": ""})
if frag.id:
slot["id"] = frag.id
if frag.function and frag.function.name:
slot["name"] = frag.function.name
if frag.function and frag.function.arguments:
slot["args"] += frag.function.arguments # fragments, not JSON
if reason == "tool_calls":
for slot in acc.values():
print(slot["id"], slot["name"], json.loads(slot["args"]))go deeper
Know that with streaming enabled the tool call does not arrive whole — it comes as fragments you have to stitch together before you can use it.
Describe the accumulator: a map keyed by index, id and name captured from the first fragment, arguments concatenated, parsed only after the stream closes.
Handle the failure modes — truncation under a length finish reason, mid-stream disconnects, interleaved prose and call deltas — and explain why a partially assembled call must be discarded rather than repaired.
Weigh whether streaming tool calls earns its complexity for a given product: what progress signal the user actually gets from an early function name, and what the retry semantics are when a half-streamed turn is thrown away.
## Why streaming complicates tool calls Streaming text is easy: append `delta.content` to a buffer and render. Tool calls are structured, and the API streams them the same token-by-token way — which means you receive a JSON object in pieces, out of a shape you can validate, until the very end. Any code that assumes a whole `tool_calls` array will crash the first time someone flips `stream=True`. ## The delta shape Each streamed chunk has `choices[0].delta`. During a tool call, `delta.content` is null and `delta.tool_calls` is a list of partial entries. A partial entry has: - **`index`** — the position of this call within the eventual `tool_calls` array. This is the join key. It is present on every fragment. - **`id`** — present only on the first fragment for that index. - **`function.name`** — likewise typically arrives with the first fragment. - **`function.arguments`** — a *fragment* of the JSON string, arriving across many chunks. So the accumulation logic is: for each fragment, look up (or create) a slot keyed by `index`; set `id` and `name` if the fragment carries them; append `function.arguments` to the slot's buffer if present. Nothing else is safe to assume — in particular, do not key on `id` (it is absent on most fragments) and do not rely on chunks for different indices arriving in contiguous runs, since with parallel calls enabled fragments for several indices can interleave. ## When the call is complete The last chunk for the choice carries `finish_reason`. `"tool_calls"` means the model has finished asking; `"stop"` means it produced prose instead; `"length"` means it hit the token ceiling — and that last one matters more here than for text, because a truncated arguments buffer is unparseable JSON, not a shortened sentence. Treat `"length"` during a tool call as a hard error: raise the ceiling or shrink the tool surface, do not try to repair the JSON. Only after the stream closes do you `json.loads` each accumulated buffer. Then reassemble the assistant message — with the full `tool_calls` array — and append it to your transcript exactly as you would in the non-streaming case, followed by one tool message per `tool_call_id`. The streaming path changes only how you obtain the call, not the protocol for answering it. ## What you can show the user The practical reason to stream tool calls at all is UI feedback. The function *name* arrives almost immediately, well before the arguments finish, so an agent UI can render "Searching the knowledge base…" within a few hundred milliseconds instead of after the whole call materializes. Streaming the partial arguments to the user is usually a mistake — half a JSON object is noise — but the name and the count of calls are genuinely useful progress signals. ## Interleaving with text A turn can begin with prose and then switch to tool calls, or come back after tool results and stream ordinary text. Your consumer therefore has to handle both `delta.content` and `delta.tool_calls` on every chunk, rather than branching once at the start of the stream. In an agent loop this repeats per iteration: stream, accumulate, dispatch, stream again — and the final iteration usually ends `"stop"` with content, which is what you render as the answer. ## Errors mid-stream A stream can fail after chunks have already been delivered — a dropped connection, or a server-side error emitted partway through. You are then holding a half-built arguments buffer that must be discarded, not executed: a partially-parsed call is exactly the case where a `delete` with a truncated filter does damage. Make the accumulator's output atomic — either every call parses and you dispatch the batch, or you discard the whole assistant turn and retry the request. Retrying is safe because you have not yet appended anything to the transcript. ## Practical checklist - Key accumulation on `index`, never on arrival order or `id`. - Concatenate arguments; never parse a fragment. - Treat `finish_reason: "length"` during a tool call as fatal. - Handle content and tool_calls deltas in the same loop. - Discard, don't salvage, a truncated call on stream failure. - Once assembled, rejoin the normal path: append the assistant message, then one tool message per id.
- Why key the accumulator on index rather than on the tool call id?Because the id is only sent on the first fragment for each call; every later fragment carries just the index and an arguments piece. Keying on id would leave you with no key for the bulk of the stream. Index is also what defines the call's position in the final tool_calls array, so accumulating by it reconstructs the array directly. With parallel calls, fragments for different indices can interleave, which is exactly the case ordering-based logic gets wrong.
- The stream ends with finish_reason 'length' partway through a tool call. What now?Treat it as a failed turn, not a partial success. The arguments buffer is truncated JSON, so parsing fails, and any attempt to repair it risks dispatching a call with a half-specified filter or path. Discard the assistant turn — nothing has been appended to the transcript yet — and retry with a higher output-token ceiling or a smaller tool surface. If it recurs, the schema is probably encouraging oversized argument payloads.
- Does streaming change what you append to the conversation afterwards?No. Once you have assembled the calls, the protocol is identical to the non-streaming case: append the assistant message carrying the complete tool_calls array, then one message with role 'tool' per tool_call_id, then request again. Streaming is purely a transport concern for obtaining the turn. The one practical difference is that you construct the assistant message yourself from the accumulator rather than passing back an object the SDK handed you.
saying these in an interview costs you the question
- Expecting a complete tool_calls array on the first streamed chunk
- Calling json.loads on an individual arguments fragment
- Matching fragments by arrival order instead of the index field
- Executing a call assembled from a stream that failed midway
- Handling only delta.content and ignoring delta.tool_calls