skip to content

How does a LangChain agent turn an AIMessage tool call into a ToolMessage?

level: middleimportance: must knowfreq 65%

answer

  1. three message types and one id
  2. the join key between call and result
  3. providers reject orphaned pairs
  4. every turn re-sends the whole list
  5. no regex parsing of prose any more

basics

~20 s

The model returns an AIMessage carrying tool_calls, each with a name, an args dict and an id. The agent looks the name up in its tool list, invokes it with args, and appends a ToolMessage whose tool_call_id matches, then calls the model again.

solid answer

~50 s

The loop is entirely message-shaped. `model.bind_tools(tools)` attaches the tool specifications to every request. When the model decides to act, it returns an `AIMessage` whose `.tool_calls` is a list of dicts with `name`, `args` and `id`. The agent resolves each `name` against its tool registry, invokes the tool with `args`, and appends a `ToolMessage` carrying the result plus a `tool_call_id` equal to the original call's `id`. That id is the join key — providers reject a follow-up request where a tool call has no matching result, or a result has no matching call. Those messages are then sent back as the next request, and the model either emits more tool calls or a plain content-only `AIMessage`, which terminates the loop. In LangChain v1 `create_agent` runs exactly this cycle for you over a `messages` state; the important consequence is that every tool result is permanently appended to the conversation, so verbose tools inflate cost on every subsequent turn.

code

python · 24 lines
python
from langchain_core.messages import HumanMessage, ToolMessage
from langchain_core.tools import tool


@tool
def word_count(text: str) -> int:
    """Count the words in a piece of text."""
    return len(text.split())


# Shape of what the model returns, and how it is answered.
tool_call = {
    "name": "word_count",
    "args": {"text": "a b c"},
    "id": "call_001",
    "type": "tool_call",
}

result: ToolMessage = word_count.invoke(tool_call)
print(result.tool_call_id, result.content)

messages = [HumanMessage("how many words in 'a b c'?")]
# messages.append(ai_message_with_tool_calls)
messages.append(result)

go deeper

for a junior

Know the three message types by name: an AIMessage can carry tool_calls, each answered by a ToolMessage, and the ids must match.

for a middle

Walk the full cycle out loud — bind_tools, tool_calls with name/args/id, tool invocation, ToolMessage with tool_call_id, resend — and say what ends the loop.

for a senior

Bring the operational consequences: orphaned call/result pairs cause provider 400s, parallel calls are a real concurrency surface, and tool output size compounds because every turn re-sends the list.

for a principal

Frame context growth as the scaling limit of agentic systems and own the policy: what tools may return, where artifacts live instead of the transcript, and how history is compacted without breaking call/result grouping.

## The loop is a message protocol, not magic Everything an agent does is visible in the message list. Understanding the three message types and the id that binds them is enough to debug almost any agent problem. 1. **Binding.** `model.bind_tools(tools)` returns a model runnable that serialises each tool's name, description and JSON argument schema into the provider's tool-calling field on every request. Nothing is executed at bind time. 2. **The model's move.** The provider responds either with content (a normal answer) or with one or more tool calls. LangChain normalises the provider-specific shape into `AIMessage.tool_calls`: a list of dicts with `name`, `args` (already parsed into a Python dict), `id` and `type`. Calls whose arguments failed to parse land in `AIMessage.invalid_tool_calls` instead, each carrying an `error` string. 3. **Execution.** The agent maps `name` to a tool object and invokes it. Passing the whole tool-call dict — `tool.invoke(tool_call)` — is the idiomatic form because it returns a `ToolMessage` already stamped with the right `tool_call_id`. 4. **The result message.** `ToolMessage(content=..., tool_call_id=...)` goes back into the message list. It also has a `status` field that can be set to `"error"` to mark a failed call, and a `name` field for readability. 5. **Repeat.** The full list — system, human, assistant-with-tool-calls, tool results — is sent again. The model sees its own previous calls and their outcomes, which is what lets it chain, retry with different arguments, or stop. ## Why the tool_call_id matters so much Providers validate the pairing. Every tool call in an assistant message must be answered by exactly one tool result before the next assistant turn, and every tool result must reference a real call. Break that invariant and you get a hard 400 from the provider, not a degraded answer. This is the number one cause of mysterious agent crashes when people hand-roll message trimming: a summarisation or truncation step drops the `AIMessage` that contained the calls while leaving the orphaned `ToolMessage` behind, or vice versa. Any trimming logic in an agent must move whole call-and-result groups. ## Parallel tool calls A single `AIMessage` can carry several tool calls; modern providers emit them routinely when the requests are independent. Each needs its own `ToolMessage`, and the agent may execute them concurrently. That is a real concurrency surface in your code: two tools running against the same session or the same non-thread-safe client will race. It also means the model committed to all of those calls without seeing any of their results, so a failure in one does not stop the others. ## What this costs The loop is stateless from the provider's point of view: every turn re-sends the entire accumulated message list. A tool that returns a 40 KB JSON blob is not paid for once — it is paid for on every subsequent turn of that run. Three practical consequences: return the smallest useful projection from tools rather than raw API payloads; use `response_format="content_and_artifact"` when your code needs the full payload but the model does not; and treat context growth as the main scaling limit of long agent runs, addressed by summarising middleware or by keeping runs short. ## In LangChain v1 `create_agent(model, tools)` from `langchain.agents` builds a graph that runs precisely this cycle and returns a runnable you invoke with `{"messages": [...]}`. The result contains the full message list, so the trace of what happened is the output — reading `tool_calls` and `ToolMessage` contents in that list is the primary debugging technique, before any tracing tool. The loop terminates when the model returns an `AIMessage` with content and no tool calls, or when a limit fires. ## Where people go wrong Assuming the framework parses free-text "Action: search" out of the model's prose is the classic outdated mental model — that was the pre-1.0 text-scraped ReAct parser, and it is why those agents broke on formatting drift. Today the structured `tool_calls` field comes from the provider, and no regex is involved. Equally common is assuming the model "knows" a tool ran: it knows only because the `ToolMessage` is in the list you send. Drop it and the model will happily call the same tool again.

  • What breaks if a context-trimming step removes the AIMessage but keeps its ToolMessage?
    The next request contains a tool result referencing a call the provider cannot see, and most providers reject it outright with a 400. Trimming inside an agent must operate on whole groups: an assistant message and every tool result answering its calls move or stay together. This is why naive last-N-messages truncation is unsafe in agents even though it is fine for plain chat.
  • A single AIMessage carries three tool calls. What are the implications?
    The model committed to all three before seeing any results, so they must be independent. Each needs its own ToolMessage keyed by its own id, and executing them concurrently exposes any shared, non-thread-safe client or session to a real race. Failure of one does not cancel the others, so you must decide whether to return partial results with error statuses or fail the whole turn.
  • How do you tell a malformed tool call from a tool that failed?
    A malformed call — arguments that did not parse or did not match the schema — appears in `AIMessage.invalid_tool_calls` with an `error` field, and never reaches your code. A tool that ran and failed produces an exception in your process, which you surface as a ToolMessage with `status="error"`. The first is a model or schema problem; the second is a systems problem, and they need different fixes.

saying these in an interview costs you the question

  • Thinking LangChain regex-parses 'Action:' lines out of model prose
  • Mismatching or omitting tool_call_id and blaming the provider
  • Assuming the model remembers a tool ran without the ToolMessage
  • Believing tool output is charged only once rather than every turn
  • Trimming individual messages inside an agent's history

context