skip to content

How does AutoGen's CodeExecutorAgent decide which code from a message it actually runs?

level: middleimportance: must knowfreq 65%

answer

  1. it parses text, not tool calls
  2. fenced blocks with a language tag
  3. blocks run in order, one work directory
  4. sources narrows who can supply code
  5. unsupported language returns an error result

basics

~20 s

CodeExecutorAgent scans incoming messages for markdown fenced code blocks, keeps the ones whose language tag its executor supports, and runs them in order in the same working directory. The reply is the combined output plus an exit code.

solid answer

~50 s

The agent is a text-in, text-out wrapper around an executor. It reads the messages it receives, pulls out fenced markdown blocks, and hands the supported ones to the executor as code blocks with a language. The command-line executors handle `python` and the shell family (`bash`, `shell`, `sh`, `pwsh`, `powershell`, `ps1`); a block with no language tag or an unsupported one is not run and comes back as an error result rather than silently succeeding. Every block in the message runs in order, in the same `work_dir`, and the agent replies with one message containing the accumulated output and the exit code. If it finds no code blocks at all, it says so rather than doing nothing. Two arguments shape this: `sources` limits which agents' messages are scanned, and if you give the agent a `model_client` it can write the code itself and, with `max_retries_on_error`, retry after a nonzero exit.

code

python · 26 lines
python
import asyncio

from autogen_agentchat.agents import CodeExecutorAgent
from autogen_agentchat.messages import TextMessage
from autogen_core import CancellationToken
from autogen_ext.code_executors.local import LocalCommandLineCodeExecutor


async def main() -> None:
    executor = LocalCommandLineCodeExecutor(work_dir="coding", timeout=30)
    agent = CodeExecutorAgent("executor", code_executor=executor, sources=["coder"])
    message = TextMessage(
        content=(
            "Here is the script.\n"
            "```python\n"
            "# filename: hello.py\n"
            "print('hello from the executor')\n"
            "```"
        ),
        source="coder",
    )
    response = await agent.on_messages([message], CancellationToken())
    print(response.chat_message.content)


asyncio.run(main())

go deeper

for a junior

Know that the agent reads markdown fenced code blocks out of chat messages and runs them, and that the block needs a language tag such as python or bash to be executed at all.

for a middle

Explain the full path: extraction from message text, language filtering by the executor, sequential execution in one work_dir, and a single reply carrying combined output and exit code. Mention the # filename: convention and the sources argument.

for a senior

Show that you treat markdown parsing as an attack surface: constrain sources so retrieved or user text cannot supply code, keep supported languages minimal, and set max_retries_on_error deliberately because each retry is another model call plus another execution.

for a principal

Be ready to argue the tradeoff of a text-parsing execution interface against a typed tool contract — where prompt discipline substitutes for a schema, what that costs in reliability, and when you would put the execution capability behind a narrower interface entirely.

## The contract `CodeExecutorAgent` sits between a conversation and an executor. Its input is chat messages; its output is one chat message containing whatever running the code produced. Understanding it means understanding three steps: extraction, filtering, and reporting. ## Extraction: markdown is the interface There is no structured tool call here. The agent parses the *text* of the messages it receives and looks for fenced code blocks — triple-backtick sections with a language tag. That has direct consequences for prompt design: the coding agent's system message must tell the model to emit a single fenced block with an explicit language, because a model that describes code in prose, or fences it without a tag, produces nothing runnable. Most "my executor does nothing" bugs are exactly this. A useful convention the command-line executors honour is a first-line comment of the form `# filename: script.py`. When present, the block is written to that name in `work_dir` instead of a generated one, which lets a model create a module in one block and import it in the next, or write a file whose name matters (`requirements.txt`, `Dockerfile`). ## Filtering: language tags and sources Each extracted block becomes a code block object carrying `code` and `language`. The executor decides what it can run. The command-line executors accept `python` and the shell family — `bash`, `shell`, `sh` on Unix-like systems and `pwsh`, `powershell`, `ps1` for PowerShell. A block tagged `javascript`, or fenced with no tag at all, is not executed; you get a failing result explaining the language is unknown, which is important because it means the failure is *visible* in the conversation and a model with a retry budget can correct itself. The second filter is `sources`. By default the agent will consider code from any message it sees, which in a group chat means any participant — including, transitively, text a tool pulled off the internet. Setting `sources=["coder"]` restricts extraction to messages from named agents, and it is the cheapest hardening step available: it removes the path where a retrieved document containing a fenced block gets executed because it happened to flow through the conversation. ## Execution and reporting Blocks run sequentially, in the order they appeared, all in the same `work_dir`. That matters both ways: a block can rely on a file an earlier block wrote, but it cannot rely on an earlier block's in-memory variables, because each command-line execution is a fresh process. Execution stops being useful the moment one block fails; the result carries the exit code, and a nonzero code is how downstream logic (or the model itself) learns something broke. The agent's reply is a single text message containing the accumulated stdout and stderr plus the exit status. When nothing was extractable it replies that it found no code blocks — an explicit statement rather than an empty turn, so a group chat does not stall silently. ## Generating as well as executing Since the 0.5/0.6 line, `CodeExecutorAgent` can optionally take a `model_client`. With one, the agent both writes and runs the code, which collapses the classic two-agent "coder plus executor" pattern into a single participant. Pair it with `max_retries_on_error` (default 0) and the agent will feed a failing execution's output back to the model and try again, up to that many times, before giving up. That budget is worth setting deliberately: each retry is another model call and another execution, so an unbounded value turns a syntax error into a cost incident. ## What this design costs The markdown-parsing interface is why AutoGen's code execution feels fragile compared with a typed tool call: correctness depends on prompt discipline, and anything that *looks* like a fenced block is a candidate for execution. The mitigations are all at this layer — constrain `sources`, keep the system message explicit about fencing, keep the executor's supported languages minimal, and remember that the executor choice (host versus container) decides what a mistake actually costs.

  • A retrieval step drops a document containing a fenced bash block into the chat. What stops it from running?
    By default, nothing — the agent extracts fenced blocks from the messages it sees, and it does not care that the text originated from a fetched document. The fix is `sources`: name only the agents you trust to produce code, so messages from the retrieval participant are never scanned. Beyond that, keep the executor container-backed so a bad block costs you a disposable container rather than the host.
  • Why does a model that explains code in prose cause the executor to appear broken?
    Because extraction is purely syntactic: no triple-backtick fence with a supported language tag means no code block, and the agent replies that it found none. The remedy is in the coding agent's system message — instruct it to output exactly one fenced block with an explicit language and no commentary inside the fence. This is prompt discipline standing in for a typed tool schema.
  • What does max_retries_on_error buy you, and what does it cost?
    With a `model_client` attached, a nonzero exit code is fed back to the model so it can fix its own code, up to `max_retries_on_error` attempts (default 0, meaning no retry). It converts transient mistakes — a missing import, a typo — into self-healing turns. The cost is real: each retry is another model call plus another execution, so a generous budget on a persistently failing task multiplies both latency and spend.
  • If two blocks appear in one message, can the second use a variable from the first?
    Not with the command-line executors. Both blocks run in order in the same `work_dir`, but each is a separate process, so only files persist between them — in-memory variables do not. A model that splits a script across two blocks and expects continuity will fail with a NameError. Either instruct it to emit one block, or use a kernel-backed executor that keeps state alive.

saying these in an interview costs you the question

  • Thinking the agent receives structured tool calls rather than parsing text
  • Assuming an untagged fenced block still runs as Python
  • Believing variables carry from one block to the next
  • Not realising any participant's message can supply code by default
  • Expecting execution to continue normally after a nonzero exit code

context