skip to content

AutoGen

You will learn Microsoft's framework for conversational multi-agent systems: configurable agents and model clients, teams and group chat, sandboxed code execution, and the layering between the Core and AgentChat APIs. Interviewers ask because agents-talking-to-agents raises precisely the questions they care about: termination conditions, runaway cost, and executing generated code safely.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

explore

questions

28

In AutoGen, what is the difference between LocalCommandLineCodeExecutor and DockerCommandLineCodeExecutor?

level: juniorimportance: must knowfreq 80%

answer

  1. one runs on your machine
  2. trust boundary, not convenience
  3. container image and mounted work directory
  4. must be started and stopped
  5. the model's code inherits your env vars

basics

~20 s

LocalCommandLineCodeExecutor runs model-generated code as a subprocess on the host machine, with your files, environment variables and network. DockerCommandLineCodeExecutor runs the same code inside a container instead, isolating it from the host. Local is for trusted prototyping only.

solid answer

~50 s

Both implement the same executor interface, so a `CodeExecutorAgent` can use either without changing anything else — the difference is entirely the trust boundary. `LocalCommandLineCodeExecutor` writes each extracted code block to a file under `work_dir` and shells out to run it as a child process of your application: same user, same filesystem, same environment variables (including API keys), same network. There is no sandbox. `DockerCommandLineCodeExecutor` starts a container from an image (`python:3-slim` by default), bind-mounts `work_dir` into it, and runs the code there, so a rogue `rm -rf` or a stray file write hits the container, not your laptop. The Docker one has a lifecycle: you must start it (`await executor.start()` or `async with`) and stop it, and by default it stops and removes its container when you do. Both take a `timeout` that defaults to 60 seconds per execution.

code

python · 19 lines
python
import asyncio

from autogen_agentchat.agents import CodeExecutorAgent
from autogen_agentchat.messages import TextMessage
from autogen_core import CancellationToken
from autogen_ext.code_executors.docker import DockerCommandLineCodeExecutor


async def main() -> None:
    async with DockerCommandLineCodeExecutor(
        image="python:3-slim", work_dir="coding", timeout=60
    ) as executor:
        agent = CodeExecutorAgent("executor", code_executor=executor)
        message = TextMessage(content="```python\nprint(2 + 2)\n```", source="coder")
        response = await agent.on_messages([message], CancellationToken())
        print(response.chat_message.content)


asyncio.run(main())

go deeper

for a junior

Be able to name both classes and say plainly that one runs code on your machine and one runs it in a container. Mention that the Docker executor has to be started and stopped.

for a middle

Explain the mechanics: code blocks are written as files into work_dir, run with the matching interpreter, and come back with an exit code and output; the Docker variant bind-mounts work_dir and defaults to the python:3-slim image with a 60-second timeout.

for a senior

Show that you treat the choice as a trust decision. Say what the local executor inherits — user, filesystem, environment variables, network position — and describe how you inject the executor so development and production differ by configuration, not by agent code.

for a principal

Own the honest limit: a container reduces blast radius but shares the host kernel, and egress, capabilities and resource caps are platform concerns the executor's constructor never touches. Be ready to say what infrastructure you would be willing to lose.

## Why AutoGen has executors at all AutoGen's signature pattern is that one agent writes code and another actually runs it. "Actually runs it" has to happen somewhere concrete, and that choice is the whole security story. AutoGen expresses it as a pluggable executor: an object with `execute_code_blocks(...)` (plus `start()`, `stop()` and `restart()`) that a `CodeExecutorAgent` holds via its `code_executor` argument. Swapping executors changes where code lands without touching agent or team wiring — which is exactly why interviewers ask: the safe and the unsafe option are one constructor call apart. ## LocalCommandLineCodeExecutor Imported from `autogen_ext.code_executors.local`. For each code block it writes a file into `work_dir` and executes it with the matching interpreter — `python` for Python blocks, the shell for `bash`/`shell`/`sh`, PowerShell for `pwsh`/`powershell`/`ps1`. The result carries an `exit_code` and the combined output. The key fact is what the child process inherits: your working user's permissions, your whole filesystem, your environment variables (so every API key exported in that shell is readable by the generated code), and your network position — including anything reachable only from inside your VPN or cloud VPC. Nothing about "the model wrote it" makes that code trustworthy; a prompt-injected web page inside the conversation can end up as code in this executor. The class exists for fast local iteration on code you are willing to eyeball, and for CI where the runner is already disposable. It does offer one useful non-security knob: `virtual_env_context`, which runs the code inside a Python virtual environment you built, so model-installed packages do not pollute your interpreter. That is dependency hygiene, not isolation — the venv shares the same user and filesystem. ## DockerCommandLineCodeExecutor Imported from `autogen_ext.code_executors.docker`. It launches a container from `image` (default `python:3-slim`), bind-mounts `work_dir` into the container (at `/workspace`) so files the code writes are visible on the host afterwards, and runs each block inside. `bind_dir` exists for the case where the path AutoGen sees is not the path the Docker daemon should mount — for example when your app itself runs in a container. Unlike the local one it owns a resource, so it has a lifecycle. The idiomatic form is `async with DockerCommandLineCodeExecutor(...) as executor:`; otherwise call `await executor.start()` before use and `await executor.stop()` after. Forgetting this is the single most common beginner error and shows up as an error about the container not running. `auto_remove` and `stop_container` default to true, so the container is stopped and deleted when the executor stops — good hygiene, but it also means anything not written into `work_dir` is gone. What you get is process and filesystem isolation from the host, plus a reproducible interpreter: the model's code sees the packages in your image, not whatever your laptop happens to have. What you do not get from the constructor is a hardened sandbox — the container shares the host kernel, and egress rules, capability dropping and resource caps are set at the platform layer, not by executor arguments. Treat it as "blast radius reduced to a container I am willing to lose", not "safe to run anything". ## Choosing Use Local when you are the only source of the prompts, the machine is disposable, and startup latency matters. Use Docker the moment untrusted input can influence what gets generated — any user-facing product — and pin a purpose-built image so the code has its dependencies without needing to install anything at run time. Because the executor is injected, the usual shape is one factory that returns Local in development and Docker everywhere else; the agent code does not change. ## Related surface A third option, `JupyterCodeExecutor`, keeps a live kernel so state persists between blocks, and a Docker-backed Jupyter variant combines that with container isolation. Whichever you pick, `timeout` (default 60 seconds) bounds a single execution — not the number of executions an agent loop can request.

  • What happens if you pass a DockerCommandLineCodeExecutor to an agent without starting it?
    Execution fails: the executor has no running container to exec into, and you get an error rather than a code result. Either wrap it in `async with`, which starts and stops it around the block, or call `await executor.start()` yourself and `await executor.stop()` in a finally. Because `auto_remove` and `stop_container` default to true, the container is also destroyed on stop, so anything outside the mounted `work_dir` does not survive.
  • Does switching from the local executor to the Docker one require changing the agent code?
    No — that is the point of the executor abstraction. `CodeExecutorAgent` takes any object satisfying the executor interface through `code_executor`, so the swap is one construction site. The practical differences are lifecycle (the Docker executor must be started and stopped) and image contents: the generated code now sees the packages baked into your image rather than your host interpreter's, so pin an image that has what your agents typically need.
  • If the generated code writes a CSV file, where does it end up in each case?
    With the local executor it lands in `work_dir` on your filesystem directly. With the Docker executor the code writes inside the container, but `work_dir` is bind-mounted in, so a file written to the container's working directory appears in the host `work_dir` too and survives the container being removed. Anything written elsewhere in the container is lost when the container is stopped and auto-removed.

The local executor is handing a stranger the keys to your desk; the Docker one gives them a rented booth you can burn down afterwards.

saying these in an interview costs you the question

  • Assuming the local executor sandboxes code in some way
  • Believing generated code is safe because a model wrote it
  • Using the Docker executor without start/stop or async with
  • Claiming the container blocks all network access by default
  • Treating work_dir as an isolation boundary rather than a shared folder

context

open as a page

In AutoGen, how does RoundRobinGroupChat pick the next speaker and what does each agent see?

level: juniorimportance: must knowfreq 68%

basics

~20 s

RoundRobinGroupChat cycles participants in the exact order passed to its constructor, one response per turn, with no model deciding anything. Every response is broadcast to all participants, so each agent sees the whole shared conversation.

open as a page

How do you configure AutoGen's OpenAIChatCompletionClient for an unknown model?

level: middleimportance: must knowfreq 56%

basics

~20 s

Pass an explicit model_info dictionary alongside model and base_url. The client keeps a capability table for models it recognizes; for anything outside it — a self-hosted or newly released model behind an OpenAI-compatible endpoint — it cannot infer capabilities and raises unless you declare them.

open as a page

In AutoGen, what happens in one AssistantAgent turn when the model calls tools?

level: middleimportance: must knowfreq 70%

basics

~20 s

AssistantAgent makes one model call with its system message plus its context, executes every requested tool, then by default returns a ToolCallSummaryMessage holding the raw tool results. It only sends those results back to the model when reflect_on_tool_use is True.

open as a page

In AutoGen v0.4+, what do autogen-core, autogen-agentchat and autogen-ext each provide?

level: middleimportance: must knowfreq 72%

basics

~20 s

autogen-core is the event-driven actor runtime plus base abstractions. autogen-agentchat is the opinionated agent-and-team API built on top of it. autogen-ext holds the concrete integrations: model clients, tool adapters, code executors and the gRPC runtime.

open as a page

How does AutoGen's CodeExecutorAgent decide which code from a message it actually runs?

level: middleimportance: must knowfreq 65%

basics

~20 s

CodeExecutorAgent scans incoming messages for markdown fenced code blocks, keeps the ones whose language tag its executor supports, and runs them in order in the same working directory. The reply is the combined output plus an exit code.

open as a page

In AutoGen, what does Console(team.run_stream(...)) give you over awaiting team.run()?

level: middleimportance: must knowfreq 68%

basics

~20 s

run_stream yields each agent message and event as it is produced; Console consumes that async generator and prints items as they arrive. run() blocks and returns only the final TaskResult. Console returns that same TaskResult at the end.

open as a page

What does AutoGen's TaskResult contain, and where do you read a run's token usage?

level: middleimportance: must knowfreq 62%

basics

~20 s

TaskResult holds messages, the full ordered list of the run's messages and events, and stop_reason, a string saying why the run ended. Token usage is not a top-level field: each message carries models_usage with prompt_tokens and completion_tokens, and you sum those.

open as a page

In AutoGen AgentChat, how do termination conditions end a team run, and what do | and & do?

level: middleimportance: must knowfreq 84%

basics

~20 s

Termination conditions from autogen_agentchat.conditions are evaluated as messages flow through a team; when one fires the team stops and run() returns a TaskResult whose stop_reason names it. The | operator means stop when either condition fires; & means stop only once all of them have.

open as a page

How do you persist and resume an AutoGen team across restarts with save_state and load_state?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Call await team.save_state() to get a JSON-serializable mapping of the conversation state, store it, then rebuild an identically configured team in the new process and call await team.load_state(state). Configuration — model clients, tools, system messages — is not in the state and must be reconstructed in code.

open as a page

What does AutoGen's UserProxyAgent do, and when should you use it?

level: juniorimportance: should knowfreq 42%

basics

~20 s

UserProxyAgent represents a human as a participant. It has no model client and no system prompt: when it is its turn, it calls its input_func to obtain a human reply and returns that as a message. By default it blocks on console input.

open as a page

In AutoGen, how does AssistantAgent.on_messages() differ from run()?

level: middleimportance: should knowfreq 52%

basics

~20 s

on_messages() is the low-level agent protocol: pass a sequence of messages plus a CancellationToken and get back a Response holding chat_message and inner_messages. run() is the task-runner wrapper: pass a task string and get a TaskResult with the full message list and a stop_reason.

open as a page

In autogen-core, how does RoutedAgent decide which @message_handler runs?

level: middleimportance: should knowfreq 48%

basics

~20 s

RoutedAgent dispatches on the type annotation of each handler's message parameter, so one handler per message type. An optional match predicate disambiguates several handlers of the same type, and an unhandled type falls through to on_unhandled_message.

open as a page

In AutoGen, what does JupyterCodeExecutor persist between executions that the command-line executors do not?

level: middleimportance: should knowfreq 45%

basics

~20 s

JupyterCodeExecutor keeps a live kernel, so imports, variables and loaded data survive from one execution to the next. The command-line executors start a fresh process each time, so only files written into work_dir carry over.

open as a page

How does ListMemory change what an AutoGen AssistantAgent sends to the model?

level: middleimportance: should knowfreq 45%

basics

~20 s

An AssistantAgent given memory=[ListMemory(...)] calls each memory's update_context before every model call, and ListMemory appends all of its stored MemoryContent items to that call's context. It ignores the query, so the injected block grows with everything you have added.

open as a page

In AutoGen's Swarm team, how does one agent actually hand control to another agent?

level: middleimportance: should knowfreq 54%

basics

~20 s

You configure an agent with handoffs=["other_agent"], which gives its model a transfer_to_other_agent tool. Calling that tool makes the agent emit a HandoffMessage whose target names the next speaker, and the Swarm team makes that agent speak next.

open as a page

How do you stop an AutoGen AssistantAgent's model context from growing unbounded?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Pass a bounded model_context when constructing the agent — BufferedChatCompletionContext keeps only the last N messages, HeadAndTailChatCompletionContext keeps the first and last slices. The default is UnboundedChatCompletionContext, which re-sends the entire history on every model call.

open as a page

In autogen-core, when do you publish_message to a topic instead of send_message?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Use send_message for request/response to one known AgentId — it awaits a reply. Use publish_message when any number of subscribed agents should react and you want no reply and no knowledge of who is listening; the publisher does not receive its own message.

open as a page

How do you migrate an AutoGen v0.2 ConversableAgent codebase to v0.4+?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Treat it as a rewrite against a new API, not an upgrade. Imports move to autogen-agentchat / autogen-core / autogen-ext, llm_config dicts become model client objects, initiate_chat becomes async run on an agent or team, and implicit stopping becomes explicit termination conditions.

open as a page

How do you require human approval before AutoGen's CodeExecutorAgent runs generated code?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Pass an approval callback to CodeExecutorAgent via its approval_func argument. Before each execution the agent calls it with the pending code and conversation context; returning a denial stops the run and the reason goes back into the conversation instead of any output.

open as a page

How do you wire OpenTelemetry tracing into an AutoGen team, and what do the spans show?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Build an OpenTelemetry TracerProvider with an exporter, pass it as tracer_provider to a SingleThreadedAgentRuntime, and hand that runtime to the team constructor. The runtime then emits spans for message dispatch and agent processing, so one multi-agent run becomes a single nested trace.

open as a page

Why does an AutoGen team's second run() see the first task's history, and what does team.reset() do?

level: seniorimportance: should knowfreq 45%

basics

~20 s

AutoGen teams are stateful: the group-chat manager and every participant keep the message thread between run() calls, so a second task is appended to the first conversation. await team.reset() clears that accumulated state on the team and its participants, returning them to their initial condition.

open as a page

In AutoGen's SelectorGroupChat, how is the next speaker chosen and how do you override it?

level: seniorimportance: should knowfreq 56%

basics

~20 s

SelectorGroupChat calls its model_client every turn with a rendered selector_prompt and asks it to name the next participant. Supplying selector_func short-circuits that: return a participant name to force the choice, or None to fall back to the model.

open as a page

When do you build directly on autogen-core instead of autogen-agentchat?

level: principalimportance: should knowfreq 38%

basics

~20 s

Drop to autogen-core when the conversation-and-team model stops matching the problem: event-driven topologies, agents deployed as separate services, non-chat message types, or bespoke control flow. The cost is that you write the message protocol, routing and stopping rules yourself.

open as a page

How would you run AutoGen code execution safely when untrusted users drive the prompts?

level: principalimportance: should knowfreq 32%

basics

~20 s

Never the local executor. Give each session its own container-backed executor with a purpose-built image, its own working directory, a tight per-execution timeout, and a bounded retry budget — then accept that the container is a blast-radius control, not a security boundary.

open as a page

How do you bound cost and runtime for an AutoGen team running unattended in production?

level: principalimportance: should knowfreq 34%

basics

~20 s

Layer the stops: one semantic condition that marks real success, ORed with hard fuses on messages, turns, tokens and wall clock, plus limits outside the framework at the model client. Then alert on TaskResult.stop_reason, because a fuse firing routinely means the design is not converging.

open as a page

What changes when an autogen-core app moves to the distributed gRPC runtime?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Agents stop sharing a process: messages must be serializable and serializers registered, a host process relays traffic and subscriptions between workers, and shared Python objects, in-process exceptions and instant delivery all stop being available. The gRPC runtime is experimental.

open as a page

In a multi-user AutoGen service, how do you decide what to persist and what to discard per session?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Separate three stores: configuration rebuilt from code or component definitions, resume state from save_state kept per session and bounded, and the transcript kept for audit under its own retention. Persist only what a resume genuinely needs; everything else is observability data.

open as a page