skip to content

In AutoGen, what is the difference between LocalCommandLineCodeExecutor and DockerCommandLineCodeExecutor?

level: juniorimportance: must knowfreq 80%

answer

  1. one runs on your machine
  2. trust boundary, not convenience
  3. container image and mounted work directory
  4. must be started and stopped
  5. the model's code inherits your env vars

basics

~20 s

LocalCommandLineCodeExecutor runs model-generated code as a subprocess on the host machine, with your files, environment variables and network. DockerCommandLineCodeExecutor runs the same code inside a container instead, isolating it from the host. Local is for trusted prototyping only.

solid answer

~50 s

Both implement the same executor interface, so a `CodeExecutorAgent` can use either without changing anything else — the difference is entirely the trust boundary. `LocalCommandLineCodeExecutor` writes each extracted code block to a file under `work_dir` and shells out to run it as a child process of your application: same user, same filesystem, same environment variables (including API keys), same network. There is no sandbox. `DockerCommandLineCodeExecutor` starts a container from an image (`python:3-slim` by default), bind-mounts `work_dir` into it, and runs the code there, so a rogue `rm -rf` or a stray file write hits the container, not your laptop. The Docker one has a lifecycle: you must start it (`await executor.start()` or `async with`) and stop it, and by default it stops and removes its container when you do. Both take a `timeout` that defaults to 60 seconds per execution.

code

python · 19 lines
python
import asyncio

from autogen_agentchat.agents import CodeExecutorAgent
from autogen_agentchat.messages import TextMessage
from autogen_core import CancellationToken
from autogen_ext.code_executors.docker import DockerCommandLineCodeExecutor


async def main() -> None:
    async with DockerCommandLineCodeExecutor(
        image="python:3-slim", work_dir="coding", timeout=60
    ) as executor:
        agent = CodeExecutorAgent("executor", code_executor=executor)
        message = TextMessage(content="```python\nprint(2 + 2)\n```", source="coder")
        response = await agent.on_messages([message], CancellationToken())
        print(response.chat_message.content)


asyncio.run(main())

go deeper

for a junior

Be able to name both classes and say plainly that one runs code on your machine and one runs it in a container. Mention that the Docker executor has to be started and stopped.

for a middle

Explain the mechanics: code blocks are written as files into work_dir, run with the matching interpreter, and come back with an exit code and output; the Docker variant bind-mounts work_dir and defaults to the python:3-slim image with a 60-second timeout.

for a senior

Show that you treat the choice as a trust decision. Say what the local executor inherits — user, filesystem, environment variables, network position — and describe how you inject the executor so development and production differ by configuration, not by agent code.

for a principal

Own the honest limit: a container reduces blast radius but shares the host kernel, and egress, capabilities and resource caps are platform concerns the executor's constructor never touches. Be ready to say what infrastructure you would be willing to lose.

## Why AutoGen has executors at all AutoGen's signature pattern is that one agent writes code and another actually runs it. "Actually runs it" has to happen somewhere concrete, and that choice is the whole security story. AutoGen expresses it as a pluggable executor: an object with `execute_code_blocks(...)` (plus `start()`, `stop()` and `restart()`) that a `CodeExecutorAgent` holds via its `code_executor` argument. Swapping executors changes where code lands without touching agent or team wiring — which is exactly why interviewers ask: the safe and the unsafe option are one constructor call apart. ## LocalCommandLineCodeExecutor Imported from `autogen_ext.code_executors.local`. For each code block it writes a file into `work_dir` and executes it with the matching interpreter — `python` for Python blocks, the shell for `bash`/`shell`/`sh`, PowerShell for `pwsh`/`powershell`/`ps1`. The result carries an `exit_code` and the combined output. The key fact is what the child process inherits: your working user's permissions, your whole filesystem, your environment variables (so every API key exported in that shell is readable by the generated code), and your network position — including anything reachable only from inside your VPN or cloud VPC. Nothing about "the model wrote it" makes that code trustworthy; a prompt-injected web page inside the conversation can end up as code in this executor. The class exists for fast local iteration on code you are willing to eyeball, and for CI where the runner is already disposable. It does offer one useful non-security knob: `virtual_env_context`, which runs the code inside a Python virtual environment you built, so model-installed packages do not pollute your interpreter. That is dependency hygiene, not isolation — the venv shares the same user and filesystem. ## DockerCommandLineCodeExecutor Imported from `autogen_ext.code_executors.docker`. It launches a container from `image` (default `python:3-slim`), bind-mounts `work_dir` into the container (at `/workspace`) so files the code writes are visible on the host afterwards, and runs each block inside. `bind_dir` exists for the case where the path AutoGen sees is not the path the Docker daemon should mount — for example when your app itself runs in a container. Unlike the local one it owns a resource, so it has a lifecycle. The idiomatic form is `async with DockerCommandLineCodeExecutor(...) as executor:`; otherwise call `await executor.start()` before use and `await executor.stop()` after. Forgetting this is the single most common beginner error and shows up as an error about the container not running. `auto_remove` and `stop_container` default to true, so the container is stopped and deleted when the executor stops — good hygiene, but it also means anything not written into `work_dir` is gone. What you get is process and filesystem isolation from the host, plus a reproducible interpreter: the model's code sees the packages in your image, not whatever your laptop happens to have. What you do not get from the constructor is a hardened sandbox — the container shares the host kernel, and egress rules, capability dropping and resource caps are set at the platform layer, not by executor arguments. Treat it as "blast radius reduced to a container I am willing to lose", not "safe to run anything". ## Choosing Use Local when you are the only source of the prompts, the machine is disposable, and startup latency matters. Use Docker the moment untrusted input can influence what gets generated — any user-facing product — and pin a purpose-built image so the code has its dependencies without needing to install anything at run time. Because the executor is injected, the usual shape is one factory that returns Local in development and Docker everywhere else; the agent code does not change. ## Related surface A third option, `JupyterCodeExecutor`, keeps a live kernel so state persists between blocks, and a Docker-backed Jupyter variant combines that with container isolation. Whichever you pick, `timeout` (default 60 seconds) bounds a single execution — not the number of executions an agent loop can request.

  • What happens if you pass a DockerCommandLineCodeExecutor to an agent without starting it?
    Execution fails: the executor has no running container to exec into, and you get an error rather than a code result. Either wrap it in `async with`, which starts and stops it around the block, or call `await executor.start()` yourself and `await executor.stop()` in a finally. Because `auto_remove` and `stop_container` default to true, the container is also destroyed on stop, so anything outside the mounted `work_dir` does not survive.
  • Does switching from the local executor to the Docker one require changing the agent code?
    No — that is the point of the executor abstraction. `CodeExecutorAgent` takes any object satisfying the executor interface through `code_executor`, so the swap is one construction site. The practical differences are lifecycle (the Docker executor must be started and stopped) and image contents: the generated code now sees the packages baked into your image rather than your host interpreter's, so pin an image that has what your agents typically need.
  • If the generated code writes a CSV file, where does it end up in each case?
    With the local executor it lands in `work_dir` on your filesystem directly. With the Docker executor the code writes inside the container, but `work_dir` is bind-mounted in, so a file written to the container's working directory appears in the host `work_dir` too and survives the container being removed. Anything written elsewhere in the container is lost when the container is stopped and auto-removed.

The local executor is handing a stranger the keys to your desk; the Docker one gives them a rented booth you can burn down afterwards.

saying these in an interview costs you the question

  • Assuming the local executor sandboxes code in some way
  • Believing generated code is safe because a model wrote it
  • Using the Docker executor without start/stop or async with
  • Claiming the container blocks all network access by default
  • Treating work_dir as an isolation boundary rather than a shared folder

context