How would you run AutoGen code execution safely when untrusted users drive the prompts?
answer
- local executor is disqualified immediately
- bake the image, do not install at run time
- one working directory per session
- timeout bounds a call, not a loop
- kernel is shared; platform sets egress and caps
basics
~20 sNever the local executor. Give each session its own container-backed executor with a purpose-built image, its own working directory, a tight per-execution timeout, and a bounded retry budget — then accept that the container is a blast-radius control, not a security boundary.
solid answer
~50 sStart from the honest threat model: user text reaches a model, the model emits code, and something runs it. `LocalCommandLineCodeExecutor` is disqualified immediately — it inherits your user, filesystem, environment variables and network position, so a single injected prompt exfiltrates every API key you exported. Use `DockerCommandLineCodeExecutor` with an image you build, containing the dependencies your agents need, so nothing has to be installed at run time. Give every session a fresh `work_dir` — never a shared or repository directory, never one holding another tenant's files — and keep `auto_remove` and `stop_container` on so containers do not accumulate. `timeout` bounds one execution, not the loop, so cap retries and total executions separately or a model in a failure loop bills you indefinitely. Add an approval gate for high-impact operations. Then be explicit about what remains: the container shares the host kernel, so network egress, capability dropping and resource caps belong to the platform you deploy on, not to the executor's constructor.
code
python · 30 linesimport asyncio
import tempfile
from autogen_agentchat.agents import CodeExecutorAgent
from autogen_agentchat.messages import TextMessage
from autogen_core import CancellationToken
from autogen_ext.code_executors.docker import DockerCommandLineCodeExecutor
async def run_session(user_code: str) -> str:
with tempfile.TemporaryDirectory() as session_dir:
async with DockerCommandLineCodeExecutor(
image="myorg/agent-sandbox:pinned",
work_dir=session_dir,
timeout=30,
auto_remove=True,
stop_container=True,
) as executor:
agent = CodeExecutorAgent(
"executor",
code_executor=executor,
sources=["coder"],
max_retries_on_error=2,
)
message = TextMessage(content=user_code, source="coder")
response = await agent.on_messages([message], CancellationToken())
return response.chat_message.content
print(asyncio.run(run_session("```python\nprint('bounded')\n```")))go deeper
Know the headline rule: never point a user-facing agent at the local executor, because generated code would run with your service's files, environment variables and network access.
Explain the concrete configuration — container-backed executor, a pinned purpose-built image, a per-session work_dir, a short timeout, containers stopped and removed after use — and why each argument matters.
Show that you bound the loop as well as the call, keep the mounted directory as the only deliberate hole, add an approval gate for externally visible actions, and log every execution because the container is destroyed by design.
Own the boundary the framework does not draw: shared kernel, no egress or capability control from the constructor, so this workload belongs on infrastructure you would be willing to treat as compromised — and be able to defend the isolation-versus-latency trade when someone proposes pooling containers.
## Frame the decision correctly The question is not "is AutoGen's code execution secure". It is "what am I willing to lose when a model, steered by an untrusted user, runs code I never reviewed". Everything else follows from that answer, and a strong response walks the layers rather than naming one control. ## Layer 1: eliminate the local executor `LocalCommandLineCodeExecutor` is a development tool. In a user-facing product it means arbitrary generated code runs as your service account, reading your service's environment — which in practice is where your model provider keys, database credentials and cloud tokens live — with your service's network reach, which typically includes internal endpoints no external attacker can otherwise touch. Prompt injection turns a helpful research agent into an internal-network scanner. Its useful siblings do not help here either: `virtual_env_context` isolates Python dependencies, not the process. The rule is simple and worth stating plainly in an interview: local executor for prototypes and disposable CI runners only. ## Layer 2: the container, and what it actually buys `DockerCommandLineCodeExecutor` moves execution into a container: separate filesystem view, separate process namespace, and a controlled interpreter. Two configuration decisions carry most of the value. **Bake the image.** Do not use a bare base image and let the model install packages at run time. Build one containing the libraries your agents realistically need, pinned. Two benefits: executions are fast and deterministic, and there is no legitimate reason for the container to reach a package registry, which makes an egress deny rule at the platform layer viable. **Own the working directory.** `work_dir` is bind-mounted, so it is the deliberate hole in the container's filesystem. One directory per session, created fresh, removed after, never pointing at a repository checkout, never containing another user's outputs, never containing credentials. `extra_volumes` should be treated the same way — every mount is another hole. `bind_dir` exists for the case where your own app runs in a container and the daemon needs a different host path; getting it wrong usually surfaces as an empty working directory rather than a leak, but it is worth knowing why the argument exists. ## Layer 3: bound the loop, not just the call `timeout` (default 60 seconds) kills one execution. It says nothing about how many executions the agent may request. Between `max_retries_on_error` on a model-backed executor agent and the surrounding conversation's own stopping rules, an agent that keeps failing keeps generating and keeps running. Cost and latency controls therefore need three separate bounds: per-execution timeout, retry budget, and a wall-clock or turn cap on the session enforced by your application. Interviewers probe this because "the sandbox held" and "we spent four thousand dollars overnight" are both true outcomes of the same unbounded loop. ## Layer 4: the approval gate For operations whose consequences are external — sending mail, writing to a bucket, calling a paid API — put a human decision on the execution path via the executor agent's approval callback, and auto-approve pure computation so the human's attention is spent where it matters. Any code-inspecting heuristic you use to decide what needs a human is advisory: it reads model-authored source text and can be evaded, so it prioritises attention rather than enforcing a boundary. ## Layer 5: what the framework does not give you The part that separates a principal answer from a competent one is naming the limits. A container shares the host kernel. Nothing in the executor's constructor drops capabilities, disables egress, caps memory or CPU, applies seccomp, or runs the workload on a hypervisor boundary. Those are decisions about the machine and the orchestrator you deploy on. So the design conclusion is: run agent code execution on infrastructure you are prepared to treat as compromised — a dedicated node pool or sandbox service, no ambient cloud credentials attached, no route to internal services, egress default-deny with a narrow allowlist, and short-lived instances. ## Layer 6: lifecycle and tenancy Containers are per-session, not per-process-lifetime. Reusing one executor across users leaks files and any state the previous session left in the mounted directory. `restart()` and stop/start exist for exactly these boundaries. Decide who pays for the cold-start cost of a fresh container per session, and if pooling is unavoidable, pool per tenant and wipe the mounted directory between uses — and say out loud that pooling trades isolation for latency rather than pretending it is free. ## Observability Record every execution: the code, the exit code, output size, duration, the approval decision, and the session it belonged to. When something goes wrong the transcript is the only forensic artifact, because the container itself is gone by design. Alert on the shapes that matter — timeout rate, repeated nonzero exits, executions attempting network calls — since they are early evidence of both a misbehaving agent and a probing user.
- Why is baking dependencies into the image better than letting the agent pip install what it needs?Three reasons. Executions become fast and deterministic instead of paying an install per run and drifting with the registry. Failures stop being confusing — a missing package is a config bug you fix once rather than something the model works around. And most importantly, it removes the only legitimate reason for the container to reach the network, which makes a default-deny egress rule at the platform layer actually deployable.
- You set timeout to 30 seconds and the bill is still enormous. What did you miss?Timeout bounds a single execution. A model-backed executor agent with a retry budget, sitting in a conversation that has not stopped, will keep generating and keep running — each attempt costing a model call plus up to 30 seconds of compute. You need three bounds, not one: per-execution timeout, a retry budget on the agent, and a session-level cap on turns or wall clock enforced by your application.
- Is reusing one container across user sessions ever acceptable?Only with eyes open. It buys latency by skipping container start, and it costs isolation: the mounted working directory carries one user's files into another's session, and anything the previous execution left behind is still there. If you must pool, pool per tenant, wipe the mounted directory between sessions, and restart the executor on a fixed cadence. Present it as a deliberate latency-for-isolation trade, not a default.
- What would you tell a team that says the Docker executor makes generated code safe?That it bounds the blast radius rather than eliminating it. The container shares the host kernel, and nothing in the executor's arguments drops capabilities, caps memory or CPU, or blocks egress — those live in the orchestrator and the host. The right posture is to run this workload on infrastructure you would be willing to treat as compromised: no ambient cloud credentials, no route to internal services, default-deny egress, short-lived nodes.
saying these in an interview costs you the question
- Calling the Docker executor a security boundary rather than a blast-radius control
- Sharing one work_dir or one container across users
- Relying on timeout alone to bound cost
- Letting generated code install packages from the network at run time
- Running the executor on a host that holds cloud credentials or internal network access
- Assuming an approval gate removes the need for isolation