Which CrewAI agents should hold CodeInterpreterTool, and how do you contain what it runs?
answer
- container by default, host on request
- the flag is not the control
- never with untrusted text in the same step
- egress and credentials inside the box
- a typed tool is often the better answer
basics
~20 sGive CodeInterpreterTool to one narrowly scoped agent, never to agents that also ingest untrusted text. It executes generated Python inside a Docker container by default; unsafe_mode=True runs it in the host process, which is a development-only setting.
solid answer
~60 s`CodeInterpreterTool` from `crewai_tools` executes model-generated Python. By default it runs that code inside a Docker container it manages, which is why the tool needs a working Docker daemon; `unsafe_mode=True` drops that and executes in the host process instead. You can also get it implicitly through `Agent(allow_code_execution=True)`, where `code_execution_mode` selects the containerised or direct path. The design question is not the flag but the assignment: keep execution on a single agent whose task list is narrow, and make sure no task gives that agent both code execution and a tool that pulls in untrusted external text — a scraped page or an unaudited tool result in context is an instruction channel, and code execution is the most valuable thing it can reach. Because tools can be set per task, you can hold execution on the agent and still deny it for steps that ingest content. Beyond the container, think about what the container itself can reach: network egress, mounted paths, and how long a runaway loop is allowed to burn.
code
python · 16 linesfrom crewai import Agent
from crewai_tools import CodeInterpreterTool, ScrapeWebsiteTool
gatherer = Agent(
role="Gatherer",
goal="Collect source material for the analysis",
backstory="Fetches pages and records what they say.",
tools=[ScrapeWebsiteTool()],
)
analyst = Agent(
role="Analyst",
goal="Compute statistics over material already gathered",
backstory="Writes small, verifiable Python.",
tools=[CodeInterpreterTool()],
)go deeper
Know that CodeInterpreterTool runs model-written Python, that it uses Docker by default, and that you attach it deliberately to one agent rather than to everything.
Explain the two routes — the tool directly, or Agent(allow_code_execution=True) with code_execution_mode — and what unsafe_mode=True gives up by executing in the host process.
Argue the containment beyond the flag: network egress from the sandbox, credentials in its environment, mounted paths, CPU and wall-clock ceilings, and bounding the retry loop when generated code fails on a missing import.
Own the policy: which agent may execute at all, an explicit rule that no task holds ingestion and execution together, the audit trail of executed source, and the standing challenge of whether a typed tool would have removed the need for general execution entirely.
## What the tool actually does `CodeInterpreterTool` takes code the LLM wrote and runs it. By default it executes inside a Docker container the tool manages, so the host filesystem and process space are not directly exposed; the practical requirement is that a Docker daemon must be available wherever the crew runs. Setting `unsafe_mode=True` skips the container and executes in the host Python process — convenient on a laptop, indefensible in anything shared. CrewAI also offers the implicit route: `Agent(allow_code_execution=True)` attaches execution to the agent, and `code_execution_mode` chooses between the sandboxed and the direct path. Same capability, less visible in a diff — which is itself a review concern, because "this agent can run arbitrary code" should be obvious to a reader. ## Assignment is the real control A sandbox limits what executed code can touch. It does nothing about *what code gets written*, and the code is written by a model whose context you do not fully control. So the first-order control is the tool list. The rule I would hold in a design review: **no task's tool list may contain both an ingestion tool and an execution tool.** Ingestion means anything that pulls text you did not author — `ScrapeWebsiteTool`, a search tool's snippets, a RAG tool over user-supplied documents, results from a third-party MCP server. All of that lands in the agent's context as text, and text in context is indistinguishable from instructions to a language model. If the same step can execute code, a successful injection has a payload. Because CrewAI resolves tools at two levels — the agent's standing list and the task's per-step list, with the task's taking precedence — you have a clean way to enforce this without duplicating agents: one analyst agent may hold the interpreter, but any task where it reads external material overrides `tools` to a list without it. Better still, split the roles: a gatherer agent that has ingestion tools and no execution, and an analyst agent that has execution and works only from what the earlier task produced. ## Contain the sandbox, not just the code The container is a boundary, not a vacuum. Questions worth asking before this ships: - **Network egress.** Can generated code reach your internal network, your metadata service, your database? A container with default networking often can. Egress restriction is usually the highest-value control after using a container at all. - **Credentials.** Anything in the container's environment is readable by generated code. Do not pass the crew's API keys into the execution environment. - **Mounts and data.** If you mount a working directory to move files in and out, that path is writable by generated code. Scope it to one run. - **Resource ceilings.** Model-generated code has an ordinary rate of infinite loops and accidental gigabyte allocations. CPU, memory and wall-clock limits belong on the container. - **Lifetime.** Treat the execution environment as disposable per run so state does not accumulate between crews. ## Cost and reliability Execution is slow relative to a model call once container startup is counted, and it fails in ways the model handles badly — a missing package, a syntax error, an import that is not installed in the image. Those failures come back as tool observations and the agent will try again, spending model turns. Bound that with the agent's iteration limit, and consider baking the libraries your workload actually needs into the image rather than letting the agent discover their absence one round trip at a time. ## When not to use it Code execution is often reached for when a deterministic tool would do. If the crew's real need is "compute a statistic over this CSV" or "reformat these dates", a purpose-built `BaseTool` with a typed `args_schema` is faster, cheaper, testable, auditable and unable to do anything else. Reserve the interpreter for genuinely open-ended analysis where you cannot enumerate the operations in advance. A principal-level answer names that tradeoff explicitly: the interpreter buys generality and costs you the ability to reason about what the system can do. ## Operating it Log the code that was executed, not just its output — it is your only record of what the system did, and it is the artefact you will want during an incident. Track failure classes (missing import, timeout, exception) separately from successes, because a rising missing-import rate is a signal that the image and the workload have drifted apart. ## The shape of a good answer Start with the mechanics — container by default, `unsafe_mode` as the escape hatch, `allow_code_execution` as the implicit route. Then make the judgment call the centre of the answer: least privilege per task, never co-locating ingestion with execution, egress and credentials inside the sandbox, and the standing question of whether a typed tool would have been the right answer instead.
- Why is 'it runs in a container' an incomplete answer to the safety question?A container bounds what executed code touches, but it does not decide what code gets written — that comes from a model whose context may include text you did not author. And the container itself usually has network access and whatever environment variables you passed it. The controls that matter are keeping ingestion and execution out of the same task's tool list, restricting egress, and keeping credentials out of the execution environment.
- When would you replace CodeInterpreterTool with a custom BaseTool?Whenever you can enumerate the operations in advance. A typed tool with an `args_schema` for 'compute these statistics over this CSV' is faster, cheaper, unit-testable, auditable, and structurally incapable of doing anything else. Keep the interpreter for open-ended analysis where the operations genuinely cannot be listed up front — that generality is exactly what you are paying for.
- What operational signals would you track on a crew that executes generated code?Log the executed source, not only its output — it is the record you will want in an incident. Then track failure classes separately: missing imports, timeouts, exceptions, resource kills. A rising missing-import rate means the container image and the workload have drifted apart, which is fixed by baking the needed libraries in rather than letting the agent rediscover their absence one model round trip at a time.
saying these in an interview costs you the question
- Treats the default container as sufficient without restricting egress
- Sets unsafe_mode=True in a deployed environment for convenience
- Puts a scraping tool and the interpreter on the same task
- Passes the crew's API keys into the execution environment
- Reaches for code execution where a typed tool would suffice