What are the limits of OpenAI's hosted code_interpreter tool container?
answer
- sandboxed Python the model drives
- nothing reaches the outside network
- short-lived scratch space
- charged per session, not per line
- download outputs or lose them
basics
~20 sThe tool runs model-written Python in a sandboxed container with no open internet access. The container is short-lived, expiring after a period of inactivity, is billed per session on top of tokens, and any files it produces are lost unless you download them first.
solid answer
~50 s`code_interpreter` gives the model a sandboxed Python environment it can write and execute code in, iterating on errors until the code runs. You attach input files to the container, and the model can emit output files — charts, converted spreadsheets — which come back as citations pointing at container files you then download through the containers endpoint. The constraints are what matter operationally: the sandbox has no general internet access, so the model cannot fetch data or install arbitrary packages at run time; containers are ephemeral and expire after roughly twenty minutes of inactivity, so nothing persists between sessions unless you re-upload it; and billing adds a **per-container-session** charge on top of token usage, which multiplies badly if every user request spins up its own container. Reuse a container across turns of one conversation, and treat it as scratch space rather than a compute platform.
code
python · 9 linesresp = client.responses.create(
model="gpt-5",
input="Plot the monthly totals from the attached CSV and save it as a PNG.",
tools=[{
"type": "code_interpreter",
"container": {"type": "auto", "file_ids": [uploaded.id]},
}],
)
print(resp.output_text) # annotations reference the generated container filego deeper
Know that this tool lets the model write and run Python in a sandbox, which is how it does real arithmetic, reads a spreadsheet, or draws a chart instead of guessing.
Explain the container: files go in, generated files come back as references you must download, execution is isolated with no internet, and the environment disappears after inactivity.
Show cost and reliability judgment — reuse a container per conversation rather than per message, persist outputs before expiry, and read the emitted code when a number looks wrong.
Draw the boundary between model-driven scratch execution and real compute: pinned dependencies, auditability and long jobs belong in your own infrastructure with the model only proposing code.
## What the tool is The hosted code interpreter is a sandboxed execution environment the model drives itself. Given a task like "compute the median deal size and plot the distribution", the model writes Python, runs it, reads the traceback if it fails, and rewrites — a loop you neither prompt for nor see the individual steps of, though they surface in the response's tool-call items. Its value is deterministic computation: arithmetic, data wrangling, format conversion and chart rendering are things a language model is bad at doing in its head and good at doing in code. ## Containers Execution happens inside a container. In the Responses API you request one implicitly with `{"type": "code_interpreter", "container": {"type": "auto"}}`, which creates a container on demand, or you create a container explicitly and pass its id to reuse it across several requests. Input files are attached to the container so the model can open them by path. In the Assistants API the equivalent attachment is the assistant's or thread's code-interpreter file resources. Containers are **ephemeral**. They expire after a period of inactivity — around twenty minutes — and once expired, everything inside is gone: uploaded inputs, intermediate state, and generated outputs. There is no snapshot, no volume, no restart with prior state. ## No open network The sandbox does not have general internet access. The model cannot fetch a URL, call your API, or install a package that is not already available in the image. This is a security property, not an oversight, and it shapes design: anything the code needs must be uploaded as a file or passed inline in the prompt. If the task genuinely requires a network call, that call belongs in one of your own functions, with the result handed to the model — the built-in interpreter is not the place for it. ## Getting files back out When the code writes a file, the response's annotations reference it as a container file. To deliver it to a user you fetch the file's content through the containers file API and store it somewhere durable — your own object storage, typically — before the container expires. Applications that render a chart by linking directly at the container path work in a demo and break the next day. ## Cost shape Billing is per **container session**, added to the token cost of the conversation. That unit is worth internalising because it changes architecture: a naive implementation that creates a fresh container for each user message pays the session charge on every turn, while one that creates a container per conversation and reuses its id pays once for the whole exchange. Reuse also preserves the uploaded dataset and any intermediate results, so subsequent turns are faster and cheaper in tokens as well. ## Failure modes to expect - **Missing package.** The model tries to import something the image lacks and, without network access, cannot install it. Symptom is a run that burns tokens looping on an ImportError. Constrain the task in your instructions to what the environment provides. - **Long computations.** The sandbox is meant for interactive-scale work; a job that takes minutes is a poor fit and risks running against limits mid-way. - **Silent misinterpretation.** The model may parse a CSV with the wrong delimiter and produce a confidently wrong number. Inspecting the emitted code — the tool-call items in the response — is the only way to catch it, which is why exposing that trace in internal tooling pays for itself. - **Large inputs.** File attachment has size limits, and a multi-gigabyte extract cannot simply be handed over; sample or pre-aggregate outside the sandbox first. ## When to use it It is an excellent fit for ad-hoc analysis over user-uploaded files, chart generation inside a chat product, format conversion, and any place where you want arithmetic to be right rather than plausible. It is a poor fit as a general compute backend, as an ETL runner, or anywhere you need reproducibility, pinned dependencies, or auditability of the executed code before it runs — those wants point at running the code yourself in infrastructure you control, with the model only proposing it.
- The model keeps failing to install a library it wants inside the code interpreter. What is going on?The sandbox has no general internet access, so package installation from a public index cannot succeed. The model will often loop on the ImportError, burning tokens. Fix it upstream: state in your instructions which libraries are available and ask for a solution using only those, or move the work into your own function tool where you control the runtime and its dependencies.
- How do you avoid paying a container session charge on every message of a conversation?Create a container explicitly, keep its id alongside the conversation, and pass the same id on subsequent requests instead of letting each call auto-create one. You pay for the session once, and the uploaded dataset plus intermediate results stay warm, so later turns are cheaper in tokens too. Watch the inactivity expiry — after a long user pause the container is gone and the next request needs a fresh one.
- A chart the interpreter generated is broken in your UI the following day. Why?You almost certainly linked at the container file rather than copying it. Containers are ephemeral and expire after inactivity, taking their filesystem with them, so any artefact you intend to show later must be fetched through the container files API and written to durable storage of your own while the container is still alive.
saying these in an interview costs you the question
- Expecting the sandbox to fetch URLs or call your API
- Treating container files as durable storage for generated charts
- Creating a fresh container for every user message
- Assuming any Python package can be installed on demand
- Using it as a general compute backend for long-running jobs