In a multi-agent system, why treat the context window as RAM and shared storage as disk?
answer
- one tier is fast and small
- pointers travel, payloads do not
- pull a slice, do not push everything
- cost multiplies once per agent
- state size stops being window-bounded
basics
~20 sThe context window is small, volatile working memory; shared storage is large and durable. Agents keep references — paths, ids, ranges — in context and pull only the bytes a step needs, so total state size stops being limited by the window.
solid answer
~50 sThe analogy sets expectations correctly. A context window behaves like RAM: fast, addressed directly by the model, small, and wiped between runs. A shared artifact store behaves like disk: slower to reach, cheap, effectively unbounded, and durable. The design consequence is that shared state is held **by reference**. Six agents working a construction project can coordinate over two hundred megabytes of drawings without a single byte of those drawings ever being resident in any context — each agent carries paths and a short index, and calls a read tool to fetch the one sheet or section it needs for the current step. Broadcasting the full state into every agent's prompt is the opposite move: it multiplies token cost by the number of agents, degrades answer quality as windows fill, and still fails the moment the state exceeds one window.
go deeper
Be able to say that the window is small and temporary while the store is large and durable, and that agents pass file paths and read content only when a step needs it.
Explain the cost and quality consequences of broadcasting state to every agent, and how indexes, small files and summaries make on-demand reads targeted.
Demonstrate that you instrument this — bytes loaded versus bytes used, repeated reloads, prompt growth without success growth — and that you restructure artifacts when reads are wasteful.
Own where the boundary sits as a platform decision: artifact granularity and indexing standards, read-cost budgets per run, and the fact that window growth changes the constants but never removes the tier.
## The two tiers Every agent has exactly one directly addressable memory: its context window. Whatever is in the window is available to the model at full fidelity, immediately, with no tool call. That is its RAM. It is also small relative to real workloads, and it is gone when the run ends. Everything else — files, buckets, tables, an append-only log — is the disk tier. Reaching it costs a tool call and a round trip. In exchange it is durable, shared across agents, and effectively unbounded. The analogy is worth taking seriously because it predicts the right design moves. You do not load a whole database into RAM to read one row; you index and seek. The same discipline applies to agents. ## Holding state by reference What lives in an agent's window should be pointers and a small working set: the task statement, the paths or ids of the artifacts in play, a short index or manifest, and the specific content the current step operates on. A construction crew of six agents may be coordinating over a two-hundred-megabyte drawing set. None of it is resident anywhere. The schedule agent holds `schedule.json`'s path and the ids of three open change requests; the field agent holds today's log path. When an agent needs sheet A-201, it reads sheet A-201 — and, ideally, drops it again once the step is done. This is what "load on demand" means concretely: the store is not pushed to agents, it is pulled from, in slices, by the agent that has a reason. ## Why broadcasting is the failure mode The naive alternative is to give every agent the whole shared state at the start of every turn so nobody is missing anything. It fails on three axes at once. - **Cost.** The same tokens are paid for once per agent per turn. With six agents that is a sixfold multiplier on the largest part of the prompt, for content most of them will not use. - **Quality.** Long windows degrade attention to any individual detail — the reason an agent misses the one line that mattered is often that it was surrounded by fifty thousand irrelevant ones. - **Ceiling.** It simply stops working when shared state exceeds one window, which for real document sets is immediately. The RAM/disk framing makes all three predictable rather than surprising. ## What the analogy does not say It is a framing device, not a specification. Three places it breaks down are worth knowing: - Reads are not free or exact. A model choosing what to load is an inference step, and it can choose wrong — retrieving the wrong file is a failure mode with no hardware analogue. - There is no automatic write-back. Nothing flushes an agent's working memory to the store; a result exists in shared state only if the agent deliberately wrote it there, and an agent that finishes without writing has produced nothing durable. - The window is not addressable by the agent the way RAM is by a program. The agent cannot free a specific region at will; the surrounding harness manages what stays. ## Practical consequences Design the store so on-demand reads are cheap and targeted: stable paths, small files rather than one giant document, an index or manifest an agent can scan before deciding what to open, and summaries at the top of long artifacts. A store that can only be read in one-hundred-thousand-token gulps is a store nobody can use selectively. And instrument reads. If an agent reliably loads five artifacts to use one, the index is wrong, not the agent. ## In an interview State the two tiers, then give the concrete consequence: agents pass paths, not payloads. The follow-up is usually about what to do when a needed artifact is genuinely too large for the window — the honest answer is that it must be split, summarized at write time, or read in ranges, because there is no version of the design where the window grows to meet it.
- What do you do when a single artifact an agent genuinely needs is larger than its context window?You change the artifact, not the plan. Split it into addressable sections, add a summary or table of contents the agent can read first, or expose a range read so the agent pulls only the relevant span. The one option that never works is hoping the window accommodates it — that turns a design problem into a runtime failure.
- How would you tell that agents are loading too much shared state?Instrument reads per step and compare bytes loaded against bytes actually cited in the output. A high ratio, or the same artifact reloaded repeatedly within one run, points at a missing index or too-coarse file granularity. Rising prompt token counts with flat task success is the same signal from the billing side.
- Does the RAM analogy imply anything is automatically written back to the store?No, and that is the sharpest place it breaks. There is no flush. A result exists durably only if the agent explicitly wrote it to the artifact store or logged it; otherwise it disappears with the run. Designs that assume implicit persistence lose work silently whenever an agent finishes without a write step.
A program does not load an entire database into RAM to read one row; it keeps an index and seeks. Agents keep paths and read only the artifact the current step needs.
saying these in an interview costs you the question
- Assumes every agent should be given the full shared state each turn
- Thinks a bigger context window removes the need for a store
- Believes work in an agent's context is persisted automatically
- Passes whole file contents between agents instead of paths
- Treats retrieval of the right artifact as a free, always-correct operation