A tool returns a 40MB log dump — how do you get it into the agent's context?
answer
- do not inline the payload
- pass a pointer, not the dump
- handle plus a short digest
- the agent must be able to re-read it
- code-generated digests, outliers not averages
basics
~20 sYou do not. Store the payload outside the window and pass a reference — a file path, object key or result ID — plus a short digest such as size, shape and the first lines. The agent reopens or searches it on demand if the digest is not enough.
solid answer
~50 sInlining a 40MB payload is impossible and, at any size that would fit, usually still wrong: it burns the budget on tokens that were relevant for one decision. The pattern is **offloading by reference**. The harness writes the payload to a store, and the context builder puts in a compact record: a handle (path, object key or result ID), the payload's size and shape, and a small digest — the first lines, a schema, counts by category, whatever lets the model judge relevance. Offloading only works if there is a way back in. Pair the handle with tools that read a span, grep, or run code over the stored artefact, so the agent can pull the exact 200 tokens it needs instead of the 40MB it does not. The failure mode to guard against is a digest that hides the anomaly the agent was looking for, so digests should be generated by deterministic code and should surface outliers rather than averages.
code
python · 17 linesimport json, hashlib, pathlib
def offload(payload: str, store: pathlib.Path, name: str) -> dict:
path = store / name
path.write_text(payload)
lines = payload.splitlines()
return {
"handle": str(path),
"bytes": len(payload),
"lines": len(lines),
"head": lines[:12],
"sha256": hashlib.sha256(payload.encode()).hexdigest()[:12],
}
dump = "\n".join(f"event {i} reason=BackOff" for i in range(100000))
record = offload(dump, pathlib.Path("/tmp"), "events.log")
print(json.dumps({k: v for k, v in record.items() if k != "head"}))go deeper
Know that large tool outputs are not pasted into the prompt. The application stores them and passes a short summary plus a reference the agent can look up if it needs more.
Explain the record you actually put in context — handle, size and shape, digest, provenance — and stress that offloading requires a reader or search tool, otherwise it is just information loss.
Show judgement on digest design: deterministic generation, surfacing outliers rather than averages, honest sampling labels, and handle lifecycle across compaction and across tenants.
Frame the context window as the top of a memory hierarchy and set the policy: default thresholds for offloading, ownership and retention of the artefact store, tenant isolation on handles, and metrics on how often handles are dereferenced.
## The problem Tools in a real system return payloads that are wildly out of proportion to the decision they inform. A cluster event dump, a full table export, a build log, a large API response — tens of megabytes, of which perhaps a hundred tokens actually matter. If the context builder inlines whatever the tool returned, one call can consume the entire budget for the session and trigger rot for everything that follows. ## Offloading by reference The harness treats large results the way an operating system treats large files: it stores them, and passes a handle. What goes into the window is a small structured record: - **the handle** — a file path in the agent's workspace, an object-store key, or an internal result ID - **shape metadata** — byte size, line or row count, columns or fields present, time range covered - **a digest** — the first N lines, a schema sample, counts grouped by a meaningful key, or the top anomalies - **provenance** — which tool produced it, with which arguments, at what time For a 40MB Kubernetes event dump that might be a path plus twelve lines: total events, distinct reasons with counts, the time span, and the three most frequent error reasons. Roughly 150 tokens standing in for 40MB. ## The re-entry path is mandatory A handle with no way to dereference it is worse than useless — the model will either hallucinate the contents or give up. Offloading is only complete when the agent has tools to go back in: read a byte or line range, search by pattern, filter by field, or run code over the artefact in a sandbox. The last of these is the strongest, because the model can write a filter that returns exactly the rows it needs without any of them passing through the window on the way. This is what makes offloading fundamentally different from truncation. Truncation throws information away permanently. Offloading moves it one level down the hierarchy and leaves a path back. ## Designing the digest The digest is where this pattern succeeds or fails, and it deserves real thought. **Generate it with code, not a model, where you can.** Counts, distinct values, ranges, first and last records are deterministic, free, and cannot hallucinate. Reserve model-generated summaries for genuinely unstructured prose. **Surface outliers, not averages.** The agent is usually hunting for something anomalous. A digest reporting mean latency is nearly useless; one reporting the slowest five requests is actionable. A digest that smooths away the single interesting line is the most common way this pattern quietly fails — the agent reads it, concludes nothing is wrong, and never reopens the handle. **Make the shape metadata honest.** If the digest covers the first 200 of 900,000 lines, say so explicitly. Otherwise the model treats a sample as the whole. ## When to inline anyway Offloading is not free — it costs a store, a lifecycle, and at least one extra round trip when the agent goes back in. Inline directly when the payload is small relative to the budget and will be referenced repeatedly across subsequent turns; the round trip would cost more than the tokens. Offload when the payload is large, single-use, or when only a small unpredictable slice will matter. A rough operational rule many teams use: anything over a few thousand tokens gets a handle by default. ## Lifecycle concerns Handles are state, and state has a lifetime. Three things to get right: **Expiry.** If the store is cleaned up mid-session, an agent that reopens a handle gets an error that reads to it like a tool failure. Handle lifetimes should outlive the session, or expiry should be reported clearly enough for the agent to recover. **Compaction interaction.** Handles must survive compaction. They are cheap to carry and they are precisely what allows a summary to be lossy safely — anything dropped from the summary is still reachable through the reference. Losing handles during compaction turns a recoverable omission into a permanent one. **Isolation.** Handles are capabilities. A handle from one user's session must not be readable from another's, and the reader tool must enforce that in code rather than trusting the model not to guess an identifier. ## The general principle Offloading generalises past tool results. The same reasoning applies to long documents, prior session transcripts and large intermediate artefacts: keep pointers and digests in the window, keep the bulk outside, and give the model instruments to inspect what it points at. The context window is the fastest, smallest and most expensive tier of the memory hierarchy, and it should hold the working set, not the archive.
- What goes wrong if the agent never reopens the handle you gave it?It reasons from the digest alone and answers confidently on partial evidence — the most dangerous failure of this pattern, because it looks like success. Mitigate by making the digest state explicitly that it is a sample of a larger artefact, by naming the reader tool in the same record, and by measuring how often handles are dereferenced. A handle that is never opened means the digest is either sufficient or misleading, and you need to know which.
- How do you decide between digest-only, handle-plus-digest, and full inline?By size relative to budget and by expected reuse. Small and repeatedly referenced: inline, since round trips cost more than tokens. Large but only a narrow slice matters: handle plus digest with a reader tool. Large and genuinely only the aggregate matters: digest alone is fine, but keep the handle anyway — it is a few tokens and it is your recovery path when the aggregate turns out to be wrong.
- Why prefer a deterministic digest over asking a model to summarize the payload?Cost, latency and fidelity. Counts, ranges and first/last records are exact, free and reproducible, while a model summary costs a call and can omit or invent detail. It also makes the digest testable — you can assert that an injected anomaly appears in the digest. Reserve model summarization for unstructured prose where no structural digest exists.
saying these in an interview costs you the question
- Pastes the whole payload into the context window
- Gives a handle with no tool that can read it
- Confuses offloading with truncating the result
- Uses averages in the digest and hides the anomaly
- Assumes handles remain valid forever without lifecycle