An incident agent's log-search tool returns 40 MB of JSON. Where do you reduce it?
answer
- reduce before it becomes context
- results are re-sent every turn
- write time, sandbox, or history
- return handles, not payloads
- clearing history is damage control
basics
~20 sReduce it before it becomes context. Aggregate inside the tool at write time, or let the agent run the query as code in a sandbox and return only the derived findings. Trimming history afterwards is damage control, not the fix.
solid answer
~50 sThere are three places to shrink a result and they are not equivalent. **At write time**, the tool wrapper aggregates before serializing — for a six-hour log window, return the 12 distinct stack-trace fingerprints with counts and first-seen times, not the raw events. **In a code-execution sandbox**, the agent writes the filter itself: the 40 MB lands in the sandbox's memory, the code reduces it, and only the derived summary crosses into context — which is what you want when you cannot predict the reduction in advance. **In history**, older tool results are cleared or replaced once their signal has been extracted; this recovers a window that is already full but cannot un-charge the tokens you already paid, and it destroys anything you failed to extract first. Prefer the first two. Also return handles — the query id or file path — so the raw data stays fetchable without being resident.
go deeper
Know that a tool result stays in the conversation and is re-sent every turn, so returning a huge payload is expensive long after the call that produced it.
Explain the three reduction points — write time, sandbox, history — and why the first two prevent cost while the third only recovers space that was already paid for.
Show the design judgment: parameterized aggregation for predictable projections, sandbox filtering for open-ended investigation, handles so the raw data stays reachable, and explicit truncation markers as a backstop.
Own per-tool result budgets and the platform decision about whether to run a code-execution sandbox at all — its security boundary, latency floor and operational cost weighed against the context savings across the whole agent fleet.
## Why 40 MB is not a formatting problem A raw six-hour log dump is roughly ten million tokens; the window it must fit in is a fraction of that. But even a result that technically fits is expensive in a way that a single-shot call is not. Tool results persist: whatever lands in the conversation is re-sent on every subsequent turn until something removes it. A 60,000-token result in an agent loop that runs thirty more turns is charged thirty more times, adds latency to every one of them, and dilutes the model's attention across a mass of low-signal text — the failure mode usually called context rot, where recall of the material that matters drops as irrelevant volume around it grows. So the real question is never "how do I fit this?" but "what is the smallest set of facts that lets the agent make its next decision?" For an incident agent, that is almost never the events. It is the distinct failure signatures, their counts, when each first and last appeared, and which service emitted them. ## Reduction point 1: at write time, inside the tool The cheapest reduction is the one that happens before the result is ever serialized. The wrapper runs the query, aggregates in code, and returns something like twelve fingerprints with counts and timestamps — a few hundred tokens carrying nearly all the decision-relevant signal. This is exact, deterministic, free of model tokens, and testable like any other code. Its limit is that you must know the reduction in advance. A fixed aggregation is perfect when the question is always "which errors spiked", and wrong when the agent needs a field you did not think to keep. Design against that by making the aggregation parameterized — group-by, filter, limit as tool arguments — so the model can ask for a different projection instead of asking for everything. Beware of doing this reduction with a model call. Summarizing at write time with an LLM costs tokens and latency on every invocation and can drop the one line that mattered. Reserve it for genuinely unstructured payloads where code cannot do the job. ## Reduction point 2: inside a code-execution sandbox The stronger version gives the agent a code environment and lets it write the reduction itself. The agent emits a short program that calls the log query, and the 40 MB materializes inside the sandbox process, not in the transcript. The program groups by fingerprint, counts, sorts, and prints the twelve lines it decided were interesting. Only that printed output crosses back into context. This keeps the flexibility of raw access — the agent can filter on any field, join two sources, compute a rate — while paying context only for the code and its output. Reported reductions of one to two orders of magnitude on large tool payloads are routine, because the ratio of raw data to decision-relevant facts in observability and analytics work is enormous. It also composes: intermediate results stay as sandbox variables or files across steps rather than being re-narrated each turn. The costs are real. You need a sandbox with a security boundary, a runtime and a lifecycle; the agent's code can itself be buggy and needs an error path back to the model; and there is a latency floor per execution. For an agent whose whole job is data-heavy investigation, it usually pays anyway. ## Reduction point 3: editing history after the fact The last lever operates on results already in the window: drop or replace tool results older than some age or beyond some count, keeping the assistant's reasoning and the extracted conclusions. Some providers expose this automatically, clearing older tool results once the window pressure crosses a threshold; you can equally implement it yourself when you own the message list. Treat it as damage control rather than design. It cannot refund the tokens already spent on the turns where the result was resident, and it is lossy in a way you may not notice: if the agent had not yet extracted the fingerprint from a cleared result, that information is simply gone and the agent may re-run the same expensive query. The discipline that makes it safe is to have the agent write its conclusion somewhere durable — a scratchpad file, a structured note — before the source result is eligible for clearing. ## Return handles, not payloads Cutting across all three: the tool should hand back a reference alongside the summary — the query id, the object-store path, the row count — so the raw data remains reachable through a follow-up call without being resident. This is the just-in-time posture: the context holds pointers and the agent dereferences on demand. It is what lets you be aggressive about reduction without ever making the underlying data unreachable. ## Choosing Aggregate at write time when the useful projection is predictable. Give the agent a sandbox when it is not, and the payload is large and structured. Clear history when a long run has accumulated results whose value has already been extracted. Whatever you choose, cap the result: a hard byte or token limit per tool result with explicit truncation wording is the backstop that keeps one unlucky query from destroying a run.
- Why is clearing old tool results weaker than filtering in the sandbox?Clearing happens after you have already paid for the result on every turn it was resident, and it is lossy — anything the agent had not yet extracted disappears with it, often causing the same expensive query to be re-run. Sandbox filtering means the bulk never entered the transcript at all, so there is nothing to pay for and nothing to lose.
- What is the risk of aggregating hard inside the tool wrapper?You are freezing a projection chosen at design time. When the agent needs a field the aggregation dropped, it has no path to it and may conclude the data does not exist. Mitigate by parameterizing the aggregation — group-by, filters, limit as arguments — and by returning a handle to the full result so a follow-up call can retrieve more.
- Should the tool call a model to summarize the payload before returning it?Only when the payload is genuinely unstructured and code cannot reduce it. LLM summarization at write time adds latency and cost to every invocation and can silently drop the one anomalous line that mattered. Deterministic code — group, count, sample, project — is cheaper and auditable, and it is what most log, metric and query payloads need.
- What should happen if a single result still exceeds your cap after reduction?Truncate at a hard byte or token limit and say so explicitly in the result: how much was returned, how much matched, and how to narrow or page. Silent truncation is worse than the size problem, because the agent then reasons over a partial set believing it is complete.
saying these in an interview costs you the question
- Says the model will just ignore the irrelevant parts
- Treats a large window as a reason not to reduce results
- Only trims history and calls that the solution
- Summarizes every payload with an LLM by default
- Truncates silently at a byte limit with no marker