skip to content

How do you truncate a 40k-token ReAct observation without losing the answer?

level: middleimportance: must knowfreq 62%

answer

  1. shrink by relevance, not by position
  2. the middle is where answers hide
  3. head+tail, top-k, structured extraction
  4. always mark what was dropped
  5. keep a handle to fetch more

basics

~20 s

Reduce by relevance, not by position. Extract the fields or rows the current step actually needs, and leave an explicit marker saying how much was dropped, so the model treats the observation as partial rather than complete.

solid answer

~50 s

There are three reduction families and each loses something different. **Head+tail truncation** is cheap and deterministic, but it cuts the middle, which on a fetched HTML page is usually exactly where the answer lives. **Selection** — top-k matching rows, the matching section, a scoped slice — keeps relevant content but needs a query to score against, and silently drops whatever that query did not anticipate. **Structured extraction**, where a cheap model or a parser turns the payload into a few named fields, gives the best fidelity per token but costs an extra call and returns only the fields you thought to ask for. To get a 40k-token page down to an 800-token observation I extract first, fall back to selection, and use head+tail only as a last-resort guard. Whichever I use, the observation carries a marker such as `[truncated: 38k of 40k tokens dropped]` plus a handle — a URL, cursor, or file path — so the agent can fetch more instead of guessing.

code

python · 6 lines
python
def head_tail(text: str, budget_chars: int = 3200) -> str:
    if len(text) <= budget_chars:
        return text
    half = budget_chars // 2
    dropped = len(text) - budget_chars
    return f"{text[:half]}\n...[truncated: {dropped} chars dropped from the middle]...\n{text[-half:]}"

go deeper

for a junior

Know that tool output usually cannot go into the transcript untouched, and that any shortened result should say it was shortened.

for a middle

Be ready to compare positional truncation, relevance selection and structured extraction, and to say concretely what each one loses on a large fetched page.

for a senior

Show you would set per-tool budgets, keep a handle back to the full payload, and measure answer retention on replayed traces rather than trusting a default truncation.

for a principal

Own the tradeoff between extraction cost per step and the cost of an agent that re-fetches: argue where the reduction belongs, what it is allowed to discard, and how the policy is audited.

## What is being reduced In a ReAct loop the agent emits a thought, takes an action against a tool, and the tool's output returns as an **observation** appended to the running transcript. That observation is text the model reads on the next step, so it competes for the same context the plan, the instructions, and every earlier step already occupy. Real tools do not return conveniently sized text: a fetch tool returns an entire HTML page, a query tool returns thousands of rows, a log tool returns a megabyte of lines. Reduction is the policy that turns that payload into something the next step can afford. The common framing — 'just truncate it' — hides the real question: reduction is lossy by definition, so the only interesting decision is *what you choose to lose*. ## Family 1: positional truncation Cut the payload at a fixed size. Head-only keeps the first N tokens; head+tail keeps both ends and drops the middle. This is deterministic, costs nothing, and needs no query, which is why almost every agent framework ships it as a default. What it loses is positional, not semantic. On a 40k-token HTML page the top is navigation, cookie banners, and boilerplate, and the tail is the footer; the paragraph that answers the question sits in the middle and is exactly what gets removed. Positional truncation is defensible for outputs whose structure puts the signal at a known end — an exception header at the top of a stack trace, the last lines of a log tail — and poor for everything else. ## Family 2: selection Keep the sub-parts that match the current need: top-k rows by a relevance or recency score, the section whose heading matches the query, the lines matching a pattern, the extracted main content of a page with markup stripped. Selection preserves verbatim text — which matters, because verbatim text is quotable and checkable on the next step — while cutting volume by one or two orders of magnitude. Its cost is that it needs a scoring signal, and its failure is silent: if the agent's query is phrased differently from the document, the matching span is dropped and the observation looks complete but is empty of the answer. Selection also destroys context around each kept fragment, so a row that only makes sense with its column header or its surrounding sentence can become misleading. ## Family 3: structured extraction Run the payload through a parser or a cheap model with a target shape: `{title, published_at, price, availability}` or 'return the three sentences that mention the refund window'. This gives the highest information density per token and produces an observation the next thought can act on directly. The costs are real: an extra model call adds latency and money on every step, extraction can hallucinate or normalise values, and the schema is a prior commitment — anything not in the schema is gone, including the surprising detail that would have redirected the plan. A good compromise is extraction plus a short verbatim excerpt, so the model has both a structured summary and quotable evidence. ## Always mark the cut An observation that was reduced must say so, inline, in the observation text. `[truncated: showing 5 of 1,842 rows, ranked by match score]` changes the model's behaviour: it can page, re-query with a narrower filter, or state that its answer is based on a sample. An unmarked reduction reads as a complete result, and the model will confidently answer 'there are 5 matching orders' from a top-5 slice. ## Keep a handle Pair the reduction with a way back to the full payload: the source URL and a byte or line offset, a cursor token, a temp-file path, a result id. This converts a one-shot lossy cut into a progressive one — the agent sees a compact summary, and pays for detail only where it decides detail matters. It also makes the reduction auditable after the fact. ## Budgeting and measurement Pick a per-observation budget as a fraction of the step budget, not as a global constant: a search tool that runs twenty times per task deserves a far smaller slice than a single deep fetch. Then measure. The honest metric is not tokens saved but *answer retention*: on a replayed set of traces, how often does the reduced observation still contain the span that the reference answer depends on? Teams that track that number usually discover their default head-truncation is dropping the answer on a large minority of fetches. ## Failure modes to name in an interview Unmarked truncation producing confident partial answers; reducing before you know what the step needs; cutting structured formats mid-record so JSON or CSV becomes unparseable; and reducing so aggressively that the model must re-call the same tool three times, which costs more than the tokens saved.

  • When would you deliberately choose head-only truncation over extraction?
    When the signal is known to sit at the front and latency matters: a stack trace whose exception type and top frames are the whole answer, a JSON response whose meaningful fields precede a large payload array, or a tool called dozens of times per task where an extra extraction call per observation would dominate cost. It is a structural bet, not a default.
  • How would you measure whether your reduction policy is losing answers?
    Replay recorded traces with the full payloads kept alongside the reduced ones, and check answer retention: for each task, does the reduced observation still contain the span the reference answer depends on? Report retention per tool, because a policy that is fine for a row-oriented query tool can be badly wrong for a page fetch.
  • What breaks if you truncate a JSON or CSV observation at a fixed character count?
    You cut mid-record and the observation becomes unparseable, so the model either gives up on the structure or, worse, repairs it by inventing the missing closing fields. Reduce structured payloads at record boundaries — drop whole rows or whole objects — and state the count you kept.

Clipping an article for a colleague: photocopying the first and last page is fast but usually misses the paragraph that matters, while highlighting the relevant passages takes effort and preserves the point.

saying these in an interview costs you the question

  • Truncating at a fixed length and never marking it
  • Assuming the answer is at the start of a page
  • Reducing structured output mid-record so it stops parsing
  • Treating a top-k slice as the complete result set
  • Compressing so hard the agent must re-call the same tool

context