skip to content

Context Engineering

Deciding what actually occupies the context window: what to include, how to compress it, what order to put it in, what to carry across turns, and how to stay inside the token budget. It is the discipline that separates a prompt working on one example from one that survives a long session.

on this pageshow

explore

questions

26

When a coding agent compacts its conversation, what actually happens to the session?

level: juniorimportance: must knowfreq 62%

answer

  1. the window fills, the task does not stop
  2. history out, summary in
  3. recent turns kept verbatim
  4. one-way: nothing is reconstructed
  5. system prompt and tools stay put

basics

~10 s

Compaction replaces the accumulated history with a written summary plus the most recent turns, then continues the same task in a nearly empty window. Work carries on, but the raw earlier messages are gone.

solid answer

~50 s

An agent session grows monotonically: every user message, model turn, tool call and tool result is appended, and the whole list is resent on each request because the model holds no state between calls. Eventually the window fills. Compaction is the reinitialization step: the agent has the model write a structured summary of the session so far — goal, decisions taken, artifacts in flight, position in the plan — then rebuilds the context as system prompt, tool definitions, that summary, and the last few turns kept verbatim. The same task continues; it is not a new chat. The property that matters is that compaction is **lossy and one-way from the model's point of view**: anything not captured in the summary and not recoverable by re-reading a file or re-running a query is gone. What goes into the summary therefore matters far more than the compression ratio.

go deeper

for a junior

Be able to say plainly that compaction swaps a long history for a summary plus the newest turns, and that the same task continues afterwards.

for a middle

Explain the sequence — threshold, summarizing call, rebuilt message list — and why the recent turns are kept raw while the middle of the conversation is condensed.

for a senior

Show that you treat it as lossy and instrument for the loss: which facts existed only in the transcript, and how you detect an agent acting as though a constraint no longer exists.

for a principal

Own the position that compaction is one of three answers alongside file-backed notes and sub-agent isolation, and argue when a session should be redesigned so it never needs compacting at all.

## Why a session needs compacting An agent conversation is append-only by construction. The harness keeps a running list of messages — the user's request, the model's reasoning and tool calls, and the output those tools returned — and resends that entire list on every turn, because the model itself carries no state between calls. A long coding or migration session therefore does not merely get slower and costlier; it eventually exceeds the context window outright. Well before that hard ceiling it enters the degraded region practitioners call context rot, where a very full window measurably hurts recall and planning quality even though nothing has technically overflowed. Compaction is the standard response. Alongside externalizing notes to files and isolating work in sub-agents, it is one of the three accepted answers to a filling window. ## The mechanic, step by step 1. **Trigger.** The harness notices utilization crossing a threshold — a fraction of the usable window, not the advertised number. 2. **Summarize.** It issues one extra model call whose prompt is roughly "write down everything a fresh instance of you would need to continue this task," usually against a fixed template so the output is structured rather than narrative. 3. **Rebuild.** It constructs a new message list: the unchanged system prompt and tool definitions, then the summary as a single message, then the most recent turns copied verbatim — the ones the agent is actively working within. 4. **Continue.** The next turn runs against that new context. Utilization drops from, say, 90% to 30%, and the loop proceeds on the same goal. The verbatim tail matters. A summary alone loses the fine texture of what just happened — the exact error string, the half-finished edit — and the agent's next action usually depends on it. Keeping the last few turns raw is cheap and prevents the most common post-compaction stumble, which is an agent that redoes the step it had just completed. ## What compaction is not - **Not a new session.** The goal, the user's constraints and the plan survive by design. If the agent asks the user to restate the task after compacting, the summary was inadequate. - **Not lossless.** Nothing reconstructs the removed messages. Detail that existed only in the conversation — a reason for a decision, a constraint the user gave in passing — is unrecoverable once dropped. - **Not free.** It costs a model call, and it rewrites the front of the context, which throws away the prompt-cache discount the session had been enjoying. - **Not the only lever.** If the bulk of the window is old tool output rather than reasoning, clearing those results is mechanical, cheaper, and preserves the reasoning thread intact. ## Why the loss is asymmetric Some material is recoverable from outside the window and some is not. A file's contents can be re-read; a query can be re-run; a repository can be re-grepped. Those things do not need summarizing — carrying a path or a query identifier is enough, and re-reading gives the current version rather than a stale copy. But a decision the agent made three hours ago and the reason behind it exist nowhere except in the transcript. Same for a constraint the user stated once. That asymmetry is the whole design principle: summarize what only the conversation knows, and keep pointers to everything the world still holds. ## The failure everyone has seen The canonical bad outcome is a summary that captured the mechanics but dropped a standing instruction. A user says early on "never run this against production"; four hours later the session compacts, that sentence is not in the summary, and the agent's next proposal is exactly the forbidden action. It has not malfunctioned — the instruction simply no longer exists anywhere it can see. This is why mature harnesses treat standing constraints as a required, always-carried section of the summary template rather than leaving their survival to the model's judgement.

  • Why keep the last few turns verbatim instead of summarizing everything?
    Because the agent's immediate next action usually depends on fine detail the summary flattens — the exact error text, the half-applied edit, the last tool result. Keeping the tail raw is cheap and prevents the most common post-compaction stumble, where the agent repeats a step it had already finished or misreads where it left off.
  • How is compaction different from just letting the oldest messages be truncated?
    Truncation drops content blindly, so whichever facts sat oldest — usually the goal and the user's constraints, which arrive first — are exactly what disappears. Compaction spends a model call to decide what to keep, so early, load-bearing material survives in condensed form while bulky middle content goes.
  • What signal tells you a compaction went badly?
    The agent immediately redoes completed work, asks the user for information already given, or proposes something a stated constraint forbids. Those are all one symptom: something load-bearing existed only in the removed history. Instrumenting for them is how you tune the summary template.

saying these in an interview costs you the question

  • Thinks compaction starts a new task or clears the goal
  • Believes the original messages can be restored afterwards
  • Assumes compaction is free rather than costing a model call
  • Thinks a bigger context window removes the need for it entirely

context

open as a page

Why do agents load context through tools on demand instead of preloading it?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Preloading spends the attention budget on material the model mostly will not use, and quality degrades as windows fill. Just-in-time loading carries lightweight identifiers — paths, IDs, queries — and pulls full content through a tool only when a step actually needs it.

open as a page

Why does an LLM assistant forget facts between sessions, and how do you fix it?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Everything the model can use in a turn is text in its context window, and that window is rebuilt from scratch for each new session. A fact survives only if the application writes it to a store outside the window and puts it back in later.

open as a page

In an LLM prompt containing a long document, where should the instruction go?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Put the instruction after the document, or at both ends. Language models attend most reliably to the beginning and the end of a prompt, so an instruction stated once before a long body is the easiest part to miss.

open as a page

Why reserve output tokens when budgeting an LLM context window?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A context window is shared by the prompt and everything the model generates. Fill it with input and the answer has nowhere to go: responses get truncated mid-sentence or the request is rejected. Reserve output space first.

open as a page

What must a compaction summary carry forward in a long agent session?

level: middleimportance: must knowfreq 68%

basics

~20 s

Four things only the conversation knows: the goal with any standing constraints, the decisions already made and why, the artifacts still open mid-edit, and the approaches already ruled out. Anything readable from a file or a query should be a pointer, not a copy.

open as a page

How does progressive disclosure let an agent work over 300 warehouse tables?

level: middleimportance: must knowfreq 56%

basics

~20 s

Progressive disclosure gives the model identifiers and thin metadata first — 300 table names with row counts — and full DDL only for the two or three tables a query actually joins. Selection happens over cheap descriptors; expensive content is expanded on request.

open as a page

A memory file loaded on every turn has grown to 8k stale tokens — how do you fix it?

level: middleimportance: must knowfreq 55%

basics

~20 s

Split the file by how often its content is actually needed. Keep only always-true conventions in the always-loaded file, move task-specific procedures into on-demand skill files that load when the task matches, and delete anything stale outright rather than leaving it in place.

open as a page

In a 20-document prompt, why is accuracy lowest when the answer sits in the middle?

level: middleimportance: must knowfreq 72%

basics

~20 s

This is the lost-in-the-middle effect: retrieval accuracy traces a U-shaped curve against position. Models use information best when it appears first or last in the prompt and measurably worst when it is buried in the middle, even though every document is inside the window.

open as a page

When should an agent grep a repository instead of querying an embedding index?

level: middleimportance: must knowfreq 64%

basics

~20 s

Grep when the query is an exact string — a symbol, an import, an error code — because lexical search returns every match and reads the current files. Embedding search earns its cost only when the wording is unknown.

open as a page

Why is a model's advertised 1M-token window not 1M usable tokens?

level: middleimportance: must knowfreq 78%

basics

~20 s

Advertised length is an acceptance limit, not a quality guarantee. Retrieval accuracy and instruction-following degrade as the window fills — the effect commonly called context rot — so the usable ceiling is typically a fraction of the maximum, and it must be measured.

open as a page

When is clearing old tool results cheaper than compacting an agent conversation?

level: middleimportance: should knowfreq 48%

basics

~20 s

When the window is full of bulky tool output rather than long reasoning. Dropping old tool results is mechanical, costs no model call, and leaves the reasoning thread intact — compaction is for when the conversation itself, not its payloads, has grown long.

open as a page

Why does compacting an agent conversation wipe out its prompt-cache savings?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Cache reuse requires the request prefix to be identical to a previous one. Compaction rewrites the history near the front, so everything after the edit is new and the next turn is billed and computed as an uncached prefill.

open as a page

At what window utilization should a long agent compact, and how do you prevent thrash?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Trigger on a high-water mark measured against the usable ceiling, not the advertised window, with headroom for the largest plausible next tool result. Then compact deep, to a low-water target, so the session cannot re-trigger within a few turns.

open as a page

How do you stop just-in-time loading from invalidating a cached prompt prefix?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Keep the prompt append-only. Put stable material — system instructions, tool definitions, the catalog — at the front, never reorder or edit it, and append everything fetched mid-session at the end. Anything inserted above a cache boundary invalidates every cached token after it.

open as a page

Preload 40k tokens of API reference, or take three round trips to fetch one endpoint?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Count round trips, not dollars. Preloading costs one hop and a permanently degraded window; on-demand fetching costs a hop per hop and a cleaner window. Decide on hit rate, how serial the hops are, and the task's latency budget — then measure both.

open as a page

A user revises a fact you already stored in memory — how do you resolve the write?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Compare the candidate fact against what is already stored and choose one outcome: add it as new, update the existing entry in place, delete an entry it invalidates, or do nothing when it is already known. Blind appending leaves both versions alive.

open as a page

A 40-page RFP review ignores the output template stated at the top — how do you fix it?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Two levers exist: restate the template at the end of the prompt, or reorder so the instruction follows the document instead of preceding it. Restating costs a duplicate instruction per call; reordering is free but changes the prompt shape. Shrinking the document reduces the exposure behind both.

open as a page

Does position bias get worse as you fill more of a million-token context window?

level: seniorimportance: should knowfreq 31%

basics

~20 s

Yes. Position effects are mild when a prompt uses a small fraction of the window and grow sharply as utilization rises. As of mid-2026, reported behaviour is that past roughly half the window the U-shape flattens into a distance-based bias favouring the end, so early material loses its protection.

open as a page

How do you set chunk size and top-k when retrieved passages share the context window?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Chunk size times top-k is the token bill. Pick the smallest k at which end-to-end answer accuracy stops improving — not the k that maximises retrieval recall — and deduplicate near-identical passages before they crowd out the one that matters.

open as a page

In late chunking, why is the whole document encoded before chunk vectors are pooled?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Late chunking encodes the entire document first so every token attends to the rest of it, then pools token embeddings per chunk. Each chunk vector therefore carries document context that independent chunk-by-chunk embedding throws away.

open as a page

How would you measure the usable context ceiling for your own workload?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Hold one graded task set fixed and re-run it at several fill levels — for example 20%, 50% and 80% of the window — padding with realistic in-domain material. Plot score against utilization, then budget below where the curve breaks.

open as a page

What does a passing needle-in-a-haystack score fail to prove about long context?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Needle-in-a-haystack plants one distinctive sentence and asks for it back, which is close to a keyword lookup. A perfect score says nothing about tracking many facts, reconciling contradictions, or reasoning across a long document — the things real workloads need.

open as a page

How do you stop an agent's long-term memory store from growing into noise?

level: principalimportance: should knowfreq 33%

basics

~20 s

Persist by exception rather than by default, consolidate asynchronously between sessions to merge duplicates and promote repeated patterns, and expire facts by class instead of one global lifetime. Growth is a policy problem: without one, every store eventually costs more than it returns.

open as a page

When is agentic chunking worth its cost for an internal document corpus?

level: principalimportance: should knowfreq 26%

basics

~20 s

Agentic chunking pays when documents lack reliable structural markers, the corpus is small and slow-changing, and retrieval failures trace to boundaries cutting through procedures. Otherwise its per-document model call, nondeterminism and re-embedding cost outweigh cheaper structural splitting.

open as a page

Your agent lists a case-file directory, then opens every document anyway — why?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Almost always because the listing does not discriminate. Entries like doc_final_v3.pdf give the model nothing to choose on, so it expands everything — paying a listing round trip on top of loading the whole corpus, which is worse than preloading.

open as a page