Your LangChain chat sessions overflow the context window; how do you bound history?
answer
- bound the prompt, not the store
- trim before the prompt, every turn
- pin the system message
- keep tool call and result together
- truncation is lossy; summarize what matters
basics
~20 sInsert trim_messages from langchain_core.messages as a step before the prompt, so every invocation cuts the loaded history to a token budget. Keep the system message pinned, trim from the oldest end, and start the kept window on a human message so tool replies are never orphaned.
solid answer
~50 sBounding happens at prompt-build time, not in the store. `trim_messages` in `langchain_core.messages` takes `max_tokens` plus a `token_counter` — a chat model, or `count_tokens_approximately` for a cheap local estimate — and a `strategy`, normally `"last"` so you keep the newest turns. Set `include_system=True` to pin instructions, leave `allow_partial=False` so half a message never ships, and set `start_on="human"` so the kept window cannot begin with a tool reply whose originating tool call was cut, which providers reject outright. Compose it as a runnable step ahead of the prompt so it runs per invocation against freshly loaded history. Trimming is a prompt-side view: the store still holds everything, which is what you want for audit, so prune or expire the store separately. When old context genuinely matters, summarize the dropped span and carry the summary in the system message.
code
python · 20 linesfrom langchain_core.messages import AIMessage, HumanMessage, SystemMessage, trim_messages
from langchain_core.messages.utils import count_tokens_approximately
history = [
SystemMessage("You are terse."),
HumanMessage("first question " * 40),
AIMessage("first answer " * 40),
HumanMessage("latest question"),
]
trimmed = trim_messages(
history,
max_tokens=60,
strategy="last",
token_counter=count_tokens_approximately,
include_system=True,
allow_partial=False,
start_on="human",
)
print([type(m).__name__ for m in trimmed])go deeper
Know that history must be capped before it is sent, and that LangChain provides trim_messages with a max token budget for exactly that.
Explain the parameters that matter — strategy last, include_system, a token counter — and why the step has to sit before the prompt inside the chain.
Name the orphaned-tool-message failure and its fix, separate prompt budget from store retention, and describe how you would test it with a synthetic long session.
Decide policy: what fidelity the product owes at turn eighty, whether durable facts move into structured state, and how the prompt budget interacts with cost and latency targets.
## Where the budget is enforced A chat message history is append-only, so an unbounded conversation eventually produces a request the provider rejects. The fix is to treat the stored transcript and the prompt as two different things: the store keeps everything, and a trimming step builds a bounded view of it on every turn. That step belongs inside the chain, ahead of the prompt, so it re-runs each invocation — trimming once at startup does nothing for turn fifty. ## trim_messages `trim_messages` (exported from `langchain_core.messages`) is the built-in. The parameters that matter: - `max_tokens` — the budget for history. Leave headroom for the system prompt, retrieved documents, tool schemas and the response. - `token_counter` — how tokens are measured. Pass a chat model to use its own accounting, or `count_tokens_approximately` from `langchain_core.messages.utils` for a fast local estimate. A crude counter is fine if your headroom is honest. - `strategy` — `"last"` keeps the newest messages, which is what conversations want; `"first"` keeps the oldest, occasionally useful for fixed preambles. - `include_system=True` — keeps the system message even when trimming from the front, so instructions are never the thing you drop. - `allow_partial` — leave it False; shipping half a message wastes budget on an incoherent fragment. - `start_on` / `end_on` — constrain which message types the kept window may begin and end with. ## The tool-call hazard This is the failure senior candidates are expected to name. A tool-using turn is a pair: an `AIMessage` carrying `tool_calls`, followed by `ToolMessage` results referencing those call ids. Trim naively at a token boundary and the window can start with a `ToolMessage` whose originating `AIMessage` was cut. Providers reject that as malformed — a hard error, not a degraded answer — and it only reproduces in long sessions that used tools, so it reaches production. Setting `start_on="human"` makes the kept window begin on a user turn, keeping call and result together. ## Trimming is not deletion What you trim is the prompt, not the record. The store still holds the full transcript, which you usually want for audit, evaluation and replay. That means two separate policies: a token budget on the prompt path, and retention or TTL on the store. Conflating them is how teams end up either paying for prompts they thought were bounded, or losing transcripts they were required to keep. ## When trimming is not enough Dropping the oldest turns is lossy by construction. If the beginning of the conversation carries constraints the model must honour at turn eighty — a chosen plan, a stated budget, an explicit refusal — you need compression, not truncation: summarize the span you are about to drop and carry that summary forward in the system message, or better, extract the durable facts into structured state your prompt renders deterministically. Structured state does not drift the way summarized prose does. ## Verifying it Measure prompt tokens per turn and watch them plateau instead of climbing. Add a test with a synthetic long history including a tool-call pair and assert the trimmed output still starts on a human message and stays under budget. This is the kind of test that pays for itself, because the bug it catches only appears deep in a session.
- Why does start_on='human' matter when the conversation used tools?A tool turn is a pair: an AIMessage carrying tool_calls and the ToolMessage results that reference those call ids. If the trim boundary lands between them, the kept window opens with an orphaned tool result and the provider rejects the request as malformed. Starting the window on a human message keeps each pair intact.
- Should trimming also delete messages from the store?No, keep them separate. Trimming builds a bounded prompt view; the store remains the record you need for audit, evaluation and replay. Apply retention to the store on its own terms, such as a TTL or an archival job, so a prompt-budget change never silently destroys transcripts you were required to keep.
- How do you keep constraints stated at turn one alive at turn eighty?Do not rely on truncation to preserve them. Extract durable facts into structured state your prompt renders deterministically, or summarize the span you are dropping and carry that summary in the system message. Structured state is preferable because summarized prose drifts and quietly loses numbers and negations.
saying these in an interview costs you the question
- Trims once at startup instead of per invocation
- Drops the system message while cutting the oldest messages
- Splits a tool call from its tool result at the trim boundary
- Deletes from the store to bound the prompt
- Sizes the budget with no headroom for the response