Why does interleaved thinking require passing earlier thinking blocks back to the model?
answer
- The API is stateless; the transcript is the state
- Thinking after each observation, not just before
- Return unmodified, not summarized
- Editing history breaks the cache prefix
- What the harness deletes, the model forgets
basics
~20 sBecause the reasoning lives in the transcript, not on the server. Interleaved thinking means the model thinks again after each tool result; if the earlier thinking is stripped before the next call, that chain is gone and the model re-derives or contradicts its own plan.
solid answer
~50 sInterleaved thinking is a model deliberating **between** tool calls, not only before the first one: read the disruption feed, think, query seat inventory, think again about what the inventory implies. That mid-loop reasoning is emitted as part of the assistant turn, and the API is stateless — everything the model knows on the next call is what you send it. Strip the thinking blocks and you have handed back tool results with no record of why those tools were called; the model reconstructs a plan from the outputs alone, which is where you see repeated calls, abandoned constraints and reversals mid-task. Providers commonly require the blocks be returned **unmodified**, and may verify them, so summarizing or rewriting them is not an option — some are returned opaque precisely so you cannot. Editing history also breaks the prompt-cache prefix, so a well-meant trim costs correctness and money at once.
go deeper
Know that the model has no memory between API calls: whatever your code sends is all it sees. If the earlier thinking is not in the request, it did not happen as far as the model is concerned.
Explain the loop mechanics — assistant turn with thinking and tool calls, results appended, call again — and why the thinking must be carried unmodified rather than summarized or dropped.
Diagnose from symptoms: repeated tool calls, violated earlier constraints and mid-task reversals point at a harness discarding state. Connect it to prompt-cache prefix invalidation so you can explain why the naive trim costs both correctness and money.
Frame the transcript as the agent's working memory and your harness as its memory manager. Own the policy for when compaction happens, what survives it, and how that policy stays correct across providers with different retention and opacity rules.
## What interleaved thinking is A non-interleaved thinking model deliberates once, up front, then acts. Interleaved thinking means the model can think **again after every observation**: it reads a tool result, reasons about what that result implies, and only then decides the next action. The difference shows up on exactly the tasks where an agent earns its keep. Take an airline irregular-operations agent handling a cancelled leg. It reads the disruption feed and thinks about which passengers are connection-critical. It queries seat inventory on the alternates. Now it must think *again* — the inventory came back thinner than expected, so the plan of rebooking everyone on one flight is dead and a split across two carriers is on the table. Without a thinking step after that observation, the model is choosing its next tool call reflexively, at the exact moment the situation changed. ## Why the blocks have to travel back The chat completion interface is stateless. The provider is not holding your agent's train of thought between calls; the conversation *is* the state. Everything the model will know on turn N+1 is precisely what you serialize into the request. So when your loop does the standard thing — take the assistant's turn, execute the tools, append the results, call again — whatever you dropped from that assistant turn is permanently gone from the model's view. If the thinking is what you dropped, the model now sees tool results attached to calls whose *rationale* has vanished. The observable failure modes are consistent: - **Redundant work.** It re-calls a tool because nothing in context records that it already reasoned about that data. - **Constraint amnesia.** A constraint it derived in step two ("this passenger has a minimum connection time we can't violate") existed only in the thinking, so step five violates it. - **Reversals.** It reaches a conclusion, loses the reasoning, and reasons its way to the opposite conclusion two steps later. - **Cost inflation.** Every re-derivation is thinking tokens you pay for twice. These look like model flakiness in a bug report. They are usually a harness that is throwing away state. ## Unmodified, not merely present A subtlety worth knowing, because it defeats the obvious optimization. Providers commonly require the thinking blocks to be returned **byte-for-byte as issued**, and may attach an integrity signature that is checked on the next call. Some providers return the thinking as an opaque or encrypted block you cannot read at all — you are expected to store it and hand it back. That rules out the tempting move of summarizing long thinking to save context. Rewriting a block either fails validation or degrades the continuation. Treat these blocks as **opaque state you carry**, not as text you own. ## The prompt-cache interaction Providers cache on a **prefix** basis: reuse is granted only for the longest identical leading stretch of the request. An agent loop is the ideal cache workload — each iteration is the previous request plus an appended result — and that is why long tool loops are affordable at all. Edit anything earlier in the transcript and you invalidate the cache from that point. Trimming thinking blocks from the middle of a live loop therefore does something worse than saving nothing: it destroys the reasoning chain *and* forces a full re-read of the prefix at uncached rates. Two regressions from one well-intentioned change. ## Where the context actually goes Interleaved thinking obviously grows the transcript, and on a twenty-step loop that growth is real. Two things keep it manageable. First, the retention requirement is scoped to the **live loop**. Within one assistant turn's chain of tool calls, thinking must be preserved because it is the working state. Once a turn completes and the user speaks again, providers commonly stop counting earlier turns' thinking against context — the deliberation that got you to a finished answer is no longer load-bearing. Second, when you genuinely must shrink a long-running session, the right move is a **deliberate compaction step** that rebuilds the transcript into a fresh, coherent summarized state and continues from there — not a silent trim of blocks out from under an in-flight loop. ## The design principle underneath A thinking model's reasoning is not a log. It is working memory, and it lives in a place your code controls. Anything your harness deletes, the model forgets. That reframing — the transcript is the agent's memory, and you are its memory manager — is the answer a senior interviewer is listening for, with the thinking blocks as the specific case.
- The transcript is getting long. Why not summarize the thinking blocks instead of carrying them whole?Because providers commonly require them returned unmodified and may verify integrity, and some are opaque so you cannot summarize them anyway. Any edit also invalidates the prompt-cache prefix from that point, so you lose continuity and pay full price for the re-read. If size is the real problem, run an explicit compaction step that rebuilds the state and continues, rather than trimming a live loop.
- How would you tell from production traces that interleaved thinking is being lost?Look for signatures of amnesia: the same tool called twice with identical arguments, constraints established early being violated later, plans reversing without new evidence, and thinking-token counts that stay high late in a loop because the model keeps re-deriving. Diff what your harness actually sends on iteration N against what the model emitted on N-1 — the gap is usually right there.
- Does interleaved thinking change how you handle a tool that returns an error?It makes it more valuable. With a thinking step after the observation, the model can reason about what the failure means — bad arguments, a genuinely unavailable resource, a transient fault — and pick a different strategy instead of retrying reflexively. That only works if the failure and the reasoning about it both stay in the transcript for subsequent iterations.
saying these in an interview costs you the question
- The provider remembers the reasoning between calls
- Thinking blocks can be summarized to save context
- Dropping thinking only costs tokens, not correctness
- Interleaved thinking means one long think up front
- Trimming old blocks is a free context optimization