Why rebuild an agent's context each turn from a durable session event log?
answer
- context is derived, not stored
- the log is the system of record
- replay to rebuild the window
- checkpoints give you a rewind point
- inputs replay, outputs still vary
basics
~20 sBecause the context window is derived state, not the system of record. An append-only event log lets the harness resume after a crash, rewind to a checkpoint, re-derive a compacted window, change the context policy without losing history, and reconstruct exactly what the model saw.
solid answer
~50 sTreat the window as a projection: the durable truth is an append-only log of user turns, model outputs, tool calls and results, checkpoints and approvals, and the context builder derives the next request from it every turn. The payoff is operational. A crashed or paused run resumes by replaying the log. A checkpoint before an irreversible action gives you a rewind point, so an operator can go back and take a different branch rather than starting over. Compaction becomes safe, because a summary that dropped something is a lossy *view* and the underlying events are still there. You can change the context policy — keep more, summarize differently, reorder — and re-derive old sessions under the new policy to compare. And in a post-incident review you can answer the only question that matters: what was in the window on the call that went wrong. What you do not get is deterministic reproduction; model sampling still varies, so replay reconstructs inputs, not outputs.
code
python · 21 linesevents = [
{"type": "user", "text": "restart the payments deployment"},
{"type": "tool_result", "handle": "/runs/7/out.log", "digest": "3 pods pending"},
{"type": "checkpoint", "id": "cp-2"},
{"type": "assistant", "text": "rolling back to revision 41"},
]
def build_context(log, keep_full_after):
past_checkpoint, out = False, []
for e in log:
if e["type"] == "checkpoint":
past_checkpoint = e["id"] == keep_full_after
continue
if past_checkpoint:
out.append(e)
else:
out.append({"type": e["type"], "summary": e.get("digest") or e["text"][:40]})
return out
for entry in build_context(events, "cp-2"):
print(entry)go deeper
Know that the conversation an application sends to the model is assembled by code from stored history, and that storing that history is what lets a session survive a restart.
Explain that the window is derived state while the log is the record, and name what the log holds: user turns, model outputs, tool calls and results, checkpoints. Describe resume as replay.
Show what the arrangement enables in production — rewind from a checkpoint before an irreversible action, safe compaction, and reconstructing exactly what the model saw during an incident — and be honest that outputs still vary on replay.
Own the tradeoff: storage, retention, privacy and schema versioning against resumability, auditability and the ability to re-derive history under a new context policy. Set the fidelity boundary by session class and version the context builder.
## Derived state versus system of record The easiest way to build an LLM app is to keep a growing array of messages, append to it, and send it. That array then *is* the application's state. Every serious harness eventually moves to the opposite arrangement: an append-only **session event log** is the system of record, and the context window is recomputed from it before each model call. The log records what happened, not what was sent. Entries include user turns, model outputs (including which tools it asked for), tool invocations with arguments, tool results or handles to them, errors, checkpoints, human approvals, compaction boundaries, and metadata such as model identity, prompt version and token counts. The context builder is a pure-ish function over that log: `context = build(log, policy)`. ## What the arrangement buys you **Resume.** Sessions outlive processes. An agent working a long task will span deployments, restarts and network failures. If state lives in a process's message array, a restart is a lost session; if it lives in the log, resume is replay. **Rewind.** This is the one that justifies checkpointing. Write a checkpoint before a consequential step — say, before an agent executes a destructive rollback in a cluster — and you have a point the run can return to. When the decision turns out wrong, an operator rewinds to the checkpoint, adds a correction, and continues from there, rather than restarting a six-hour session. Without a log there is nothing to rewind *to*. **Safe compaction.** Compaction is lossy by construction. That is tolerable precisely because the summary is a view: whatever the summary dropped is still in the log and still reachable through handles. If the compacted window were the state, compaction would be destruction and every summarization bug would be permanent data loss. **Policy evolution.** Context strategy is the highest-leverage thing you tune, and tuning it requires comparison. With logs you can re-derive historical sessions under a new builder policy — keep more recent turns, summarize differently, drop tool chatter earlier — and see what the model would have been shown. Without them, every policy change is untestable against the past. **Forensics.** When an agent does something surprising, the first question is always what it saw. A log plus a deterministic builder answers it exactly. Reconstructing it from application logs and guesswork does not, and this is usually the moment teams discover their builder was string concatenation scattered across three files. ## What it does not buy you Be precise here, because overclaiming is a common interview stumble. **Not deterministic output.** Replaying the log rebuilds the same *inputs*; the model may still produce different outputs, and providers change models underneath you. If you need bit-identical replay — for a regression suite, say — you must record model responses and replay those, which is a different mechanism (record/replay) layered on the same log. **Not freedom from context management.** A log makes compaction safe; it does not make it unnecessary. The window still has a budget. **Not correctness of tools.** Replaying a tool call from the log means replaying a *recorded result*. Actually re-executing it hits a world that has moved on, which is why non-idempotent tools need explicit thought about whether a resumed run re-runs them or reuses the recorded result. ## Cost and design choices The arrangement is not free. **Storage and retention.** Full logs with raw tool results are large. In practice teams store events with handles to payloads rather than payloads inline, and set retention by session class — long for anything auditable, short for casual chat. **Privacy.** A durable log of everything a user typed and everything a tool returned is a sensitive asset. Redaction at write time, encryption, tenant scoping and deletion-on-request need to be designed in, not retrofitted. **Schema evolution.** Logs outlive code. Event types will change; version them, and keep old builders able to read old events or accept that historical re-derivation has a horizon. **Write amplification.** Every turn writes several events. For high-volume, low-value sessions, teams often log a reduced event set and reserve full fidelity for sessions with side effects or compliance weight. ## Where the boundary sits A reasonable default: log everything that has a side effect or that a human might have to explain later, at full fidelity; log everything else at reduced fidelity; keep payloads behind handles; and make the context builder an explicitly versioned component whose version is recorded on every call. That last detail is small and repeatedly saves incidents — knowing that a session ran under builder v7 rather than v8 is often the whole answer. The framing to carry into a design round: this is the same event-sourcing tradeoff that mature stateful systems make, applied to a window of tokens. You accept storage and schema cost in exchange for resumability, rewind, auditability and the ability to change your mind about what the model should see.
- Does an event log give you deterministic replay of an agent run?It gives deterministic reconstruction of inputs, not of outputs. Sampling is stochastic and providers update models, so replaying the same context can produce a different action. For bit-identical replay you additionally record each model response and serve it from the recording — record/replay layered on the log. Say this plainly; claiming full determinism from event sourcing alone is a common overreach.
- What must a logged tool call contain to stay useful on resume?The tool name and version, the exact arguments, a timestamp, and either the result or a durable handle to it, plus whether the call had side effects. The side-effect flag is what lets a resumed run decide between re-executing and reusing the recorded result — re-running a non-idempotent action during resume is a classic way to double-charge or double-deploy.
- What are the real costs of keeping a full session event log?Storage volume, privacy exposure, and schema evolution. Full logs with raw payloads grow fast, so store handles rather than payloads and set retention by session class. A durable record of user input and tool output is sensitive, so redact at write time and scope by tenant. And because logs outlive code, event schemas need versioning or historical re-derivation stops working.
saying these in an interview costs you the question
- Treats the message array in memory as the session state
- Claims event sourcing makes agent runs fully reproducible
- Says a log removes the need for compaction
- Replays non-idempotent tool calls on resume without thought
- Stores raw multi-megabyte payloads inline in the log