Why do multi-agent systems keep shared state in an artifact store plus an append-only event log?
answer
- state must outlive any one context
- two stores, two different jobs
- artifacts are nouns, events are verbs
- one JSON record per line, appended
- replay rebuilds any derived view
basics
~20 sAn artifact store plus an append-only event log gives every agent one durable, inspectable source of truth that lives outside any context window. State then survives restarts, and no agent's view depends on which messages it happened to see.
solid answer
~50 sShared state that lives only inside contexts or in the messages agents exchange is invisible to any agent that starts late, restarts, or runs in parallel. So the state design that shipped in practice is two stores. An **artifact store** — a directory or bucket — holds the durable nouns: for a construction crew that is `/site/daily-logs/`, `/site/rfis/`, drawings, the schedule. An **append-only event log** — one JSON record per line — holds the verbs: RFI 118 opened, moved to in_review, answered. Artifacts are addressed by stable path and read on demand; the log is ordered, cheap to scan, and never overwritten, so history stays recoverable and any derived view can be rebuilt by replay. Together they give durability across crashes, a shared vocabulary between agents, and the raw trace you need to debug a run afterwards.
code
python · 14 linesimport json, time
def append_event(path, event):
event["ts"] = time.time()
with open(path, "a", encoding="utf-8") as f:
f.write(json.dumps(event) + "\n")
append_event("events.jsonl", {
"agent": "field-agent",
"task": "rfi-118",
"from": "open",
"to": "in_review",
"artifact": "/site/rfis/rfi-118.md",
})go deeper
Be able to say that shared state lives in files and a log outside the model, not in the chat history, and that agents read it by path when they need it.
Explain the split — artifacts for content, an ordered append-only log for transitions — and why appending rather than overwriting preserves history, supports replay, and makes many producers safe.
Show the operational consequences: recovering a crashed run from the log, using it for failure attribution, snapshotting derived views, and setting retention before the log becomes unscannable.
Own the schema and lifecycle as a platform concern — event versioning across agents built months apart, what is authoritative versus derived, retention and cost, and how much structure to impose before it slows delivery.
## What "shared state" means here In a multi-agent system, several model-driven agents work on one task. Shared state is everything they must agree on that outlives a single model call: the documents being produced, the facts discovered, and the status of each unit of work. The design question is where that state lives. There are three candidate homes — inside each agent's context window, inside the messages agents pass to each other, or in an external store all of them read and write. Only the third holds up under a real workload. ## Why contexts and messages are the wrong home A context window is per-agent, per-run, bounded and volatile. Two agents working the same task have different windows, and neither can see the other's. If a fact exists only because agent B read agent A's message, then an agent spawned in parallel, restarted after a crash, or added later has an incomplete picture — and nothing in the system can tell you it is incomplete. Messages are also a poor archive: they are ordered by delivery, not by meaning, and they mix reasoning with results. State pinned to a conversation dies with the conversation. ## The two stores **Artifact store.** A filesystem directory, object bucket, or document table holding the durable outputs. For a construction-project crew: `/site/daily-logs/2026-07-11.md`, `/site/rfis/rfi-118.md`, the drawing set, `schedule.json`. Artifacts are large, addressed by a stable name, and read on demand — an agent puts the *path* in its context, not the bytes. **Append-only event log.** A single ordered stream — commonly one JSON object per line in `events.jsonl` — recording transitions: `rfi-118 opened by field-agent`, `rfi-118 -> in_review`, `rfi-118 answered, artifact /site/rfis/rfi-118-answer.md`. Records are small, timestamped, and attributed to the agent that wrote them. A useful split: **artifacts are nouns, events are verbs**. A forty-page inspection report is an artifact; the fact that it was produced, by whom, and that review is now unblocked is an event. Never put the report body in the log — put the path and a one-line summary. ## Why append-only specifically Overwriting a mutable status field destroys the reason the status changed, and two agents editing the same field race. Appending does neither. Each writer only ever adds its own record, so many producers are safe without coordination on the log itself. Three properties follow: - **Replay.** Any derived view — a per-task status table, a dashboard, a summary for a new agent — can be rebuilt from the log, so you are never stuck trusting a cache. - **Attribution.** When a run goes wrong, the log is the raw material for working out which agent's step introduced the error. Step-level attribution remains hard even with full traces, and it is impossible without them. - **Audit.** "Who changed this and when" is answerable by construction rather than by instrumentation added after an incident. ## Consuming it Agents are not handed the store; they are handed pointers and tools to read it. A reviewer agent asked to check open RFIs scans the tail of the log for `status=open`, then opens only the two artifacts it needs. This is what keeps six agents and hundreds of megabytes of drawings workable when no single window could hold a fraction of it. ## What it costs You now own a schema. Event records need a shape and a version, or agents six weeks apart write incompatible lines. The log grows without bound, so it needs retention or periodic snapshotting into a compacted view. And free-form prose written into shared files by one agent is frequently unparseable to another — which is the argument for constraining writes to typed operations rather than raw text dumps. None of these costs is a reason to go back to conversation-resident state; they are the ordinary price of durability. ## In an interview The answer that lands is not "use a database". It is the separation: a big, lazily-read artifact store for content, plus a small ordered log for facts about progress, both outside every context window, both readable by an agent that just booted with no history. As of mid-2026 that is the shape essentially every production multi-agent system converged on.
- How does a newly spawned agent get up to speed without replaying the whole log?Give it a compacted view rather than raw history: a periodically snapshotted status table derived from the log, plus the paths of the artifacts relevant to its assignment. The full log stays available for debugging and replay, but routine reads hit the derived view. The rule is that the snapshot is always reconstructible from the log, never the only copy.
- What belongs in the event log versus the artifact store when an agent produces a large analysis?The analysis document goes to the artifact store under a stable path. The log gets a small record naming the agent, the task, the resulting status, and the artifact path — plus at most a one-line summary. Putting document bodies in the log makes it expensive to scan and destroys its value as a cheap ordered index of what happened.
- Does the append-only log make concurrent writes safe in general?Only for the log itself, because each producer appends its own distinct record and nothing is overwritten. It says nothing about two agents editing the same artifact — that still needs an ownership rule for the file. The log records that both attempted a change; it does not merge them.
saying these in an interview costs you the question
- Says shared state is just the conversation history all agents see
- Puts full documents inside the event log instead of paths
- Treats a mutable status field as equivalent to an append-only log
- Assumes every agent can hold the whole shared state in context
- Believes append-only logging makes concurrent artifact edits safe