skip to content

Why should agents keep private scratchpads and promote only verified results to shared state?

level: seniorimportance: must knowfreq 58%

answer

  1. shared state is somebody else's input
  2. a guess in a file reads as fact
  3. one messy zone, one settled zone
  4. the gate is what makes it worth anything
  5. short, grounded, attributable crosses over

basics

~20 s

Anything written to shared state is read by other agents as established fact, so an unverified guess propagates as truth. Keeping exploration in a private scratchpad and promoting only checked, compressed results puts a verification gate between reasoning and everyone else's input.

solid answer

~50 s

Shared state is an input to every other agent, and agents have no way to tell a hypothesis from a conclusion once it is sitting in a shared file. So the boundary matters more than the storage. Each agent works in a private scratchpad — dead ends, partial reads, tentative interpretations, notes to itself — and **promotes** only what passed a check: a schema-valid artifact, a claim backed by a tool result or citation, a summary a verifier accepted. In a construction crew, the field agent may work through six ambiguous drawing notes privately and promote one verified RFI summary into `/site/rfis/`. Promotion is also compression: what crosses the boundary is short, typed and attributable, not the reasoning transcript. This buys two things at once — other agents' contexts stay clean, and one agent's mistaken guess cannot silently become the whole system's premise.

go deeper

for a junior

Know the two zones: private notes an agent keeps for itself, and shared files everyone reads. Say plainly that only checked results should go into the shared ones.

for a middle

Explain both gains — other agents' contexts stay small, and unverified claims cannot propagate — and describe what a promotion actually contains: a short grounded summary plus an artifact path.

for a senior

Rank the verification gates honestly, argue for a reviewer with a clean context, and describe over-promotion, under-promotion and confidence laundering as the concrete failure modes you design against.

for a principal

Decide where the promotion bar sits per surface and who pays for it: which outputs justify an independent verifier, how much compression is acceptable before findings are lost, and how to keep the boundary enforced structurally rather than by prompt.

## The asymmetry that makes this necessary Within one agent, a wrong intermediate thought is usually recoverable: the next step contradicts it, a tool returns something inconsistent, and the agent revises. Once that same thought is written into shared state, the recovery path closes. Other agents read shared files as ground truth — there is no marker in a file that says "the author was only guessing" — and they build on it. One agent's speculation becomes three agents' premise, and by the time the error surfaces, several artifacts encode it. This asymmetry is why the scratchpad boundary is a first-class design element rather than a tidiness preference. ## Two zones **Private scratchpad.** Per-agent, per-run working notes: partial file reads, candidate interpretations, ruled-out options, a running plan. Nobody else reads it. It can be messy, contradictory and long, because its only consumer is the agent that wrote it and it disappears with the run. **Shared state.** The artifact store and event log every agent reads. Everything here is treated as settled. It should be small, typed, attributable, and true. A field agent on a construction project might read twelve drawing notes, form and discard three readings of an ambiguous dimension, and end with one defensible question. The twelve notes and three discarded readings stay in the scratchpad. The one question, formatted as an RFI with the sheet reference attached, is promoted to `/site/rfis/`. ## What "verified" has to mean Promotion is only worth something if the gate is real. Ranked roughly by strength: 1. **Programmatic verification.** A test passes, a schema validates, a compiler accepts, a total reconciles. The strongest gate, available more often than teams assume. 2. **Grounding in a tool result.** The claim carries a citation — this file, this line, this query result — so a downstream agent or human can check it without redoing the work. 3. **Independent review.** A second agent with its own clean context checks the candidate before it is promoted. Clean context matters: a reviewer that shared the author's reasoning inherits the author's blind spot. 4. **Self-assessment.** An agent judging its own output. Weakest by a wide margin, and on its own it does not justify promotion of anything consequential. A design that promotes on "the agent said it was done" has a scratchpad boundary in name only. ## Promotion is also compression The second gain is contextual. If everything an agent produced flowed into shared state, every other agent would pay to read the reasoning of every agent — the token cost and the attention dilution that isolated subagents exist to avoid. Promotion forces a decision about what is worth the other agents' attention. A subagent that explored for twenty thousand tokens should typically hand back one or two thousand: the finding, its grounding, and the artifact path. That compression is lossy on purpose, and the loss is the point — but it is also the main risk, discussed next. ## Where it goes wrong - **Over-promotion.** Agents dump raw notes into shared files because it is easier than deciding what matters. Shared state becomes an unreadable log, and downstream agents cannot distinguish signal. - **Under-promotion.** A finding that other agents needed stays in the scratchpad and dies with the run. This is the more expensive failure: the work was done, paid for, and lost. - **Confidence laundering.** A hedged scratchpad note ("possibly a dimension conflict") becomes a flat assertion in the promoted summary. Carry the uncertainty across the boundary explicitly, or the gate is actively harmful. - **Unattributed promotion.** A promoted artifact with no record of which agent produced it, from what inputs, is unusable in a post-incident review. ## How to make the boundary enforceable Do not rely on instructions alone. Give scratchpads a location no other agent reads and no shared index scans. Make promotion an explicit operation with a typed shape — a summary plus a grounding reference plus an artifact path — rather than an ordinary file write into a shared directory. Log the promotion as an event so the trail exists. Structure beats exhortation here, because the failure is quiet and shows up several agents downstream. ## In an interview Lead with the propagation argument: shared state is other agents' input, so writing to it is an assertion about truth. Then name the gate you would use and be honest that self-assessment is the weakest one. Being able to describe both the over- and under-promotion failure modes is what separates a considered answer from a slogan.

  • What is the risk of under-promotion, and how would you detect it?
    A finding other agents needed stays private and dies with the run, so the work is paid for and lost — often surfacing as a later agent rediscovering the same fact. Detect it by looking for duplicated effort across traces, and by checking whether promoted artifacts answer the questions downstream agents actually asked. Explicit hand-back schemas that require naming findings reduce silent loss.
  • Why should a reviewer agent that gates promotion get a clean context rather than the author's?
    A reviewer that inherits the author's reasoning inherits its blind spots and tends to ratify rather than check. Given only the candidate artifact and the source material, it re-derives independently and catches assumptions the author never examined. The deliberate context separation is the source of the signal — sharing history to save tokens usually destroys the value of the review.
  • How do you keep uncertainty from being lost when a hedged note is promoted?
    Make the promoted shape carry it: a confidence or status field, a required grounding reference, and an explicit open-question section. Without those, models tend to flatten hedged reasoning into confident prose during summarization, which is worse than not promoting at all because downstream agents now build on false certainty.

saying these in an interview costs you the question

  • Treats shared state as a general dumping ground for agent notes
  • Promotes on the agent's own claim that it finished
  • Says the scratchpad is only about saving tokens
  • Flattens hedged findings into confident assertions when promoting
  • Promotes artifacts with no record of which agent produced them

context