skip to content

Why is agent procedural memory usually stored as code or skill files, not prose?

level: seniorimportance: should knowfreq 30%

answer

  1. procedures have a pass condition, facts do not
  2. execute and verify, then keep
  3. diffs, review, rollback
  4. load the name, expand on use
  5. Voyager's growing skill library

basics

~20 s

Procedures need to be verified, repeated exactly, reviewed and versioned. Code and structured skill files give all four — you can test them, diff them and roll them back. Prose "lessons learned" cannot be tested, accumulate contradictions, and are recalled fuzzily rather than executed.

solid answer

~60 s

Procedural memory is know-how the agent reuses: how to run a deployment check, how to craft a tool in a game world, how to structure a Socratic hint sequence. Storing it as code or as a structured skill file rather than as retrieved prose buys four things. **Verifiability** — a function either passes its test or does not, so the agent can confirm a skill works before keeping it; a paragraph has no pass condition. **Determinism** — a stored program runs the same way twice, whereas re-deriving a procedure from a remembered description reintroduces variance every time. **Reviewability** — skills live in version control, get diffed, reviewed and rolled back, which matters because a skill is executable and therefore a security surface. **Token economy** — the agent can hold only a name and a short description for each skill and load the full instructions and bundled scripts on activation. Voyager, the Minecraft agent, is the canonical demonstration: it wrote programs, verified them against the environment, kept the successful ones in a growing skill library, and composed them into harder tasks later. Prose still has a place for judgment-shaped guidance that cannot be executed.

code

python · 20 lines
python
SKILLS = {}


def skill(name, description):
    def register(fn):
        SKILLS[name] = {"description": description, "fn": fn}
        return fn

    return register


@skill("craft_stone_pickaxe", "Gather wood and stone, then craft a stone pickaxe")
def craft_stone_pickaxe(world):
    world.append("stone_pickaxe")
    return world


print(sorted(SKILLS))
print(SKILLS["craft_stone_pickaxe"]["description"])
print(craft_stone_pickaxe([]))

go deeper

for a junior

Know that procedural memory is reusable how-to and that it is normally kept as code or skill files, because code can be run and tested while remembered advice cannot.

for a middle

Give the reasons in order — verifiability, determinism, reviewability, token cost — and describe the tiered loading that keeps a big skill library out of the context window until a skill is used.

for a senior

Treat agent-authored skills as a security surface: version control, review, rollback, sandboxed execution and scoped credentials. Be ready to describe how a verified skill library compounds capability while prose lessons decay into contradictions.

for a principal

Own the boundary between what becomes code and what stays written guidance, and the governance around it — who reviews a new skill, how a bad one is revoked across every running agent, and what that promotion process costs the team.

## What procedural memory is Of the four agent memory types, procedural memory is the one about capability rather than content. Working memory is the current task, episodic memory is what happened, semantic memory is what is true — procedural memory is how to do something, packaged so it can be done again without rediscovering it. The interesting design question is not whether to keep it but what form to keep it in, and the field converged on executable artifacts rather than remembered text for reasons worth being able to list. ## Verifiability is the deciding property A procedure has a success condition; a fact does not. That asymmetry drives everything else. If a skill is code, you can run it in the environment it targets and observe whether it worked, and only then write it to the library. If a skill is a paragraph of advice, there is no equivalent check — the only test is trying it later, on a real task, where failure is expensive and diagnosis is hard. This is why the strongest reflection and self-improvement results come from procedures with external verifiers attached: tests, compilers, schema validators, or an environment that reports success. Verified skills accumulate into an asset; unverified prose accumulates into noise. ## Voyager and the growing skill library The Minecraft agent Voyager remains the clearest illustration. It operated in an open world with no fixed task list, generated programs to accomplish goals, executed them, used environment feedback and self-verification to iterate until a program worked, and then stored the successful program in a skill library indexed by a description of what it does. Later tasks retrieved relevant skills and composed them, so the agent's capability compounded instead of resetting each episode. Two properties made that possible and neither survives a prose formulation: the stored artifact was executable, so success was checkable, and it was composable, so a skill could be called from inside another. ## Determinism and drift A stored program executes identically each time. A remembered description has to be re-interpreted by the model on every use, and interpretation drifts — the agent that reads "remember to check the staging environment first" will do something slightly different on each occasion, and there is no way to pin it. For anything with side effects, that variance is the difference between a procedure and a suggestion. ## Review, versioning and safety Because a skill is executable, it is a security surface. A memory system that lets an agent write arbitrary new procedures and invoke them later is a self-modifying execution path, and treating skills as files in source control is what makes that governable: changes are diffs, diffs can be reviewed, a bad skill can be reverted to a known-good version, and the whole library is auditable. None of that is available for a fuzzy body of remembered advice, where you cannot even enumerate what the agent currently believes it should do. ## Token economy and progressive disclosure A large skill library cannot sit in context. The standard packaging is tiered: a short name and description per skill — on the order of tens of tokens — kept available for selection, with the full instructions loaded only when a skill is activated, and any bundled scripts read or executed only if that skill actually needs them. The open Agent Skills format standardises exactly this three-tier disclosure around a skill document plus its bundled files, and providers ship it as a portable way to hand an agent new procedural capability. The design goal is that adding the hundredth skill costs almost nothing until it is used. ## The failure mode of prose procedural memory Systems that store how-to as accumulated natural-language lessons run into a predictable decay. Entries contradict each other as practice changes, with no mechanism to decide which wins. Retrieval is by similarity, so the agent gets an approximately relevant instruction rather than the right one. Nothing is testable, so nothing is ever confidently deleted. And the library's behaviour cannot be reasoned about ahead of time, because it is only realised through the model's interpretation at read time. ## When prose is still right Not everything procedural can be code. Judgment-shaped guidance — how to phrase a refusal, when to escalate to a human, what tone this customer segment expects — has no executable form, and forcing it into code produces brittle rules that miss the point. The defensible split is that deterministic, checkable steps become code, and judgment becomes written guidance that is still versioned, reviewed and scoped like code even though it cannot be tested. Structured skill documents are exactly the hybrid: prose instructions with executable scripts alongside, both living in the same reviewed, versioned artifact.

  • How does a large skill library avoid eating the context window?
    By tiered disclosure. Only a name and a one-line description per skill stays available for selection, typically tens of tokens each; the full instructions load when a skill is activated, and bundled scripts are read or executed only if that skill needs them. The intended property is that adding another skill costs almost nothing until something actually calls it.
  • What is the security concern with an agent writing its own procedural memory?
    It is a self-modifying execution path: a skill written today runs with the agent's privileges tomorrow, possibly triggered by untrusted input. Keeping skills as reviewed files in version control makes changes visible as diffs and revertible, and the same least-privilege limits you would apply to any tool — scoped credentials, sandboxed execution, no unrestricted network egress — apply to skills the agent authored itself.
  • Is there procedural knowledge that genuinely should stay as prose?
    Yes — judgment-shaped guidance with no executable form: how to phrase a refusal, when to escalate to a human, what tone a customer segment expects. Forcing it into code yields brittle rules that miss the intent. Keep it as written guidance that is still versioned, reviewed and scoped like code, which is exactly what a structured skill document combining instructions with bundled scripts gives you.

saying these in an interview costs you the question

  • Says procedural memory means keeping a pile of past prompts
  • Stores how-to as accumulated prose lessons with no way to test them
  • Retrieves a named skill by similarity search instead of by name
  • Ignores that agent-authored skills are executable and need review
  • Claims code form is only about saving tokens

context