skip to content

Why package an agent role as a skill file that loads only when the role activates?

level: middleimportance: should knowfreq 42%

answer

  1. not everything needs to be resident
  2. metadata always, body on demand
  3. a role as a reviewable artifact
  4. the description is the activation trigger

basics

~20 s

Packaging a role as a file — short metadata always visible, full instructions loaded only on activation — keeps the baseline context small while making the role portable, versionable and reviewable like code instead of buried in one giant system prompt.

solid answer

~50 s

Two problems get solved at once. The first is context: if every role's full instructions sit in the system prompt, ten roles means ten roles' worth of tokens on every single turn, and long contexts degrade quality. **Progressive disclosure** fixes that — only a name and a one-line description stay resident, and the body plus any bundled scripts load when the role is actually selected. The second is portability and governance. A role that lives in a file — as of mid-2026 the common form is a `SKILL.md` with YAML frontmatter plus optional supporting scripts — can be version-controlled, diffed, code-reviewed, tested and shared across teams and across different agent harnesses. A role that lives in a prompt string inside application code cannot. The cost is that the description in the frontmatter becomes load-bearing: if it does not describe when to use the role, the role never activates.

code

markdown · 13 lines
markdown
---
name: exploitability-verifier
description: Use when triaging static-analysis findings to decide whether a reported vulnerability is reachable from untrusted input.
---

# Exploitability verifier

1. Read the reported file and locate the sink.
2. Trace backwards to any entry point that accepts untrusted input.
3. If a path exists, reproduce it inside the sandbox.
4. Return: verdict (exploitable | not_exploitable | unknown), file, line, one-line justification.

Prefer "unknown" over an unsupported claim.

go deeper

for a junior

Know that role instructions can live in a file rather than the prompt, and that only a short description stays loaded until the role is needed. Say why keeping the prompt small matters.

for a middle

Explain progressive disclosure concretely — resident metadata versus deferred body and bundled scripts — and why a long always-resident prompt degrades quality as well as costing tokens.

for a senior

Treat role files as artifacts: versioned, reviewed, evaluated against pinned versions. Be ready to diagnose an activation failure back to a description that names an identity instead of a trigger condition.

for a principal

Own the library — how roles are shared across teams and harnesses, how overlapping activation descriptions are prevented, and how a role change becomes a measurable, reviewable event rather than an untracked prompt edit.

## The problem with roles in the system prompt The naive way to define agent roles is to write them into one system prompt: a paragraph for the reviewer, a paragraph for the writer, a paragraph for the researcher, plus their rules and examples. This works for two or three thin roles and stops working past that. It fails on three axes. **Context cost**: every role's full instructions are resident on every turn, including the turns where nine of ten roles are irrelevant. **Quality**: long contexts degrade model attention — the standard term for this is context rot — so paying that tax buys you worse behaviour, not just a bigger bill. **Governance**: a prompt string embedded in application code is not reviewable as an artifact. Nobody diffs it, nobody tests it, and it cannot travel to another team or another harness. ## Progressive disclosure as the fix Progressive disclosure means putting a small amount of information in front of the model always, and the rest only when it becomes relevant. Applied to roles, the resident part is metadata — a name and a description of *when this role applies*. The deferred part is everything else: the detailed procedure, the examples, the output contract, the edge cases, and any scripts or reference files bundled alongside. As of mid-2026 the widely adopted concrete form is a skill file: a `SKILL.md` document with YAML frontmatter carrying the name and description, a Markdown body carrying the instructions, and optional adjacent files the role can read or execute. The format is portable — the same directory works across several coding-agent harnesses — which is precisely the point: a role stops being a property of one application and becomes an artifact. ## What this buys you **Baseline context stays flat as the role library grows.** Adding a thirtieth role costs one line of metadata, not another page of instructions. This is the difference between a role library that scales and one that quietly poisons every request. **Roles become reviewable.** Because the role is a file in a repository, it goes through the same pipeline as code: pull request, diff, review comment, revert. A regression in agent behaviour becomes traceable to a commit rather than to "someone edited the prompt". **Roles become testable and reusable.** You can pin a version, run an eval suite against a role in isolation, and share it with another team without shipping your application. A bundled script inside the role package is often the sharpest upgrade: instead of describing a deterministic procedure in prose and hoping the model reproduces it, the role carries the script and instructs the agent to run it. **Roles compose with tool scoping.** The skill file describes behaviour; the allowlist grants capability. Keeping them as separate layers means a role's instructions can be shared while each deployment applies its own permission scope. ## The costs and the failure modes The description in the metadata is now load-bearing in a way people underestimate. It is the *only* thing the model sees when deciding whether to activate the role, so it must describe the triggering situation, not the role's identity. "Security expert" is a bad description; "use when triaging static-analysis findings to decide whether a reported vulnerability is reachable" is a good one. A brilliant role body with a vague description never fires, and the symptom is silence — the agent simply does its own thing and nobody sees an error. Second failure mode: descriptions that overlap. Two roles whose trigger conditions cannot be distinguished from their descriptions produce nondeterministic activation, and debugging that is unpleasant because the choice happens before any of the detailed instructions are visible. Third: over-decomposition. Splitting a coherent procedure across five skill files means the agent must activate several to do one job, and the ordering between them is unspecified. A role should be one job. ## When not to bother If you have two roles and they are three sentences each, put them in the prompt and move on. Progressive disclosure is a scaling technique; the machinery costs something to maintain and pays off when the library is large, when several teams share roles, or when the role instructions are long enough to hurt if resident. ## Practical guidance Write the description as an activation condition. Keep the body focused on procedure and output contract rather than personality. Bundle a script for anything deterministic. Version the file and run your evals against pinned versions, so a role change is a reviewable event with measurable consequences rather than an untracked prompt edit.

  • What is the most common reason a well-written role file never gets used?
    Its description does not describe a triggering situation. The metadata line is the only thing visible at selection time, so a description like "security expert" gives the model nothing to match against. Rewriting it as a condition — "use when triaging static-analysis findings to judge reachability" — usually fixes activation immediately. The failure is silent, which is why it lingers.
  • Why bundle a script alongside a role's instructions instead of describing the procedure in prose?
    Because deterministic steps should be executed, not re-derived. Prose describing a fixed procedure is reproduced approximately and differently each run; a script runs the same way every time, costs a handful of tokens to invoke rather than a page to describe, and can be tested independently. Keep prose for judgment and hand the mechanical parts to code.
  • When is putting roles in the system prompt still the right call?
    When there are few of them and they are short. Progressive disclosure is a scaling technique with real maintenance cost — a file layout, activation descriptions, version pinning. With two three-sentence roles, none of that pays. Reach for packaging when the library grows, when instructions are long enough to hurt if always resident, or when several teams need the same role.

saying these in an interview costs you the question

  • Thinking skill packaging changes what the model can do
  • Writing descriptions that name the role instead of its trigger
  • Keeping every role's full instructions resident always
  • Splitting one coherent job across several role files
  • Assuming a role file grants the tools it mentions

context