skip to content

A memory file loaded on every turn has grown to 8k stale tokens — how do you fix it?

level: middleimportance: must knowfreq 55%

answer

  1. paid on every turn, every session
  2. stale instructions are followed, not ignored
  3. sort lines by how often needed
  4. procedures become on-demand skills
  5. pointers beat pasted prose

basics

~20 s

Split the file by how often its content is actually needed. Keep only always-true conventions in the always-loaded file, move task-specific procedures into on-demand skill files that load when the task matches, and delete anything stale outright rather than leaving it in place.

solid answer

~50 s

Treat an always-loaded memory file as a permanent tax: its tokens are paid on every turn of every session, whether or not they are relevant. Stale content is worse than expensive — the model reads the wrong build command as authoritative and follows it, which is a context-clash failure, not a cosmetic one. The fix is triage into three buckets. **Always-relevant conventions** ("dates are ISO 8601", "never push to main") stay in the always-loaded file, ideally a few hundred tokens. **Procedural know-how** used occasionally — "run the quarterly close" — moves into a named skill file whose one-line description is cheap to keep visible and whose body loads only when the task matches; that is procedural memory that costs nothing until used. **Everything untrue or unused** is deleted, not commented out. Then keep the file under review, because it will grow back.

code

markdown · 11 lines
markdown
---
name: quarterly-close
description: Run the quarterly financial close: reconcile ledgers, generate the close pack, file the summary.
---

# Quarterly close

1. Reconcile the ledger export against the bank statement for each entity.
2. Post accrual adjustments listed in the close checklist.
3. Generate the close pack and attach the reconciliation evidence.
4. Circulate for sign-off before filing.

go deeper

for a junior

Know that a project memory file is prepended to every turn, so its size and accuracy matter more than an ordinary document's. Be able to say that out-of-date instructions get followed.

for a middle

Explain the triage: always-true conventions stay, occasional procedures move to on-demand skill files, dead content is deleted. Be ready to argue why staleness is a correctness problem, not just a cost one.

for a senior

Show that you would measure the file's share of the turn budget before editing, and that you would design the split around discovery — a skill nobody loads is as useless as a fact nobody reads. Name context clash and confusion as the concrete failure modes.

for a principal

Own the policy: who reviews the always-loaded tier, what the admission bar is, and how you stop it re-growing across many repositories and teams. Frame the always-loaded budget as a shared resource that needs an explicit ceiling.

## Why an always-loaded file is a special case Most context is situational: a retrieved document appears for one question, a tool result for one step. A project memory file — the `AGENTS.md`/`CLAUDE.md` convention that coding agents read at startup — is different. It is prepended to essentially every turn of every session by every user of that repository. Its cost is therefore multiplied by total usage, and its content is read by the model as standing instruction rather than as evidence. That gives it two distinct failure modes, and a good answer names both. ## Failure one: the budget Every token spends a finite attention budget. An 8,000-token memory file is 8,000 tokens that are unavailable to the actual task and are re-paid on every single turn. The practical goal in context engineering is the smallest high-signal token set, not the most complete one. A memory file drifts the other way by default: nobody is ever punished for adding a line to it, so it accretes. There is a second-order effect too. Large blocks of marginally relevant material make the model harder to steer — the well-known *context confusion* mode, where irrelevant content degrades output quality even when the model does not act on it directly. ## Failure two: staleness is active, not passive The more dangerous problem is that a stale instruction is not ignored. If the file says the build command is `make release` and the project moved to a different toolchain six months ago, the model will confidently emit the old command, because the file is a high-authority instruction and the model has no independent way to know it is out of date. When the file contradicts what the model can observe in the repository, you get *context clash*: two incompatible facts in one window, with the outcome depending on wording and position rather than on truth. This is why "just leave it, tokens are cheap" is the wrong instinct. The token cost is the smaller of the two problems. ## The triage Go through the file line by line and sort each item by **how often it is needed** and **whether it is still true**. **Bucket 1 — always-loaded.** Content that is true on every task and cheap to state: naming conventions, the repository layout in one sentence, hard prohibitions, the tone or format the team expects. These earn their permanent slot because the cost of *not* having them is a wrong answer on an arbitrary turn. Target a few hundred tokens, not thousands. **Bucket 2 — on-demand procedural memory.** Multi-step procedures used occasionally: the quarterly-close runbook, the release checklist, the data-migration dance. In current practice these become skill directories — a `SKILL.md` with short frontmatter (a name and a description) plus a body and any supporting scripts. Only the name and description need to be visible for the agent to know the skill exists; the body enters the context when the task matches. A ten-step close procedure used four times a year is the textbook case: as an always-loaded block it wastes tokens on 99% of turns; as a skill it costs almost nothing until the day it is needed. **Bucket 3 — delete.** Anything untrue, superseded or never referenced. Do not archive it inside the same file behind a heading like "legacy", because the model reads headings as organisation, not as a disclaimer, and a plausible stale instruction remains plausible. If you need the history, that is what version control is for. ## The decision is a policy, not a cleanup The reason interviewers like this question is that the one-off cleanup is easy and the durable answer is not. Ask who owns the file, what the review trigger is (every release? every time a build instruction changes?), and how content earns its way into the always-loaded tier. A useful rule of thumb: a line belongs in the always-loaded file only if a turn that lacks it would plausibly go wrong. Everything else is a skill, a document to be fetched, or deleted. A second practical rule: prefer pointers to prose. "Deployment steps live in `docs/deploy.md`" is a handful of tokens and stays correct as the document changes; pasting the steps inline is hundreds of tokens that go stale silently. ## What a good answer sounds like "I'd measure it first — what fraction of the turn budget is this file, and which lines are still true. Then split by frequency of need: conventions stay, procedures become skills that load on match, dead content gets deleted. And I'd put a review trigger on it, because this file grows back within a quarter if nobody owns it."

  • How do you decide a line belongs in the always-loaded file rather than a skill?
    Apply a needed-on-an-arbitrary-turn test: if a turn that lacks the line would plausibly go wrong, it stays; otherwise it moves. Conventions and hard prohibitions pass because you cannot predict which turn will violate them. A procedure only matters when its task comes up, and its trigger is recognisable from the request, so it can be loaded on match instead.
  • Is there a downside to moving too much into on-demand skills?
    Yes — discovery becomes the failure mode. The agent only loads a skill it knows exists and correctly matches, so a vague or overlapping description means the procedure is silently never used. You also add a step of latency on the turn where it loads. Descriptions have to be written as selection prompts, and overlapping skills need disambiguating just as overlapping tools do.
  • Someone argues the stale lines are harmless because the model can check the repository. Why is that wrong?
    Because the file is framed as authoritative instruction while the repository is evidence the model has to go and gather. With both present you get contradictory content in one window and an outcome that depends on wording and position rather than truth. Empirically the model often emits the memorised command without verifying, which is the exact failure the file was supposed to prevent.

saying these in an interview costs you the question

  • Says tokens are cheap so file size does not matter
  • Assumes the model ignores instructions that conflict with the repository
  • Archives stale content under a legacy heading in the same file
  • Pastes long procedures inline instead of pointing to them
  • Treats the cleanup as one-off with no owner or review trigger

context