Why does a routing supervisor's context balloon by turn 15, and how do you fix it?
answer
- everything flows back through the centre
- cost per decision rises every turn
- long windows also route worse
- define what each specialist returns
- feed a ledger, not the transcript
basics
~20 sEvery specialist reply flows back through the supervisor, so its window accumulates work it never needed to read. Cost per routing decision climbs, routing quality degrades with length, and the run eventually stalls. Fix it with short return contracts, a task ledger, and externalized artifacts.
solid answer
~50 sCentralization is the cause. Because control returns to the supervisor after every specialist, the router's context grows monotonically with the full text of work it does not need - a 311 supervisor can sit at 30k tokens by turn 15 purely from specialist replies. Three things then hurt: every routing decision re-reads the whole window, so cost and latency per decision rise superlinearly across the run; long contexts degrade attention, so the router starts forgetting the goal and re-dispatching agents it already consulted; and the window eventually caps the run. The fixes are all about what crosses the boundary back. Give each specialist a **return contract** - a short structured summary, typically one to two thousand tokens, not its transcript. Replace the raw history in the routing prompt with a **task ledger**: goal, completed steps with one-line outcomes, open questions. Keep large outputs in an artifact store and pass references. Compact periodically, always preserving the goal and the ledger.
go deeper
Know that in a supervisor system every specialist's answer comes back through the supervisor, so its prompt gets longer with each turn and eventually gets expensive.
Explain the mechanics: each routing decision re-reads the whole window, so cost and latency climb across a run, and long contexts degrade the model's ability to follow its own routing instructions.
Show the fix set and the diagnosis. Define return contracts per specialist, replace raw history with a task ledger, externalize large artifacts, and prove the problem with tokens-per-decision and early-versus-late routing accuracy.
Frame it as the cost of centralization you accepted for auditability and coherent termination. Set the budget - what a specialist may return, when compaction fires, when a workload should stop being routed centrally at all.
## Why the window grows A supervisor topology has one property that is simultaneously its strength and its scaling limit: everything passes through the centre. Specialist A answers, the answer lands in the supervisor's context so it can decide what is next; specialist B answers, that lands too. Nothing is ever removed. By turn 15 of a moderately chatty run the router is carrying the accumulated output of a dozen specialist turns, most of which mattered only for the one decision that followed them. A concrete shape: a city 311 supervisor consults potholes, permits, sanitation and housing agents over a complex mixed-use complaint. Each returns a few thousand tokens of detail because nobody defined what it should return. The supervisor's own context passes 30k tokens by turn 15 - and it is still only choosing a name from a list of four. ## Why it hurts more than the token bill **Cost is quadratic-ish across a run.** Each routing decision re-sends the whole accumulated context. Turn 15 costs roughly fifteen times turn 1 for a decision of identical difficulty. Prompt caching softens the repeated prefix but does not remove the growth, since every turn appends new tail content. **Latency follows.** Time to the routing decision grows with input length, and the supervisor sits on the critical path of every hop. **Quality degrades - this is the part candidates miss.** Long contexts suffer what practitioners call context rot: retrieval and instruction-following weaken as the window fills. A degrading router forgets its completion criteria, re-dispatches an agent that already answered, or starts answering the user itself. The failure looks like a bad model when it is a bad context budget. **And then the run stops.** Window exhaustion turns a slow system into a broken one, usually mid-task. ## The return contract The highest-leverage fix is defining, per specialist, what it is allowed to hand back. A good return contract is a small structured object: a one-paragraph outcome, a status, any facts the router genuinely needs for the next decision, and references to anything large. The specialist's internal reasoning, tool calls and intermediate results stay in its own context and die with it. The test is simple: the router needs enough to decide *who acts next and whether we are done*. It does not need the specialist's working notes. If your summaries are being written for the end user rather than for the router, they are the wrong artifact - produce the user-facing text once, at the end, from stored artifacts. ## The task ledger Even with tight returns, raw history accumulates. The stronger move is to stop feeding history at all and feed a maintained ledger instead: the goal verbatim, the completed steps with one-line outcomes, the open questions, and the completion criteria. The supervisor's prompt then grows with the *number of steps*, not with their verbosity, and the goal stays near the top of the window where it survives. This also makes the supervisor closer to stateless per decision, which has a pleasant operational side effect: you can replay a routing decision from its ledger alone when debugging. ## Externalize the bulk Large outputs - a permit history, a full inspection report, a generated document - belong in a store with a handle, not in the conversation. The supervisor routes on the handle plus a description; whichever specialist actually needs the content fetches it. This is the same discipline as passing a file path instead of a file. ## Compaction as the backstop When the window still fills on genuinely long runs, compact: summarize the older middle of the transcript while preserving the goal, the ledger and the most recent turns verbatim. Compaction is a backstop, not a strategy - if you are compacting every few turns, your return contracts are too loose. ## Knowing it is happening Instrument tokens-in per routing decision as a time series across a run, not just totals per run. A rising line with a flat decision difficulty is the signature. Watch alongside it: hops per run, repeat dispatches of the same specialist, and routing accuracy late in a run versus early. If accuracy at turn 12 is materially worse than at turn 2, you are watching context rot, not model variance. ## The structural caveat All of this mitigates a property you chose deliberately. A single router that sees everything is also the thing that makes the system auditable and its termination decision coherent. The trade is real, and the honest senior answer names it as a trade rather than pretending tighter summaries make centralization free.
- Doesn't prompt caching make the growing supervisor context a non-issue?No. Caching cuts the price of re-sending the unchanged prefix, which helps the bill and time-to-first-token, but the context still grows, still gets read by the model, and still degrades instruction-following as it lengthens. Caching is a cost optimization layered on top of a context budget, not a substitute for one.
- How would you tell context bloat apart from a genuinely hard routing problem?Compare routing accuracy early versus late within the same runs. If the router picks correctly at turn 2 and poorly at turn 12 on comparable requests, length is the variable. If accuracy is uniformly poor from turn 1, the specialist descriptions or the roster design are at fault, and shrinking the context will not help.
- What do you lose by summarizing specialist output before it reaches the supervisor?Detail the router might have needed for a later decision, and the ability to reconstruct exactly why a choice was made from the conversation alone. Mitigate by keeping the full specialist output in an artifact store keyed by a handle, so the summary carries a pointer and nothing is truly discarded - only kept out of the hot window.
It is air-traffic control: one controller sees the whole sky, which is exactly why the airspace stays coherent - and exactly why every aircraft's chatter has to fit through one pair of ears.
saying these in an interview costs you the question
- Treats growing supervisor context as only a cost problem
- Assumes prompt caching removes the growth
- Lets specialists return full transcripts to the supervisor
- Blames the model when late-run routing degrades
- Compacts constantly instead of tightening return contracts