skip to content

How do you set a per-request context budget instead of pasting every available document?

level: principalimportance: should knowfreq 42%

answer

  1. context is not free space
  2. billed on every call, forever
  3. processing the prompt precedes the first token
  4. accuracy is a hump, not a ramp
  5. ablate each block against the eval

basics

~20 s

Derive the budget from measured effective context length rather than the advertised limit, then require each additional block of context to earn its place by improving a measured outcome — because input tokens are billed on every call and prefill time grows with prompt size.

solid answer

~50 s

Start from measurement, not from the spec sheet: run a length sweep for each of your question categories and find where accuracy breaks. That break point, minus a margin, is the ceiling. Then treat everything below it as an economic decision rather than free space. Input tokens are charged on every single call, so a protocol assistant that grows its prompt from 20K to 180K tokens multiplies its per-question input bill roughly ninefold across every question it will ever answer, and the time spent processing the prompt before the first token appears grows with it too. Against that, extra context has diminishing and eventually negative returns as confusable material accumulates. The discipline is an ablation: measure accuracy with and without each context block, keep what pays, and enforce the budget in the assembly code with per-section caps so no single retrieved document can crowd out the rest.

code

python · 9 lines
python
def input_cost(tokens: int, price_per_mtok: float) -> float:
    return tokens / 1_000_000 * price_per_mtok

PRICE = 3.0  # placeholder unit price per million input tokens
DAILY_QUESTIONS = 5_000

for prompt_tokens in (20_000, 60_000, 180_000):
    per_call = input_cost(prompt_tokens, PRICE)
    print(prompt_tokens, round(per_call, 4), round(per_call * DAILY_QUESTIONS, 2))

go deeper

for a junior

Know that every token you put in a prompt is paid for on every request and makes the response slower to start, so context should be chosen rather than pasted wholesale.

for a middle

Explain the three costs — accuracy past the effective length, per-call input spend, and time spent processing the prompt before output begins — and describe measuring each context block's contribution before keeping it.

for a senior

Show the procedure: sweep to find the break point per question category, allocate the budget by section with caps, ablate optional blocks against the eval, and enforce truncation in the assembly code with logging.

for a principal

Own the explicit tradeoff and its expiry date. Make the paste-everything versus retrieve decision with accuracy, unit cost and latency numbers attached, note when a stable cacheable prefix changes the arithmetic, and re-open the decision as volume or the model changes.

## Why "just paste everything" is a real proposal now With million-token windows widely available, the simplest architecture is genuinely tempting: skip retrieval, skip chunking, put the whole corpus in the prompt and let the model sort it out. It removes an entire subsystem and its failure modes. Any principal engineer should be able to say precisely why it is usually still wrong, without falling back on "long prompts are bad". There are three costs, and they are different in kind. ## Cost one: accuracy The window is a capacity limit, not a quality guarantee. Effective context length — the point where accuracy degrades — sits well below the advertised limit and differs by task shape, with retrieval holding out furthest, multi-hop breaking earlier and whole-document aggregation earliest. Adding context past that point does not merely stop helping; it hurts, because additional material brings additional confusable near-misses that compete with the correct evidence. So the accuracy-optimal prompt size is a hump, not a ramp, and the top of the hump is an empirical number you must measure per task. ## Cost two: money, per call, forever Input tokens are billed on every request. This is the point most often underestimated, because it is recurring rather than one-off. A clinical-trial protocol assistant that answers eligibility questions can either retrieve the two or three relevant sections — call it 20K tokens — or paste the full protocol plus amendments at 180K. The 180K design costs roughly nine times as much *per question*, on every question, for the entire life of the product. At a few hundred questions a day the difference is a line item; at enterprise volume it dominates the budget. The comparison worth putting in a design document is not "is 180K affordable" but "is nine times the unit cost buying nine times the value" — and the answer is almost always no, because the marginal 160K tokens contain material irrelevant to the specific question. One real mitigation deserves a mention: when a large prefix is stable across many requests, providers commonly expose a prefix-caching mechanism that makes repeated processing of that prefix substantially cheaper. That changes the arithmetic for a fixed corpus with many queries against it, and it is the strongest argument for a large static context. It does not help when the large part of the prompt varies per request, which is the common case for retrieved evidence. ## Cost three: latency The model must process the whole prompt before it produces anything. That processing time grows with prompt length, so a nine-fold larger prompt shows up directly as a longer wait before the first word appears. For an interactive assistant this is often the constraint that decides the design, ahead of cost — users tolerate a somewhat worse answer far better than they tolerate a long silence. ## Setting the number A workable procedure: 1. **Measure the ceiling.** Sweep input length per question category and find the break point for each. Take the lowest category you must serve in one prompt, subtract a margin, and that is the hard ceiling. 2. **Allocate by section.** Divide the budget into named allotments — system instructions, task description, retrieved evidence, conversation history, output scaffolding — with a cap on each. Caps prevent the common failure where one enormous retrieved document silently evicts the conversation history. 3. **Ablate.** For each optional block, run the eval with and without it. Blocks that do not move a measured outcome are removed. This is the step teams skip, and it is where most of the savings are: prompts accumulate context by accretion, and nobody ever checks whether the paragraph added six months ago still earns its tokens. 4. **Enforce in code.** The budget must live in the prompt-assembly layer with explicit truncation and a logged warning when a section is trimmed, not in a design document. Anything unenforced drifts upward. 5. **Re-measure on model change.** Effective length, cost per token and prefill behaviour all change with the model. The budget is a per-model constant, not a permanent one. ## The honest exceptions Sometimes pasting everything is right. When the corpus is genuinely small, when the question requires whole-document coverage that retrieval would fragment, when the prefix is stable and cacheable across many queries, or when engineering time is the scarcer resource than inference spend during a prototype — a big prompt is the correct call. The principal's job is not to ban large contexts; it is to make the choice explicit, with the accuracy, cost and latency numbers attached, and to revisit it when volume grows. A decision that was right at a hundred requests a day is often wrong at a hundred thousand.

  • When is pasting a whole corpus into every prompt actually the right call?
    When the corpus is genuinely small, when the question needs whole-document coverage that chunked retrieval would fragment, when the large prefix is stable across many queries and can be cached, or during a prototype where engineering time costs more than inference. The point is not that large prompts are forbidden — it is that the choice should carry explicit accuracy, cost and latency numbers, and should be revisited when request volume grows by an order of magnitude.
  • How do you stop a context budget from quietly drifting upward over time?
    Enforce it in the prompt-assembly code, not in a document. Give each section a hard cap, truncate deterministically when a cap is exceeded, and emit a logged warning so trimming is visible rather than silent. Add a test that fails when the assembled prompt for a fixture case exceeds the budget. Then re-run the ablation periodically, because prompts grow by accretion and nobody removes the paragraph they added last quarter.
  • Which do you optimise first when a long-context assistant is both slow and expensive?
    Usually latency, because it decides whether the product is usable at all and the fix overlaps with the cost fix. Shrinking the prompt reduces both the processing done before the first token and the per-call bill, so a single ablation pass moves both metrics. Cost-only optimisations that keep the prompt large — cheaper models, caching, batching — are worth doing afterwards, but they do not rescue an interactive experience that makes the user wait.

saying these in an interview costs you the question

  • Treating the advertised window as the budget
  • Assuming a large prompt is a one-off cost rather than per call
  • Believing more context monotonically improves accuracy
  • Setting a budget in a document but not enforcing it in code
  • Ignoring that prompt size delays the first visible token

context