skip to content

Effective vs Advertised Context

A million-token window does not mean a million usable tokens: retrieval accuracy and instruction-following degrade long before the hard limit. Interviewers ask how you would measure that ceiling for your own workload rather than trusting the spec sheet.

on this pageshow

questions

4

Why reserve output tokens when budgeting an LLM context window?

level: juniorimportance: must knowfreq 62%

answer

  1. the answer needs room too
  2. input is not the whole budget
  3. reasoning tokens spend the same window
  4. subtract the largest expected response
  5. truncated mid-sentence is the symptom

basics

~20 s

A context window is shared by the prompt and everything the model generates. Fill it with input and the answer has nowhere to go: responses get truncated mid-sentence or the request is rejected. Reserve output space first.

solid answer

~50 s

On essentially every deployed model the advertised context length covers the whole exchange — system prompt, tool definitions, conversation history, retrieved documents, any reasoning the model emits, and the final answer. So the budget is not "how much can I stuff in", it is "how much is left after I subtract the largest response this task can legitimately produce". A concrete failure I have seen repeatedly: a contract-review agent packs 190k of a 200k window with clauses and precedent, then is asked for a 12k-token structured diff — and either the provider rejects the request or the JSON stops halfway through, which downstream parsing reports as a malformed model rather than a budgeting error. The fix is arithmetic done up front: window minus fixed overhead minus maximum output minus a reasoning allowance minus a small safety margin gives the ceiling for history and retrieved context, and the retrieval layer is capped to that number.

code

python · 9 lines
python
WINDOW = 200_000
SYSTEM = 1_800        # instructions
TOOLS = 6_400         # tool names, descriptions, schemas
MAX_OUTPUT = 12_000   # largest structured report this task can produce
REASONING = 8_000     # tokens the model may emit before answering
SAFETY = 2_000        # tokenizer drift and rounding

allowance = WINDOW - (SYSTEM + TOOLS + MAX_OUTPUT + REASONING + SAFETY)
print(allowance)  # 169800 tokens left for history + retrieved documents

go deeper

for a junior

Be ready to say plainly that the window covers prompt and response together, and that you subtract the expected answer size before deciding how much context to include.

for a middle

Explain the full budget breakdown — system prompt, tool schemas, history, retrieved context, reasoning, answer — and show the subtraction that yields the retrieval allowance, including a safety margin for tokenizer differences.

for a senior

Demonstrate that you can diagnose this from production signals: stop reasons, token counts per request, and failures that correlate with input size. Show how you bound the elastic parts so oversized requests degrade evidence, never the answer.

for a principal

Own the design tradeoff: which tasks get big outputs at all, when to split a job into plan-plus-sections instead of one giant call, and what per-request token envelope each product surface is allowed so budgets stay predictable across teams.

## The window is a shared ledger A context window is a single fixed-size ledger, not an inbox with a separate outbox. Every token the model reads and every token it writes in a turn is drawn from the same budget. That means the number in the spec sheet — 200k, 1M, 2M — is the total for the turn, and the useful question is never "does my prompt fit" but "does my prompt plus the answer I am asking for fit". Providers typically expose a second, smaller cap on how many tokens a single response may contain. That cap is a ceiling on generation, not a reservation: setting it does not carve space out of the window for you. If the prompt already consumes nearly the whole window, a generous output cap changes nothing — generation stops when the window is exhausted. ## What competes for the space Before any retrieved content arrives, several fixed costs are already charged: - **System prompt and instructions.** Usually stable, usually a few thousand tokens, and easy to forget because it is invisible in the application code. - **Tool and function definitions.** Each tool's name, description and input schema is text in the prompt. A handful of richly-documented tool servers can consume tens of thousands of tokens before the user has typed anything. - **Conversation history.** Grows monotonically unless something evicts it, and each tool call adds both the call and its result. - **Retrieved documents.** The elastic part, and therefore the part that should absorb the squeeze. - **Reasoning or thinking tokens.** On models that emit an internal reasoning pass before the answer, those tokens are generated and charged like any other output. A long reasoning pass can be several times the size of the visible answer. - **The answer itself.** The one item that must never be squeezed, because a truncated answer is worse than a short prompt. ## The failure signature in production Output starvation does not announce itself as "you ran out of context". It looks like: JSON that stops mid-key and fails schema validation; a summary that ends in the middle of a sentence; a code edit that emits half a function; a stop reason indicating the maximum length was reached rather than a natural end. Teams frequently misdiagnose this as model quality — "it keeps producing invalid JSON" — and reach for a stricter prompt or a bigger model when the real cause is that the request left 900 tokens of room for a 4,000-token object. A second signature is intermittent failure: the same pipeline works on short documents and fails on long ones, because input size is variable and only the long tail crosses the line. Anything that varies with input size — retrieved chunk count, history length, tool output — should be treated as a variable in the budget equation, never as a constant that "usually fits". ## Reserving output first The reliable pattern is to compute the input allowance as a subtraction, once, at the top of request construction: 1. Start from the model's window. 2. Subtract measured fixed overhead: system prompt plus tool definitions. 3. Subtract the **maximum** output the task can legitimately produce — not the average. If the structured report can be 12k tokens for a complex contract, reserve 12k, not the 3k median. 4. Subtract a reasoning allowance if the model emits one. 5. Subtract a safety margin for tokenizer drift between your counter and the provider's. 6. Whatever remains is the ceiling for history plus retrieved context, and the retriever and history trimmer are both bounded by it. The margin matters because token counts are model-specific. Counting with a different tokenizer than the serving model uses gives an estimate, not a guarantee, and the error is usually a few percent — enough to matter only when you have budgeted to the last token, which is exactly why you should not. ## Designing so the answer stays small Budgeting is also a design lever on the output side. Asking for a patch rather than a rewritten file, for identifiers rather than quoted passages, or for a decision plus a pointer rather than a full narrative, can shrink the reservation by an order of magnitude and free that space for evidence. When a task genuinely needs a very large output, the answer is usually to split it: produce a plan or an outline in one call and each section in its own call, each with its own comfortable reservation. ## A working rule Treat the response as the first line item in the budget and the retrieved context as the last. If something must be dropped when a request is oversized, drop evidence and record that you did — never let the answer be the thing that gets cut, because a silently truncated answer is a correctness failure that looks like a model failure.

  • Your structured outputs keep failing schema validation on long inputs but pass on short ones. Where do you look first?
    At the stop reason and the token arithmetic, not the prompt. Intermittent truncation that correlates with input size is the classic signature of output starvation: the prompt grew with the document, the response ran out of window, and generation stopped mid-object. Log prompt tokens, output tokens and stop reason per request, then confirm whether the failures cluster near the window limit before touching the schema or the instructions.
  • Does raising the provider's maximum-output-tokens setting create room for a longer answer?
    No. That setting caps generation; it does not reserve space. If the prompt already fills the window, raising the cap changes nothing — the model still stops when the window is exhausted. Room is created only by making the prompt smaller: trimming history, retrieving fewer chunks, or shrinking tool definitions.
  • How do you handle a task whose output genuinely cannot fit alongside the evidence it needs?
    Split it. Produce a plan, outline or index in one call, then generate each section in a separate call that carries only the evidence that section needs. Each call then has a comfortable reservation. Asking for deltas rather than full rewrites, or references rather than quoted text, is the other lever — it shrinks the reservation without losing information.

saying these in an interview costs you the question

  • Thinks the context limit applies to the prompt only
  • Sizes the prompt to the advertised maximum and hopes
  • Reads truncated output as the model having nothing more to say
  • Ignores reasoning tokens when computing the budget
  • Assumes the provider reserves output space automatically

context

open as a page

Why is a model's advertised 1M-token window not 1M usable tokens?

level: middleimportance: must knowfreq 78%

basics

~20 s

Advertised length is an acceptance limit, not a quality guarantee. Retrieval accuracy and instruction-following degrade as the window fills — the effect commonly called context rot — so the usable ceiling is typically a fraction of the maximum, and it must be measured.

open as a page

How would you measure the usable context ceiling for your own workload?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Hold one graded task set fixed and re-run it at several fill levels — for example 20%, 50% and 80% of the window — padding with realistic in-domain material. Plot score against utilization, then budget below where the curve breaks.

open as a page

What does a passing needle-in-a-haystack score fail to prove about long context?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Needle-in-a-haystack plants one distinctive sentence and asks for it back, which is close to a keyword lookup. A perfect score says nothing about tracking many facts, reconciling contradictions, or reasoning across a long document — the things real workloads need.

open as a page