How would you measure the usable context ceiling for your own workload?
answer
- one variable: how full the window is
- same questions, different fill levels
- pad with real in-corpus documents
- repeat runs, nondeterminism is real
- budget below the knee of the curve
basics
~20 sHold one graded task set fixed and re-run it at several fill levels — for example 20%, 50% and 80% of the window — padding with realistic in-domain material. Plot score against utilization, then budget below where the curve breaks.
solid answer
~50 sThe procedure is a controlled sweep. Freeze a task set large enough to be statistically meaningful — a couple of hundred questions with programmatic grading — and keep the questions, the gold evidence and the grader identical across runs. The only variable is how much additional in-corpus material pads the prompt to a target utilization: 20%, 50%, 80%, and one run near the limit. Repeat each cell several times, because sampling is nondeterministic and a single run at each level cannot distinguish a real cliff from variance. Score not just accuracy but failure type: missed facts, wrong-but-plausible passages chosen, refusals, truncated outputs. The result is a curve with a knee; you set your production budget below the knee, leave the reserved output space on top, and re-run the sweep whenever the model, the prompt skeleton or the retrieval depth changes. Vendor numbers cannot substitute — the knee moves with your task shape.
code
python · 25 linesdef build_padded_prompt(question, gold_docs, filler_docs, window, utilization, count_tokens):
"""Fill a prompt to a target utilization with in-corpus filler, keeping gold evidence."""
target = int(window * utilization)
docs = list(gold_docs)
used = count_tokens(question) + sum(count_tokens(d) for d in docs)
for doc in filler_docs:
n = count_tokens(doc)
if used + n > target:
break
docs.append(doc)
used += n
return question, docs, used
def words(text):
return len(text.split())
q, docs, used = build_padded_prompt(
"Which clause caps liability?",
["clause 7 caps liability at fees paid"],
["neighbouring clause text " * 20] * 40,
window=1000,
utilization=0.5,
count_tokens=words,
)
print(used, len(docs))go deeper
Know that you find the usable limit by testing rather than by reading the spec: run the same questions at different prompt sizes and see where the answers get worse.
Describe the controlled sweep — fixed question set and grader, only fill varies, realistic in-corpus padding, several levels — and explain why the padding material choice changes the result.
Show the full procedure and its judgment calls: repeats for nondeterminism, failure-type buckets rather than one score, the knee-minus-margin cap enforced in the request builder, and production instrumentation of utilization.
Own the policy: which decisions require a fresh sweep, what evidence gates a model migration or a higher cap, and how the cost of running the harness is traded against the risk of silent quality loss across products.
## Why you have to measure it yourself Effective context is not a property of a model alone; it is a property of the model, the task shape and the material you surround the answer with. A single-fact lookup over clean prose survives far more fill than a twelve-fact reconciliation over near-duplicate documents. Published long-context results tell you what shape to expect — a band that holds, then a drop — but not where your knee sits. Since the budget you enforce in code is a specific number, and a wrong number produces either wasted headroom or silent quality loss, you measure. ## The sweep The design is a controlled experiment with exactly one independent variable: utilization. **Fix the task set.** Two hundred or so items is a reasonable target: enough that a few points of difference is not pure noise, small enough to run repeatedly at cost. Draw them from real production traffic where possible, and include the failure modes you actually care about, not only the happy path. Every item needs a programmatically checkable answer — an exact value, a set of clause identifiers, a schema-valid object — because a judge model adds its own variance to an experiment that already has plenty. **Fix everything except fill.** Same questions, same gold evidence present in every condition, same system prompt skeleton, same grader, same sampling settings. If the questions change between levels, the curve is meaningless. **Pad with realistic material.** Fill to the target utilization with documents from the same corpus — plausible neighbours that could compete with the correct answer. Random characters or lorem ipsum understate the difficulty badly, because discriminating against near-misses is a large part of what breaks at high fill. Padding with text from an unrelated domain is nearly as flattering: the model can dismiss it at a glance. **Choose levels that bracket the interesting region.** 20% as a low-fill reference, 50% and 80% through the band, plus one run close to the practical limit once output reservation is accounted for. If the drop lands between two levels, add a level between them rather than interpolating — the curve is cliff-shaped, so interpolation is exactly the wrong instinct. **Repeat each cell.** Generation is nondeterministic even at low temperature in most serving stacks. Three to five runs per level gives you a spread rather than a point, and lets you tell a genuine cliff from an unlucky run. Report the spread, not just the mean. ## Score more than accuracy An aggregate score hides the mechanism. Bucket the failures: - **Missed** — the model did not surface a fact that was present. - **Decoyed** — it surfaced a plausible but wrong passage, typically one of your padding documents. - **Reasoned wrongly** — it found the right evidence and drew the wrong conclusion. - **Truncated or refused** — an output-space or safety problem, not a retrieval problem. Each bucket points at a different remedy. Rising misses at high fill is the classic context-rot signature. Rising decoys says your padding is realistically confusable and your retrieval precision matters more than window size. Truncations mean your reservation arithmetic is wrong, not that the model degraded. ## Turning the curve into a budget Once you have the curve, the budget follows mechanically. Pick the utilization at which the metric you care about is still within your tolerance, subtract a margin for the fact that production inputs will be messier than your padding, and enforce that ceiling in the request builder: cap the number of retrieved chunks, cap history length, cap the size of tool results carried forward. Layer the output and reasoning reservation on top so the ceiling applies to input only. Then instrument production to log actual utilization per request, so you can see whether real traffic is drifting toward the region you decided to avoid. ## Maintenance The curve is not a permanent artifact. Re-run the sweep when: - You change model or model version. A larger advertised window is not permission to raise the cap; it is a reason to re-measure. - You change the prompt skeleton materially — a much larger system prompt or a big new set of tool definitions shifts where fixed cost ends and evidence begins. - You change retrieval depth, chunk size or ranking, since these change how much competing material lands in the window. - Your corpus changes character, for instance after ingesting a much larger or much more repetitive document set. Keep the harness cheap enough to run on that cadence. If a full sweep is expensive, run a reduced version — fewer items, three levels, three repeats — as a regression check, and reserve the full sweep for model migrations. ## What good looks like in an interview The answer that lands is a procedure, not an opinion: what is held fixed, what varies, how the padding is chosen, how many repeats, what is scored, and how the resulting number becomes an enforced cap with headroom. Saying "we measured it" without being able to describe the control is the tell that it was never measured.
- Why not just pad with lorem ipsum or random tokens to reach the target length?Because a large part of high-fill degradation is failing to discriminate against material that looks like the answer. Meaningless filler is trivially dismissed, so the model appears to tolerate far more fill than it will in production, and your measured ceiling comes out optimistically high. Padding should be real documents from the same corpus that plausibly compete with the correct evidence.
- How many repeats per level do you need, and why does it matter?Three to five is a reasonable floor. Generation is nondeterministic in most serving stacks, so a single run per level produces a curve you cannot distinguish from noise — and since the real curve is cliff-shaped, a lucky or unlucky run at one level can invent or hide a knee. Report the spread across repeats, and add repeats rather than levels when the variance is what is ambiguous.
- You ship a model upgrade with a window four times larger. What do you do to the budget?Nothing, until the sweep is re-run. A larger advertised window is a capacity change, not evidence that the usable band scaled with it, and the knee also moves with task shape and prompt skeleton. Keep the existing cap through the migration, re-measure on the same fixed task set, and raise the cap only where the new curve supports it.
- What do you instrument in production so this stays honest after launch?Log prompt tokens as a fraction of the window per request, along with output tokens, stop reason and whatever quality signal you have — thumbs, downstream validation failures, retries. That lets you check whether real traffic is drifting into the band you decided to avoid, and whether failures cluster at high utilization. Without it, high-fill degradation is silent: nothing errors, answers are just quietly worse.
saying these in an interview costs you the question
- Trusts the vendor's effective-context figure for its own workload
- Changes the question set between utilization levels
- Pads with lorem ipsum or random tokens
- Runs each level once and calls the curve measured
- Interpolates between two levels across a cliff