Trimming LLM chat history by character count overflows the window on Japanese threads and wastes 30% on English ones — how do you rebuild the budget?
answer
- opposite failures mean wrong units
- characters are not a stable proxy
- reserve the output side first
- count the rendered request
- re-count after every removal
basics
~10 sBudget in tokens, not characters. Count the fully templated request with the deployed model's tokenizer, subtract reserved output tokens and a safety margin, then drop or summarize whole turns and re-count after each removal.
solid answer
~50 sThe bug is a unit mismatch: characters are a proxy whose conversion rate varies by several times across languages and content types, so one constant is simultaneously too generous for Japanese and too strict for English. Rebuild it as a token budget. Take the model's context limit, subtract the tokens you reserve for the response — including hidden reasoning tokens where a depth dial is enabled — and subtract a safety margin, and treat the remainder as the input allowance. Measure against the *rendered* request, so the system prompt, tool schemas and per-message template overhead are inside the number. Then trim whole turns rather than slicing mid-message, re-count after each removal, and summarize dropped history instead of discarding it silently. Instrument it: log estimated versus reported usage, alert on near-limit requests per locale, and exercise the worst-case locale in CI so the regression is caught before a user hits it.
code
python · 14 linesMAX_WINDOW = 200_000
RESERVED_OUTPUT = 8_000
SAFETY_MARGIN = 500
def fit_history(messages, count_tokens):
"""count_tokens renders the chat template, then tokenizes."""
allowance = MAX_WINDOW - RESERVED_OUTPUT - SAFETY_MARGIN
kept = list(messages)
while len(kept) > 1 and count_tokens(kept) > allowance:
del kept[1] # drop the oldest turn, keep the pinned system message
if count_tokens(kept) > allowance:
raise ValueError("single message exceeds the input allowance")
return keptgo deeper
Understand that context limits are counted in tokens, so a policy written in characters cannot be correct for every language, and the fix is to change the unit rather than the number.
Explain the mechanics: count the templated request with the model's tokenizer, subtract reserved output and a margin, trim whole turns, and re-count after each removal because template overhead disappears with the message.
Show production judgment — summarize rather than silently discard, handle the oversized single message, instrument estimate-versus-actual and per-locale headroom, and put the worst-case locale in CI so the regression cannot ship again.
Own the policy across markets: parity of effective context between locales, where the cost of counting is paid, how the trimming strategy interacts with prompt-prefix stability, and what quality signal tells you the budget is set wrong.
## Diagnosing the symptom Two opposite failures from one policy is the signature of a unit mismatch. A character limit implies a fixed characters-per-token conversion, and that conversion is not fixed: English prose runs near four characters per token, while Japanese approaches one. Set the character limit so English fills the window and Japanese overflows it by a multiple; set it so Japanese fits and English leaves a third of the window — and a third of the available history — unused. There is no constant that satisfies both, which is why tuning the number is not the fix. A second, smaller contributor hides in the same policy. Even summing message strings perfectly, the count omits the chat template: per-message role and turn-delimiter tokens, the system prompt, and tool schemas re-sent every call. On a long thread of short turns that overhead alone can be a meaningful share of the request. ## The budget, stated properly Write it as an equation in tokens: `input_allowance = context_limit - reserved_output - safety_margin` `reserved_output` is the maximum response you will permit, and on models exposing a reasoning-depth dial it must also cover hidden reasoning tokens, which are generated and billed as output and consume the same window. `safety_margin` absorbs estimator error and anything the provider injects that you did not author. Anything left is what the rendered input may occupy. Then enforce it against the rendered request. Apply the model's chat template to the message list, tokenize the result with that model's tokenizer — offline where one is published, otherwise via a counting endpoint — and compare to the allowance. Now the Japanese thread and the English thread are measured in the same unit, and both use the window fully. ## Trimming, done carefully Trim at message boundaries, never mid-message and certainly never mid-token; a half-sentence at the head of the context is noise the model must reconcile. Keep the system prompt pinned. Drop oldest-first, and re-count after each removal rather than assuming the saving equals the message's own tokens, because removing a turn also removes its template overhead and any subsequent summarization changes the total again. Prefer summarizing dropped history over discarding it. A compact recap of the removed turns preserves commitments and decisions at a fraction of the tokens, and the summarization call itself needs a budget so it does not become an unbounded second cost. Handle the pathological case explicitly: if a single message exceeds the allowance on its own, chunk or summarize that message rather than letting the trim loop spin. One subtlety worth naming: trimming from the front changes the beginning of the prompt on every turn, which defeats reuse of a stable prefix downstream. Where that matters, prefer strategies that keep the earliest portion fixed — a pinned system block plus a summary slot that changes rarely — so churn is concentrated at the tail. ## Instrumentation, because this failed silently once already The original bug shipped because nothing measured it. Fix that alongside the algorithm. Log the pre-flight estimate next to the usage the response reports and alert when the residual drifts, which is what catches a new tool, an expanded system prompt or a model swap. Emit headroom as a metric segmented by locale, so a market approaching the limit is visible before it overflows. Track the trim rate — how often history is discarded — because a rise means users in some locale are losing context and quality even though no error is thrown. And put the worst-case locale into CI: a fixture thread of dense non-Latin text plus the largest toolset, asserted to fit within the allowance. ## The judgment being tested The interviewer wants three things. That you identify the unit mismatch rather than proposing a larger character limit. That you know the count must be taken over the templated request with the deployed model's tokenizer, not over concatenated strings with a remembered ratio. And that you treat the 30% waste as a real defect, not just the safe side of the bug — unused window is context the assistant could have used, so the English path is quietly delivering worse answers than it should. A strong answer also refuses a false economy. Counting tokens costs something: an offline tokenizer costs CPU, a counting endpoint costs a round trip. Both are cheap relative to a rejected request or a truncated conversation, and where the hot path cannot afford it, the fallback is a calibrated per-locale ratio at a pessimistic percentile plus a larger safety margin — an approximation you monitor, not a constant you trust.
- Why is the 30% unused window on English threads a defect rather than a safe margin?Because unused context is discarded capability. Those tokens could have held more conversation history or more retrieved evidence, so the English path answers with less grounding than the model could have used, and the quality gap is invisible because nothing errors. A budget that is too tight and one that is too loose are the same bug seen from two sides, and both are fixed by measuring in the right unit.
- What must the reserved-output portion of the budget account for beyond the visible answer?On models with a reasoning-depth or effort dial, the hidden reasoning tokens are generated before the answer, billed as output and drawn from the same window, so a high effort setting can consume far more than the visible reply suggests. The reservation must cover reasoning plus answer plus any structured-output scaffolding, otherwise a request that passed the input check still fails or truncates when the response runs long.
- Counting tokens on the hot path is too expensive for your latency budget. What is the fallback?Calibrate offline per locale and content type from sampled real traffic, use a pessimistic percentile rather than the mean, add the fixed per-message template overhead, and widen the safety margin to absorb the estimator error. Then keep logging the estimate against reported usage and alert on drift. It is an approximation you monitor, not a constant you trust, and it must be re-derived whenever the model changes.
- How does front-trimming interact with keeping a stable prompt prefix?Dropping the oldest turns changes the start of the rendered prompt on every request, so nothing downstream can rely on the prefix being unchanged between calls. If prefix stability matters to you, restructure so churn sits at the tail: pin the system block and tool definitions, keep a summary slot that is rewritten rarely rather than every turn, and accept slightly less aggressive trimming in exchange for a prefix that stays put.
saying these in an interview costs you the question
- Proposes tuning the character limit instead of changing units
- Sums message strings and ignores chat-template overhead
- Forgets to reserve tokens for the response
- Truncates mid-message or mid-token to hit the limit
- Treats the unused 30% as acceptable slack, not a defect