How do you pack reranked chunks into a fixed RAG context budget when chunk sizes vary widely?
answer
- the window is not the budget
- subtract instructions, history and the answer
- count with the real tokenizer
- greedy fill, skip what does not fit
- cut only at a structural boundary
basics
~20 sFill in rank order against a token budget, not a character count. Reserve room for instructions, the question and the model's answer, then greedily add chunks and skip any that no longer fit rather than truncating mid-chunk.
solid answer
~50 sTreat it as a small knapsack problem and solve it greedily. First compute the real budget: total window minus system instructions, tool definitions, conversation history, the user question, and reserved output tokens — what is left is the chunk slot. Then walk the reranked list in order, counting each chunk with the actual tokenizer plus its delimiter overhead, and add it if it fits. When one does not fit, skip it and keep trying the smaller lower-ranked chunks rather than stopping at the first miss; with semantic splitting producing chunks anywhere from 90 to 1,400 tokens, stopping early can waste a third of the slot. Guarantee the top-ranked chunk is always included, even if that means dropping everything else. Truncation is a last resort and only at a structural boundary with an explicit marker — never a silent mid-sentence cut, which invites the model to confabulate the missing half.
code
python · 9 linesdef pack(chunks, budget, count_tokens, delimiter_cost=12):
"""chunks: reranked best-first. Greedy skip-and-continue fill."""
packed, used = [], 0
for i, chunk in enumerate(chunks):
cost = count_tokens(chunk["text"]) + delimiter_cost
if i == 0 or used + cost <= budget:
packed.append(chunk)
used += cost
return packedgo deeper
Know that the retrieved text competes for a fixed token budget alongside the instructions, the question and the answer, and that chunks are counted in tokens rather than characters.
Be able to describe the greedy fill in rank order, the delimiter overhead added to each chunk's cost, and why the last chunk that does not fit is skipped rather than silently cut in half.
Demonstrate the operational side: a real slot calculation with a safety margin, a guaranteed top-1 floor, a truncation policy that cuts at structural boundaries with a marker, and the per-request logging that tells you whether the slot is chronically empty or the splitter is producing outliers.
Own the budget as a tuned parameter with cost and latency consequences, not the leftover space in the window. Be ready to argue why an exact knapsack solver is the wrong objective and where chunk-size discipline at ingest removes the packing problem entirely.
## The budget is smaller than the window The first mistake is treating the model's context window as the space available for retrieved text. It is not. The window has to hold the system instructions, any tool or schema definitions, prior conversation turns, the user's question, the formatting scaffolding around the chunks, and — critically — the tokens the model is going to generate. Output shares the window. If you fill the window to the brim with chunks, the generation either truncates or the request is rejected outright. So compute the chunk slot explicitly: `slot = window - fixed_prompt - history - question - reserved_output - formatting_overhead`. Make that arithmetic a single function with a small safety margin, because every one of those terms except the reserved output varies per request. In a travel-policy assistant, for example, a 6,000-token slot is a deliberate figure derived that way, not an arbitrary constant, and it is usually far smaller than the model's nominal maximum — deliberately, because more context is not monotonically better for either answer quality or cost and latency. ## Count tokens, not characters A characters-divided-by-four estimate is fine for a rough dashboard and dangerous for a packer. Code snippets, tables, non-Latin scripts, and long identifiers all tokenize far worse than prose, and an estimate that is 20% low means the request fails at the provider rather than degrading gracefully. Use the real tokenizer for the model you are calling, cache counts on the chunk record at ingest time so packing stays cheap at query time, and add the per-chunk delimiter and header cost to each chunk's price — a dozen tokens each across twenty chunks is a couple of hundred tokens you would otherwise overspend. ## Greedy packing, and the skip-vs-stop decision The formal problem — maximize total relevance subject to a token capacity — is a knapsack, but you should not solve it exactly. Greedy in rank order is nearly always right, because relevance is heavily front-loaded and an optimal solver would happily trade the single best chunk for three mediocre ones that fit better, which is exactly the wrong trade for grounding. The one real choice inside the greedy loop is what to do when a chunk does not fit: - **Stop-at-first-miss** ends packing at the first oversized chunk. Simple, and it preserves a clean "top N contiguous" story, but with semantic splitting producing chunks from 90 to 1,400 tokens, a single 1,400-token chunk arriving with 900 tokens of slot left throws away all the remaining space. - **Skip-and-continue** drops that chunk and keeps testing the rest, filling the slot with smaller lower-ranked material. This is the usual choice. Its cost is that the injected set is no longer a prefix of the ranking, which you must account for when you evaluate and when you debug. Whichever you choose, pin a floor: the rank-1 chunk goes in unconditionally. If it alone exceeds the slot, that is a chunking bug or a budget bug worth surfacing, not something to paper over. ## Truncation policy Sometimes you genuinely must cut. Rules that hold up in production: cut at a structural boundary — a paragraph, a list item, a heading — never mid-sentence. Mark the cut explicitly in the text so the model can see that the document is incomplete rather than treating the fragment as the whole rule. Prefer truncating the tail of a front-loaded document (policy pages usually state the rule first and elaborate after) over one whose payload is at the end, like a table of figures. And prefer dropping a whole low-ranked chunk to mutilating a high-ranked one. ## Order and packing are separate steps Pack first, then order. Packing decides membership under the budget; ordering decides positions inside the block. Interleaving them — deciding placement while you still do not know the final membership — produces packers that are hard to reason about and orderings that silently change when one chunk's size shifts. ## Observability Log, per request: slot size, tokens used, number of chunks offered, number packed, number skipped for size, and whether truncation fired. Those five numbers answer most "why was the answer thin?" questions immediately. A slot that is chronically 40% empty means your chunk sizes and your budget are mismatched; a high skip rate means your splitter is producing outliers that never fit; truncation firing routinely means the budget is too small for the corpus rather than the packer being clever.
- How do you size the retrieved-context slot relative to the model's full window?Subtract the system prompt, tool definitions, conversation history, the user question and a reserved output allowance, then apply a safety margin. Cap what remains at whatever evaluation shows actually helps, rather than the maximum that fits — more context raises cost and latency and does not monotonically raise answer quality, so the slot is a tuned parameter, not a leftover.
- When is truncating a chunk acceptable rather than dropping it?When the chunk is high-ranked, long, and front-loaded — a policy page that states the rule first and elaborates afterwards. Cut at a paragraph or heading boundary, mark the cut explicitly so the model knows the document is incomplete, and never cut mid-sentence. Prefer dropping a whole low-ranked chunk over mutilating a top-ranked one.
- Why not solve the packing problem optimally as a knapsack?Because the objective is wrong. An optimal solver maximizes summed relevance under the capacity, so it will happily swap the single best chunk for several mediocre ones that pack more tightly — the opposite of what grounding needs. Relevance is heavily front-loaded, so greedy in rank order with a guaranteed top-1 floor is both simpler and better aligned.
saying these in an interview costs you the question
- Measures the context budget in characters rather than tokens
- Fills the whole window and leaves no room for the model's output
- Truncates the last chunk mid-sentence with no marker
- Assumes packing more chunks always produces a better answer
- Stops packing at the first oversized chunk and wastes the remaining slot