skip to content

In RAG chunking, why can a 1000-character chunk overflow your token budget?

level: middleimportance: should knowfreq 50%

answer

  1. characters are a proxy, tokens are the constraint
  2. the ratio moves with language and content
  3. English prose around four characters per token
  4. CJK text closer to one token per character
  5. truncation at the embedding limit is silent

basics

~20 s

Characters and tokens are not proportional. English prose averages roughly four characters per token, but code, dense punctuation and Japanese or Chinese text run far closer to one, so the same character window can produce three or four times as many tokens.

solid answer

~50 s

Character-based chunk sizes are a proxy for what actually matters — the token count the embedding model and the reader see. The conversion ratio is language- and content-dependent. With a typical byte-pair-encoding tokenizer, ordinary English prose sits around four characters per token, so 1,000 characters is roughly 250 tokens. The same 1,000 characters of Japanese product-manual text can tokenize to several hundred or a thousand tokens, because the tokenizer has far fewer multi-character merges for that script. Code and heavily punctuated text sit in between. So a chunk size that is comfortably inside an embedding model's input limit for English can overflow it on a multilingual corpus, and most embedding endpoints truncate silently rather than erroring — you lose the tail of the chunk and never see a failure. The fix is to measure length with the tokenizer the model actually uses, not with a character count.

go deeper

for a junior

Know that tokens are not characters and that the ratio depends on the text — roughly four characters per token for English prose, far fewer for code and Japanese.

for a middle

Explain why the ratio moves with script and content, and know that embedding endpoints usually truncate over-length input silently rather than failing.

for a senior

Describe how you would catch this in production: token-counting the finished chunks with the embedding model's own tokenizer during indexing, and failing the job rather than shipping silently truncated vectors.

for a principal

Treat multilingual token-density variance as a capacity and quality planning input — it changes index size, embedding spend and per-market recall — and require the pipeline to express every size constraint in the unit the constraint actually uses.

## Characters are not tokens Every model in a retrieval pipeline — the embedding model that indexes chunks, and the model that reads the retrieved ones — has a limit expressed in tokens. Tokens are the units a subword tokenizer produces, typically via byte-pair encoding: common words and word fragments become single tokens, rare sequences get broken into several. A character-based chunk size is therefore a proxy measurement. It is a convenient one, because counting characters is free and counting tokens costs a tokenizer call, but the conversion factor is not a constant. It depends on what is in the text. ## How the ratio moves Rough figures for a typical BPE tokenizer trained mostly on English web text: - **English prose** — around 4 characters per token. The tokenizer has learned merges for common English words, so most words are one token. - **Source code** — around 2.5–3.5 characters per token. Identifiers in camelCase or snake_case get split, and punctuation-dense lines produce many single-character tokens. - **Japanese, Chinese, Korean** — often close to 1 character per token, and sometimes worse than one token per character, because these scripts are encoded as multi-byte UTF-8 sequences that a mostly-English vocabulary has few merges for. The practical consequence: 1,000 characters of an English policy document is about 250 tokens, while 1,000 characters of a Japanese product manual can be 700–1,000. A single character window applied across a multilingual corpus produces chunks whose real sizes differ by three to four times, which means your "512-token" chunking is nothing of the sort on half the corpus. ## Three tokenizers in one pipeline The subtler version of the problem is mismatch. A RAG pipeline commonly involves three different tokenizations of the same text: 1. The splitter's length function, if it counts tokens at all. 2. The embedding model's own tokenizer, which enforces its input limit. 3. The reader model's tokenizer, which determines whether k chunks fit in the context window. These are frequently different vocabularies from different model families. A chunk measured at 500 tokens by the splitter's tokenizer might be 560 by the embedding model's. If the embedding model's limit is 512, you have just silently lost the tail of that chunk. ## Silent truncation is the dangerous part Most embedding APIs and libraries truncate over-length input rather than rejecting it. Nothing in your indexing job fails. The vector is produced, stored and served. It simply represents the first part of the chunk, and the truncated tail — which may be the exact sentence a user will one day search for — is unretrievable, permanently, until you re-index. Because there is no error, this class of bug is usually discovered months later during a recall investigation, and only if someone thinks to compare stored chunk text against what the model actually saw. The same silence applies at read time, though it is more visible: if k chunks plus the prompt exceed the reader's context, either the framework drops chunks or the provider errors. Either way the arithmetic you did in characters was the wrong arithmetic. ## Practical rules - **Measure in the units the constraint is expressed in.** If the embedding model's limit is in tokens, make the splitter's length function count tokens with that model's tokenizer. - **If you must count characters, calibrate per corpus and per language.** Sample a few hundred documents, compute the actual characters-per-token ratio, and set the character window from the token target divided by that ratio — separately for each language in the corpus. - **Keep a safety margin.** Even with token counting, overlap handling, appended metadata prefixes and joining separators add tokens after the size check. A 10% headroom under the hard limit is cheap insurance. - **Assert, do not assume.** During indexing, count the tokens of each finished chunk with the embedding model's tokenizer and log or fail on any chunk over the limit. That single check converts an invisible data-loss bug into a build failure. ## Why interviewers ask this It separates people who have configured a chunker from people who have debugged one. The character-vs-token gap is the most common cause of "retrieval works fine in English and badly in our other markets", and the answer requires knowing both that tokenizers are content-dependent and that truncation is usually silent.

  • Your splitter counts tokens with one tokenizer and the embedding model uses another. What can go wrong?
    Chunks that measure under the limit by your count can exceed the model's real count, and the endpoint truncates them without erroring. You store a vector representing only the head of the chunk while the full text sits in your document store, so retrieval quietly misses anything in the tail. Use the embedding model's own tokenizer for the length function, and assert on finished chunks.
  • How would you set a character-based window for a corpus that is half English and half Japanese?
    Do not use one window. Measure the characters-per-token ratio separately on a sample of each language, then derive a per-language character window from the shared token target. Better still, switch the length function to token counting so the target is expressed directly in the unit the model constrains, and the language question disappears.
  • Why does source code tokenize less efficiently than prose?
    Subword vocabularies are dominated by natural-language merges. Code splits identifiers at case and underscore boundaries, and lines carry dense punctuation — brackets, dots, operators — that often becomes one token each. Indentation whitespace adds more. The result is roughly 2.5 to 3.5 characters per token instead of four, so a character window sized for prose produces noticeably larger chunks in tokens.

saying these in an interview costs you the question

  • Assumes a fixed four-characters-per-token ratio everywhere
  • Thinks over-length input always raises an error
  • Uses a character count to enforce a token-expressed model limit
  • Ignores that splitter and embedding tokenizers may differ
  • Believes token counts are the same across model families

context