Why can an embedding model silently ignore half of a 900-token chunk?
answer
- limits are counted in tokens, not words
- the default on overflow is not an error
- the vector represents the prefix only
- jargon and identifiers inflate token counts
- assert tokenized length at ingest
basics
~20 sEvery encoder has a maximum sequence length, and the default behaviour when input exceeds it is to truncate rather than fail. The returned vector represents only the leading tokens; the tail never influences it, and nothing in the response says so.
solid answer
~60 sEmbedding models have a fixed maximum sequence length — many BERT-family sentence encoders cap at 512 tokens, while newer long-context encoders accept 8k or more. When a longer input arrives, the standard tokenizer behaviour is truncation, not an error: the model encodes the first N tokens, pools them, and returns a perfectly normal-looking vector. So a 900-token chunk sent to a 512-token model is indexed as its first half, and every fact in the tail is invisible to retrieval forever. Two things make this hard to catch. First, there is no failure signal — same dimensionality, same score ranges, no exception. Second, people size chunks in words or characters while the limit is in tokens, and domain text inflates that ratio badly: identifiers, product codes, chemical or legal terms and non-Latin script fragment into several subword tokens each, so a chunk that looks comfortably short can be well over the limit. The fix is to measure tokenized length with the model's own tokenizer at ingest, alert or split when it approaches the cap, and choose a long-context encoder if your documents genuinely need it.
go deeper
Know that embedding models have a maximum input length measured in tokens, and that longer text is cut off by default rather than rejected.
Explain that a token is a subword unit, so word counts underestimate length for jargon, identifiers and non-Latin script, and that the returned vector then represents only the leading portion of the text.
Demonstrate the diagnosis: this is a silent failure with no error signal, caught by tokenizing at ingest, tracking the chunk-length distribution, and testing whether a prefix-only embedding matches the whole-document embedding.
Own the design principle that any stage able to discard input silently needs an explicit assertion, and weigh long-context encoders against finer granularity — a bigger window costs latency and dilutes a single vector's precision.
## The mechanism An encoder is compiled around a maximum position count. Positions beyond that have no positional representation, so the model physically cannot consume them. The tokenizer therefore enforces the limit before the model runs, and the near-universal default is **truncate**: keep the first N tokens, drop the rest, proceed. The consequence is stated simply and is worth saying exactly this way in an interview: *the vector represents the prefix, not the document.* Everything after the cut had zero influence on the numbers you stored. ## Why it stays invisible This is a silent-corruption failure, the worst kind in a retrieval pipeline: - The output vector has the normal dimensionality, so no shape check fires. - Similarity scores stay in their usual range, so no threshold looks odd. - No exception, no warning field, no status code — a truncated call and a clean call are indistinguishable in logs. - Retrieval still returns results. They are simply the wrong ones, or missing the right one, for any query whose answer lived in the tail. The user-visible symptom is a system that answers well about the beginnings of documents and inexplicably "doesn't know" things that are demonstrably in the corpus. Teams usually blame the chunking strategy, the reranker or the model's reasoning long before they check token counts. ## Why token counts surprise people Chunkers are usually written against characters or words because those are what a splitter naturally counts. The model's limit is in **tokens**, and the ratio between them is not a constant. Subword tokenizers keep frequent words whole and split rare ones into pieces. For ordinary English prose the ratio is roughly three-quarters of a word per token, so a rule of thumb of ~400 words fitting in 512 tokens feels safe. Then reality intervenes: - **Domain jargon** — pharmaceutical names, legal Latin, medical terms, taxonomies — is rare in the tokenizer's training data and fragments into many pieces. - **Identifiers and codes** — SKUs, order numbers, UUIDs, error codes, file paths — fragment worst of all, sometimes one token per two or three characters. - **Source code** — indentation, punctuation and camelCase identifiers each consume tokens. - **Non-Latin scripts** — a tokenizer whose vocabulary was built mostly on English may spend several tokens per character on Japanese, Korean or Devanagari. - **Tables and markup** — pipes, tags and repeated delimiters inflate counts dramatically. So a chunk that a word-count check calls "400 words, fine" can be 700+ tokens in a catalogue of part numbers, and half of it disappears. Also remember the model may add special tokens of its own, so the budget for your content is slightly under the nominal maximum. ## Detecting it The check is cheap and belongs at ingest, before anything is indexed: 1. Tokenize each chunk with **the same tokenizer the embedding model uses** — not an approximation, not a different model's tokenizer. 2. Compare the length against the model's documented maximum sequence length. 3. Record the distribution of chunk token lengths as a pipeline metric, and alert on the tail. A histogram piling up exactly at the limit is the signature of mass truncation. 4. Re-run this whenever you change chunker, corpus or embedding model — a new model can change both the limit and the tokenizer. A useful sanity test: embed a document, then embed only its second half, and check whether the two vectors are suspiciously unrelated while the first half's vector is nearly identical to the whole document's. That pattern is truncation caught red-handed. ## Fixing it - **Chunk to the model's real limit, measured in tokens**, with headroom for special tokens. - **Use a long-context encoder** if your unit of meaning genuinely does not fit. Encoders accepting 8k tokens and beyond are widely available as of mid-2026. Note the trade-off: a longer window is not free quality — pooling a very long passage into one vector averages away specifics, so a single vector for an entire long document is often *worse* at retrieval than several vectors for its parts, even when the model can technically ingest it. - **Split and aggregate** when a document must stay one logical unit: embed sub-chunks, store them individually, and let retrieval work at that granularity. - **Fail loudly instead of truncating** in your own ingest code — treat an over-length chunk as a pipeline error rather than accepting the tokenizer's silent default. ## The general lesson An embedding pipeline has several steps that degrade quality without producing an error, and truncation is the sharpest of them. Any stage that can silently discard input deserves an explicit assertion at ingest, because the cost of finding it later is measured in a full re-index of the corpus.
- Why not just switch to an 8k-token encoder and stop chunking entirely?It removes the truncation risk but not the retrieval problem. Pooling thousands of tokens into one vector averages away specifics, so a single vector for a whole document is often less precise than several vectors for its sections — the document matches everything vaguely and nothing sharply. Long-context encoders are also slower and costlier per call. Use the longer window to stop splitting mid-idea, not to abandon granularity.
- How would you confirm truncation is the cause of a retrieval miss you have already noticed?Take a document whose tail contains the missing fact and tokenize it with the embedding model's own tokenizer. If the length exceeds the maximum sequence length, embed the whole document and the prefix alone: near-identical vectors prove the tail contributed nothing. Then check the corpus-wide distribution of chunk token lengths for a pile-up at the limit.
- Does the same limit apply to queries?Yes — the cap is a property of the encoder, not of the role the text plays. Queries are usually far shorter, so it rarely bites, but pipelines that expand a query with conversation history, prior turns or retrieved context can push it over, and the tail of the expanded query is then dropped. Assert the length on the query path too.
saying these in an interview costs you the question
- Assuming an over-length input raises an error rather than being truncated
- Sizing chunks in words or characters instead of tokens
- Believing the model compresses or summarizes the overflow
- Trusting a generic token estimate on jargon-heavy or code-heavy text
- Checking vector dimensionality as evidence the input was fully encoded