skip to content

In RAG chunking, what does chunk overlap buy you, and what does it cost?

level: middleimportance: must knowfreq 70%

answer

  1. insurance against boundary loss
  2. the tail repeats at the head
  3. the window slides by size minus overlap
  4. duplicated text you pay to embed twice
  5. near-duplicate hits crowd out top-k

basics

~20 s

Overlap repeats the tail of one chunk at the start of the next, so a fact cut by a boundary still appears whole somewhere. The cost is duplicated text: more chunks to embed and store, and near-duplicate search hits.

solid answer

~50 s

Length-based splitting cuts wherever the counter runs out, which regularly lands mid-sentence or between a claim and the number that supports it. Overlap is insurance against that: each chunk begins by repeating the last N tokens of its predecessor, so any span shorter than the overlap is guaranteed to sit intact inside at least one chunk. You pay for it in duplication. With chunk size S and overlap V, the number of chunks scales as roughly S/(S−V), so 50% overlap doubles your embedding bill, index size and query-time scan. It also hurts diversity: two adjacent chunks sharing most of their text often both land in the top-k, crowding out an independent source. In practice most teams sit around 10–20% of chunk size and deduplicate near-identical hits after retrieval. Overlap fixes boundary loss only — it cannot rescue a fact longer than the chunk itself.

code

python · 6 lines
python
def fixed_chunks(text, size=800, overlap=100):
    assert 0 <= overlap < size
    step = size - overlap
    return [text[i:i + size] for i in range(0, len(text), step)]

print(fixed_chunks("abcdefghij", size=4, overlap=1))

go deeper

for a junior

Know that overlap means each chunk repeats a little of the previous one, and be able to say why: so a sentence cut at a boundary still shows up whole in one chunk.

for a middle

Be ready to do the arithmetic — the window advances by size minus overlap, so chunk count scales as S/(S−V) — and to name the concrete costs: more vectors to embed, store and search.

for a senior

Show that you have felt the diversity cost in production: overlapping neighbours occupying multiple top-k slots, and the post-retrieval dedup or span-merge you added to fix it.

for a principal

Frame overlap as a cost knob on the whole index, not a quality knob. Own the position that it should be justified by measured retrieval gain on an eval set, since it multiplies re-indexing spend across every future model change.

## What overlap actually is Fixed-size splitting walks a document and emits a chunk every time a length counter — in characters or tokens — is exhausted. Overlap changes the *step* the window advances by, not the window's width. With a chunk size of 800 tokens and an overlap of 100, the first chunk covers tokens 0–800, the second 700–1500, the third 1400–2200, and so on. The window is always 800 wide; it advances 700 each time. That last 100 tokens of every chunk is reproduced verbatim as the first 100 tokens of the next. This is why the pattern is often described as a sliding window: the window slides by step = size − overlap. ## Why boundaries lose information A length-driven cut has no idea what it is cutting. It will happily land in the middle of a sentence, between a defined term and its definition, or between a table row's label and its value. The damage is not just cosmetic — it shows up twice. First at *index* time: the fragment "...the surcharge applies only when the shipment exceeds" is embedded as its own vector, and that vector encodes a truncated, half-meaningless idea. It will not sit near the query "when does the surcharge apply?" in embedding space. Second at *read* time: even if retrieval finds the chunk, the model reading it sees a sentence that stops mid-clause and either hedges, hallucinates the missing half, or answers from the fragment it can see. The classic failure is a freight tariff table split at a fixed token count, cutting a rate row so the lane price ends up in one chunk and its unit — per hundredweight, per pallet — in the next. Both chunks are individually plausible; the answer built from either is wrong. Overlap targets exactly this. Any contiguous span shorter than the overlap length is mathematically guaranteed to appear uncut in at least one chunk, because the window advances by less than the span's length. That is a real guarantee, and it is the whole value proposition. ## What overlap costs The cost is duplication, and it compounds across every part of the pipeline. - **Chunk count.** Chunks scale as roughly S/(S−V) for size S and overlap V. Ten percent overlap adds about 11% more chunks; fifty percent overlap doubles them. - **Embedding spend.** Every duplicated token is embedded again, at index time and again on every re-index. - **Storage and search.** More vectors means a larger index, more memory for the approximate-nearest-neighbour structure, and a slightly slower query. - **Result diversity.** This is the cost people forget. Two adjacent chunks that share 40% of their text tend to score almost identically for the same query, so both occupy slots in the top-k. Your k=8 retrieval effectively returns five distinct sources instead of eight, and an independent document that would have answered the question gets pushed off the list. ## Choosing a value There is no universal number, but the reasoning is stable. Set overlap to at least the length of the *unit you cannot afford to split* — a full sentence, a table row, a code statement — measured in the same units as chunk size. In ordinary prose that puts you at roughly 10–20% of chunk size, which is where most production systems land. Going far above that buys shrinking returns because the atomic units are already covered, while the duplication cost keeps growing linearly. Two mitigations pair well with overlap: deduplicate or merge near-identical chunks after retrieval so overlapping neighbours consume one slot rather than two, and keep byte offsets on every chunk so you can detect and collapse overlaps programmatically. ## What overlap does not fix Overlap is a boundary patch, not a sizing fix. If the fact you need spans 900 tokens and your chunks are 400 with 50 overlap, no chunk will ever contain it whole; the fix there is a larger chunk, not more overlap. Overlap also does nothing about a chunk that is semantically incoherent for reasons other than truncation, and it cannot recover context that lives elsewhere in the document, such as a section heading a thousand tokens earlier. Recognising which of these three problems you actually have — a cut boundary, an undersized chunk, or missing document context — is the distinction interviewers are listening for.

  • How would you decide the overlap value for a corpus of legal clauses rather than prose?
    Measure the atomic unit you cannot split. For legal clauses that is usually a full numbered clause including its proviso, so sample the corpus, take a high percentile of clause length in tokens, and set overlap at least that large. If that pushes overlap past roughly a quarter of chunk size, the real signal is that chunk size is too small, not that overlap is too small.
  • What retrieval-side change would you make to offset the diversity cost of heavy overlap?
    Deduplicate after retrieval. Keep source document id and byte offsets on every chunk, then collapse hits whose spans overlap substantially into a single merged passage before building the prompt. That way two neighbouring chunks consume one context slot instead of two, and the top-k keeps its intended number of independent sources.
  • If you doubled overlap and retrieval quality did not change, what would you conclude?
    That boundary loss was not your bottleneck. The overlap you already had covered the atomic units in the corpus, so the extra duplication bought nothing but cost. Roll it back and look elsewhere — chunk size relative to the answer span, the query-to-chunk embedding match, or ranking — rather than tuning overlap further.

saying these in an interview costs you the question

  • Thinks more overlap always improves recall
  • Believes overlap can hold a fact longer than the chunk
  • Ignores that overlap multiplies embedding and storage cost
  • Never considers that overlapping neighbours both land in top-k
  • Sets overlap to half the chunk size as a default

context