skip to content

In RAG, why embed small child chunks but return their larger parent sections?

level: middleimportance: must knowfreq 66%

answer

  1. searching and reading want different sizes
  2. precision from the child, context from the parent
  3. one vector per child, one id back
  4. deduplicate before you fetch
  5. the enlarged text costs context budget

basics

~20 s

Small chunks embed cleanly, so matching is precise; but a two-sentence hit often lacks the surrounding argument the model needs to answer. Small-to-big retrieval matches on the child and hands the generator the enclosing parent section instead.

solid answer

~50 s

Chunk size is pulled in two directions. Small chunks give a sharp embedding — one idea, one vector — so retrieval precision is high, but the retrieved text is frequently too thin to answer from. Large chunks carry enough context to answer but their embeddings are muddy averages of several topics, so they rank badly. Parent-document retrieval refuses the tradeoff by decoupling the unit you *search* from the unit you *read*: at index time you split into parent sections, subdivide each into small children, and store only the child vectors with a `parent_id`. At query time you retrieve top-k children, map them to parent ids, deduplicate, and fetch the parents to give the model. You pay for it in context budget and need dedupe when several children hit the same parent, so parent granularity — a subsection, not a whole chapter — is the tuning knob.

code

python · 19 lines
python
# index time: embed children, keep parents in a plain store
parent_store = {}
child_records = []

for parent_id, parent_text in enumerate(parent_sections):
    parent_store[parent_id] = parent_text
    for child_text in split_into_sentences(parent_text):
        child_records.append({"text": child_text, "parent_id": parent_id})

# query time: match on children, return deduplicated parents
def retrieve(query, k=8):
    hits = vector_search(query, child_records, k=k)
    seen, parents = set(), []
    for hit in hits:
        pid = hit["parent_id"]
        if pid not in seen:
            seen.add(pid)
            parents.append(parent_store[pid])
    return parents

go deeper

for a junior

Recall the shape: small pieces are what you search over, and the bigger section around a match is what you hand to the model. Be able to say why one size cannot do both jobs well.

for a middle

Walk through the index-time and query-time mechanics — child vectors carrying a parent id, dedupe on the way back, parents living in a plain document store — and name sentence-window expansion as the lighter variant of the same idea.

for a senior

Demonstrate the tradeoff in production terms: how parent granularity moves token cost and answer quality, how you deduplicate and re-rank when several children share a parent, and how you separated retrieval failures from context-sufficiency failures when debugging.

for a principal

Own the budget argument. Expansion trades tokens for recall on every single query, forever; decide where on that curve your product should sit given latency and cost targets, and be willing to say that for some corpora a flat index with better boundaries is the cheaper answer.

## The tension this resolves Every chunking discussion eventually hits the same wall. Make chunks small and each embedding represents one clear idea, so the nearest-neighbour search is precise — but the retrieved text may be a fragment that mentions the answer without explaining it. Make chunks large and there is enough surrounding argument for the model to work with — but the single vector now has to represent five subjects at once, and it sits in a mushy region of the embedding space where it matches everything weakly and nothing strongly. The insight of small-to-big is that these are two different jobs being forced onto one object. Searching wants a specific, narrow unit. Reading wants a complete, self-contained unit. There is no reason they have to be the same text. ## The mechanism At index time: split the document into parents — typically a subsection, a heading-bounded block, or a few paragraphs. Subdivide each parent into children — a sentence, a couple of sentences, a small paragraph. Embed the children only. Store each child vector with a `parent_id` pointing at the parent's text in a document store (which can be a plain key-value store; it never needs to be a vector index). At query time: embed the query, retrieve top-k children, collect their parent ids, deduplicate, fetch those parents, and pass the parent text to the generator. The vector index therefore holds many small, sharp vectors; the generator sees few large, complete passages. Note that this changes how much text you send, so how those enlarged passages get ordered and packed into the prompt is a separate concern from the retrieval design here. ## Sentence-window expansion: the same idea, cheaper A lighter variant does the expansion at read time instead of maintaining an explicit hierarchy. Chunk into individual sentences, embed each, but store with each sentence its position in the document. On a hit, return that sentence plus the *k* sentences before and after it. There is no parent object to define, no hierarchy to maintain, and one parameter to tune. Sentence-window is usually enough when the corpus is flowing prose and the answer's supporting context is genuinely adjacent — transcripts, articles, narrative reports. Explicit parent-document retrieval is better when the meaningful unit has real boundaries you would rather not cross, such as a documented API method or a clause in a contract, because a fixed window will happily run off the end of one and into the next. ## Where it costs you Context budget. Ten child hits can expand into ten parents. If parents are large, you have just quadrupled the tokens sent per query, with the cost and latency that implies — and pushed genuinely relevant text into the middle of a long prompt, where models attend to it less reliably. Deduplication is mandatory. Three children from the same parent must yield one copy of that parent, not three. Skip this and you burn budget on duplicates and skew any downstream reranking. Parent granularity is the real knob. Too small and you have not solved anything; too large — a whole chapter — and you are back to stuffing loosely relevant text at the model. A subsection-sized parent is the usual sweet spot. Ranking becomes indirect. Your top-ranked child may belong to a parent that, read whole, is less useful than the parent of the third-ranked child. If several children of one parent are hit, that is a reasonable signal to promote it. ## When it is the wrong tool If the child chunk already answers the question completely — short FAQ entries, product records, log lines — expansion adds tokens and no information. If your corpus is a set of tiny independent documents, there is no parent to return. And if the whole document is short enough to send anyway, the hierarchy is machinery you do not need. ## How to evaluate it Measure the two halves separately, because they fail differently. For retrieval, check whether the correct child appears in the top-k — that tells you whether the small-vector matching works. For generation, check whether the returned parent contained everything needed to answer — that tells you whether your parent is big enough. A system with high child recall and poor answers has parents that are too tight; a system with good answers but soaring token costs has parents that are too loose. ## The one-line version Embed what makes matching precise; return what makes answering possible. They are allowed to be different objects, and treating them as one is why so many first-attempt RAG systems feel like they retrieve the right thing and still answer badly.

  • Three of your top-five children come from the same parent. What do you do with that?
    Return the parent once — duplicates waste context and distort reranking. But treat the concentration as a ranking signal: several independent children matching the same parent is stronger evidence that the parent is on-topic than a single high-scoring child, so it is reasonable to promote it above a parent with one hit.
  • When is sentence-window expansion enough, and when do you need explicit parent sections?
    A fixed window works on flowing prose where the supporting context is literally adjacent — transcripts, articles, reports. You want explicit parents when the meaningful unit has hard boundaries you should not cross, such as one API method, one contract clause or one procedure step, because a symmetric window will spill into the neighbouring unit and add text that is topically close but factually about something else.
  • Does small-to-big retrieval mean you can stop caring about chunk quality?
    No. The child is still what gets matched, so a child that straddles two topics or lacks any identifying context still fails to be found. Small-to-big fixes insufficient context at generation time; it does nothing for poor boundaries or orphaned references at match time. Those are still fixed by structure-aware splitting and by prefixing chunks with their document and section identity.
  • How would you decide how large the parent should be?
    Empirically, against your own questions. Start at subsection size, then check two numbers: how often the returned parent actually contained the full answer, and the average tokens sent per query. Grow the parent while the first number is climbing; stop when it plateaus and cost keeps rising. A whole-chapter parent almost always sits past that point.

saying these in an interview costs you the question

  • Says just use big chunks, the model has a long context anyway
  • Embeds both the child and the parent into the same index
  • Returns one parent copy per matching child without deduplicating
  • Believes expansion fixes a child chunk that was never retrievable
  • Treats parent size as fixed rather than tuned against real questions

context