skip to content

How would you prove a chunking change actually improved a LlamaIndex RAG pipeline?

level: principalimportance: should knowfreq 38%

answer

  1. split the pipeline before measuring it
  2. labels for retrieval, labels for answers
  3. hold the judge constant
  4. cost and latency belong in the report
  5. small sets cannot detect small effects

basics

~20 s

Freeze an eval set, then measure retrieval and generation separately: RetrieverEvaluator with hit_rate and MRR over labelled query/node pairs, and BatchEvalRunner with the faithfulness and correctness evaluators over answers. Hold the judge model fixed, and report cost and latency next to quality.

solid answer

~50 s

The first move is to split the pipeline. `RetrieverEvaluator.from_metric_names(["hit_rate", "mrr"], retriever=retriever)` scores retrieval alone against labelled question-to-node pairs, which you can bootstrap with `generate_question_context_pairs` over your nodes and then hand-clean. Generation quality goes through `BatchEvalRunner` with faithfulness, relevancy and correctness evaluators. Measuring both matters because a chunking change acts on retrieval first: if hit rate did not move, any change in answer scores is judge noise. Then control the experiment — same eval set, same judge model at temperature 0, same generator, one variable changed — and re-run both halves for each configuration. Report token cost and p95 latency alongside the quality numbers, because smaller chunks with a higher top-k usually buy recall with prompt tokens. Finally, be honest about power: on a fifty-question set a two-point move is noise, so size the set to the effect you need to detect, or repeat runs.

code

python · 21 lines
python
import asyncio

from llama_index.core import Document, VectorStoreIndex
from llama_index.core.evaluation import RetrieverEvaluator, generate_question_context_pairs
from llama_index.core.node_parser import SentenceSplitter
from llama_index.llms.openai import OpenAI

docs = [Document(text="Refunds are accepted within 30 days of purchase. " * 40)]
label_llm = OpenAI(model="gpt-4o", temperature=0)


def hit_rate_for(chunk_size: int) -> float:
    nodes = SentenceSplitter(chunk_size=chunk_size, chunk_overlap=20).get_nodes_from_documents(docs)
    qa_dataset = generate_question_context_pairs(nodes, llm=label_llm, num_questions_per_chunk=1)
    retriever = VectorStoreIndex(nodes).as_retriever(similarity_top_k=5)
    evaluator = RetrieverEvaluator.from_metric_names(["hit_rate", "mrr"], retriever=retriever)
    results = asyncio.run(evaluator.aevaluate_dataset(qa_dataset))
    return sum(r.metric_vals_dict["hit_rate"] for r in results) / len(results)


print(hit_rate_for(256), hit_rate_for(1024))

go deeper

for a junior

Know that you need a fixed set of questions with expected answers, and that you compare configurations on that same set rather than on ad-hoc questions you thought of.

for a middle

Explain how retrieval metrics like hit rate and MRR are measured separately from answer quality, and why holding the judge model and eval set constant is what makes two runs comparable.

for a senior

Design the experiment: one variable at a time, paired comparison on flipped rows, per-row feedback for triage, and token cost and p95 latency reported next to every quality number.

for a principal

Own the evaluation programme — what the eval set must represent, its refresh and versioning policy, the statistical power needed for the decisions it gates, judge drift, and what quality a given cost increase is worth.

## Why "it feels better" is not an answer Chunk size, overlap, splitter choice, top-k and reranking all interact, and every one of them changes the text that reaches the model. Without a fixed measurement, a team cycles through configurations forever, each change validated by whoever tried three questions they remembered. The interview question behind this is whether you can design an experiment, not whether you can name a metric. ## Step one: build the eval set, and treat it as an asset You need two kinds of labels, and they are not the same. For retrieval you need question → correct-node pairs. `generate_question_context_pairs(nodes, llm=..., num_questions_per_chunk=2)` synthesises these by asking an LLM to write questions answerable from each node, producing an `EmbeddingQAFinetuneDataset` that maps query ids to the node ids that should be retrieved. Synthetic generation gets you to a usable set in an hour; a human pass to delete degenerate questions ("what does this section say?") is what makes it trustworthy. The known weakness is that synthetic questions inherit the chunking used to generate them, which biases comparisons between chunkings — mitigate by generating from the raw documents or by keeping node-level labels aligned to source spans rather than to one splitter's output. For generation you need question → reference answer pairs, which is what `CorrectnessEvaluator` consumes. These are more expensive to produce and more valuable; they are also the set real stakeholders will read when they challenge a result. Version both sets in the repo. An eval set that quietly changes between runs destroys every comparison you have made. ## Step two: measure retrieval independently `RetrieverEvaluator.from_metric_names(["hit_rate", "mrr"], retriever=retriever)` builds an evaluator over a retriever; `aevaluate_dataset(qa_dataset)` runs it across the labelled set. Hit rate answers "was the right node retrieved at all", MRR answers "how high did it rank". This half is cheap — embeddings and vector search, no judge LLM — so run it first and run it often. The reason to isolate this half is causal clarity. Chunking changes what can possibly be retrieved. If hit rate is flat between two chunk sizes, then the retrieved evidence is effectively unchanged, and any wobble in downstream answer scores is judge variance rather than an improvement. If hit rate rose but answer quality did not, the bottleneck is synthesis, and the next experiment belongs there. ## Step three: measure generation Run `BatchEvalRunner` with the faithfulness, relevancy and correctness evaluators over the same questions, end-to-end through the query engine. Keep the per-row results, not just the aggregate: the interesting output is the set of questions whose verdict flipped between configurations, and the judge's `feedback` on them. Ten flipped rows read carefully teach more than a two-decimal aggregate. ## Step four: control the experiment One variable at a time. Same eval set, same judge model, judge temperature 0, same generator and same top-k when chunking is what you are testing. Record the full configuration with the result — chunk size, overlap, splitter, embedding model, top-k, reranker, generator, judge. A result without its configuration is not reproducible, and six weeks later nobody remembers which run had reranking on. Judge drift deserves specific attention: when the judge model version changes, historical numbers are no longer comparable to new ones. Either pin the judge or re-baseline the whole history when you move it. ## Step five: report cost and latency next to quality Smaller chunks usually improve hit rate, because a precise chunk embeds more cleanly — but they also push you toward a higher top-k to retain context, which inflates prompt tokens linearly. Token counting gives you that number. A configuration that gains three points of faithfulness while doubling prompt cost and adding 400ms to p95 is a tradeoff for a product owner to make, and presenting the quality number alone is quietly dishonest. ## Step six: respect statistical power Small eval sets plus a nondeterministic judge means small deltas are noise. Before believing a result, ask what effect size the set can detect: on fifty questions, one flipped answer is two points. The practical rules are to grow the set for decisions that matter, repeat borderline runs to see the spread, and prefer paired comparison — the same questions under both configurations, looking at which rows flipped — over comparing two independent aggregates. ## Making it durable Once this exists, wire it in: the cheap retrieval half on every change to indexing code, the expensive generation half on release candidates and nightly. Publish the numbers where the team sees them. The value is less any single measurement than the fact that the next chunking argument gets settled by a run rather than by seniority.

  • Retrieval hit rate improved but correctness scores did not move. What do you conclude?
    That retrieval is no longer the bottleneck for those questions — the evidence now reaches the model, and the model is not using it. The next experiment belongs downstream: synthesis strategy, context ordering, prompt wording, or generator capability. It is also worth checking whether the flat correctness is real or noise, by looking at which individual rows changed verdict rather than at the aggregate, since a set of the usual size cannot resolve a small move.
  • What is the danger of generating your retrieval eval set from the same nodes you are testing?
    The synthetic questions are written to be answerable from those specific chunks, so they are biased toward the chunking that produced them — a comparison between two chunk sizes partly measures which one wrote the questions. Mitigate by generating questions from the raw documents, or by anchoring labels to source spans rather than to one splitter's node ids, and by keeping a human-written subset that is independent of any chunking decision.
  • How do you decide when a measured improvement is worth shipping?
    By putting quality, cost and latency in the same table and treating it as a product decision, not a technical one. A three-point faithfulness gain that doubles prompt tokens and adds 400ms at p95 may be right for a low-volume expert tool and clearly wrong for a high-volume assistant. The engineering obligation is to produce all three numbers with their confidence, and to say plainly when the delta is inside the noise band of the eval set.

saying these in an interview costs you the question

  • Judging a chunking change by spot-checking a few questions
  • Measuring only end-to-end answer quality and never retrieval
  • Changing chunk size and top-k in the same experiment
  • Comparing runs whose judge model version differs
  • Acting on a two-point move across fifty questions

context