Your Ragas testset run over 5,000 documents takes hours — what do you change?
answer
- Two costs, very different scaling
- Corpus size drives the transforms
- testset_size is comparatively cheap
- Save and reload the graph
- Cheap model for the big sweep
basics
~20 sThe cost is in the transforms, which run per node over the whole corpus, not in the questions, which scale with testset_size. Build and enrich the KnowledgeGraph once, save it with kg.save(), reload it for later runs, and use a cheaper transforms_llm.
solid answer
~50 sSeparate the two costs first. Graph transforms scale with **corpus size** — every extractor is at least one LLM call per node, and a default pipeline runs several of them, so 5,000 documents split into chunks is easily tens of thousands of calls. Query synthesis scales with **testset_size**, which is usually tens. So the corpus phase dominates and the fix is to stop repeating it: build the graph once, `apply_transforms` once, `kg.save("kg.json")`, and on every later run `KnowledgeGraph.load("kg.json")` into `TestsetGenerator(..., knowledge_graph=kg)` and call `generate(testset_size=...)` instead of `generate_with_langchain_docs`. Beyond caching: pass a cheap `transforms_llm` so summarisation and entity extraction do not use your expensive model, trim the transform list to the extractors your target query mix actually needs, and consider generating from a representative sample of the corpus rather than all of it — a testset of 50 questions does not need 5,000 documents behind it.
code
python · 21 linesfrom ragas.testset import TestsetGenerator
from ragas.testset.graph import KnowledgeGraph
from ragas.testset.transforms import default_transforms, apply_transforms
# --- build once (expensive: scales with corpus size) ---
kg = KnowledgeGraph()
# ... append Node objects for each document ...
apply_transforms(
kg,
default_transforms(docs, llm=cheap_llm, embedding_model=embeddings),
)
kg.save("kg.json")
# --- reuse many times (cheap: scales with testset_size) ---
kg = KnowledgeGraph.load("kg.json")
generator = TestsetGenerator(
llm=strong_llm,
embedding_model=embeddings,
knowledge_graph=kg,
)
testset = generator.generate(testset_size=50)go deeper
Know that generating a testset costs real model calls and that the count scales with how many documents you feed in, not just with how many questions you ask for.
Separate the two cost curves out loud — transforms scale with corpus size, synthesis with testset_size — and name the save/load pair on the knowledge graph as the way to stop repeating the expensive half.
Show the full remediation: cache the graph as a versioned build artefact, split transforms_llm from the generator model, trim the pipeline to the synthesizers you actually use, sample the corpus, and structure the run so a rate-limit failure does not cost you the whole sweep.
Own the position that testset generation is a build step and not a CI step, and set the organisation's policy on how generated sets are curated, frozen, versioned and refreshed as the corpus drifts.
## Find the real cost centre The instinct is to lower `testset_size`, and it is the wrong lever. Do the arithmetic on the two phases separately. **Phase one, transforms.** Splitting 5,000 documents into sections might produce 20,000–40,000 chunk nodes. A default transform pipeline runs several LLM-backed extractors — headlines, summaries, keyphrases, entities — plus embedding passes. Each is a full sweep over every node. Four LLM extractors over 30,000 nodes is 120,000 model calls, plus 30,000 embedding calls, plus the similarity comparisons for relationship building. That is the hours. **Phase two, synthesis.** You asked for, say, 50 questions. Each is a handful of calls. That is minutes at most. So the run is roughly all preprocessing, and `testset_size` is nearly free by comparison. Every optimisation should target phase one. ## Fix one: stop repeating the graph build This is the big one, and it is the reason `generate_with_langchain_docs` is a convenience method rather than the only entry point. That method does ingestion, transforms and synthesis in one call — which means every time you re-run it to regenerate a testset, tweak weights, or try different personas, you pay the entire corpus preprocessing bill again for a graph that has not changed. Split it. Build the `KnowledgeGraph`, `apply_transforms(kg, transforms)`, then `kg.save("kg.json")`. Treat that file as a build artefact: it is derived from a specific corpus snapshot with a specific transform pipeline, so version it alongside them. Every subsequent run is `KnowledgeGraph.load("kg.json")`, construct `TestsetGenerator(llm=..., embedding_model=..., knowledge_graph=kg)`, and call `generate(testset_size=...)`. Now iterating on the query distribution costs minutes instead of hours, which changes the whole ergonomics of tuning the mix. Rebuild the graph when the corpus changes materially or when you change the transform pipeline — not otherwise. ## Fix two: split the models `generate_with_langchain_docs` accepts `transforms_llm` and `transforms_embedding_model` separately from the generator's own models. The work these do — summarise this chunk, list the entities in it — is routine and a small fast model handles it acceptably. Question writing is the part that benefits from a strong model, and it runs `testset_size` times rather than node-count times. Putting the frontier model on the small job and the cheap model on the enormous one is exactly backwards, and it is the default if you construct the generator with one model and never think about it. ## Fix three: trim the pipeline The default transform set is built to support every synthesizer. If you only want single-hop specific questions, you do not need summaries or summary-similarity edges. If you want multi-hop abstract questions, the summary extractor is essential and the entity extractor may not be. Assemble the transform list deliberately for the question mix you are targeting and drop the passes that feed synthesizers you are not using. Each dropped extractor removes a full sweep of the corpus. Also mind ordering: transforms that do not depend on one another can run concurrently rather than as sequential sweeps, which turns wall-clock time into a provider-throughput problem instead of a serial one. ## Fix four: question whether you need the whole corpus A testset of 50 questions cannot possibly cover 5,000 documents anyway — 50 questions touch at most a hundred or so nodes. If the goal is a regression suite rather than exhaustive coverage, generating from a stratified sample of the corpus (a few hundred documents spanning the domains you care about) gives you a comparable testset for a fraction of the preprocessing. Sample deliberately rather than by taking the first N documents, so the sample spans the topical range. ## What to watch while it runs At this volume you are also hitting the provider's rate limits, and the failure mode is a long run that dies most of the way through. Run the transform phase as its own step so it can be retried independently of synthesis, and save the graph immediately when it completes. A graph you have to rebuild because the synthesis step failed afterwards is the most expensive kind of mistake here. ## The judgment to state out loud The underlying point an interviewer is listening for: synthetic testset generation is a *build* step, not a *test* step. It does not belong inside the loop that runs on every commit. You generate, you curate, you freeze the result, and the frozen file is what CI reads. Regenerating a testset per run is both expensive and, because generation is non-deterministic, actively harmful to the stability of the numbers.
- Which call replaces generate_with_langchain_docs once you have a saved graph?Load the graph, pass it to the constructor as `knowledge_graph=`, and call `generate(testset_size=...)`. That skips ingestion and transforms entirely and goes straight to query synthesis, so iterating on the query distribution or personas costs minutes rather than repeating the whole corpus sweep.
- When do you have to rebuild the saved knowledge graph rather than reuse it?When the corpus changes materially — new or substantially rewritten documents — or when you change the transform pipeline itself, since the graph's node properties and relationships are the output of that specific pipeline. Treat the saved file as a build artefact keyed to a corpus snapshot plus a pipeline definition, and version it that way.
- Why is regenerating the testset on every CI run a bad idea even if it were fast?Because generation is non-deterministic. Different questions every run means the score moves for reasons that have nothing to do with your change, and you can no longer tell a regression from resampling noise. Generate once, curate the rows by hand, freeze the file, and let CI read the frozen artefact.
- Where does the rate limit bite in a run this size, and how do you structure around it?In the transform phase, which issues tens of thousands of calls in bulk. Run it as a separate step from synthesis so it can be retried on its own, and save the graph the moment it finishes — losing a completed transform sweep because a later synthesis step failed is the most expensive failure available here.
saying these in an interview costs you the question
- Lowering testset_size to cut the cost of a corpus-scale run
- Re-running generate_with_langchain_docs for every tweak to the weights
- Using the expensive model for per-node summarisation and extraction
- Regenerating the testset inside CI on every commit
- Assuming the whole corpus must be ingested for a 50-question suite