skip to content

Which Ragas calls turn your own documents into a synthetic testset?

level: juniorimportance: must knowfreq 55%

answer

  1. Two calls, not one
  2. Construct, then generate over docs
  3. One classmethod wraps LangChain models
  4. The fused name is not real
  5. from_langchain plus generate_with_langchain_docs

basics

~10 s

Two calls. TestsetGenerator.from_langchain(llm, embedding_model) builds the generator from a LangChain model pair, then generator.generate_with_langchain_docs(docs, testset_size=N) returns a Testset of N synthetic samples with questions, reference contexts and reference answers.

solid answer

~40 s

Ragas testset generation is a two-step API. First you build the generator: `TestsetGenerator.from_langchain(llm, embedding_model)` takes a LangChain chat model and a LangChain embeddings object and wraps both for ragas internally. Then you run it: `generator.generate_with_langchain_docs(documents, testset_size=20)` takes the loaded LangChain `Document` objects and returns a `Testset`. Under the hood that second call builds a knowledge graph over the documents, applies a default set of transforms to enrich it, and then asks query synthesizers to write questions against it. The result exposes `to_pandas()` for inspection — columns like `user_input`, `reference_contexts`, `reference` and `synthesizer_name` — and `to_evaluation_dataset()` to hand straight to a ragas evaluation. There is no `from_langchain_docs`; that fused name is a common mis-memory.

code

python · 13 lines
python
from langchain_openai import ChatOpenAI, OpenAIEmbeddings
from langchain_community.document_loaders import DirectoryLoader
from ragas.testset import TestsetGenerator

docs = DirectoryLoader("./docs", glob="**/*.md").load()

generator = TestsetGenerator.from_langchain(
    ChatOpenAI(model="gpt-4o-mini"),
    OpenAIEmbeddings(model="text-embedding-3-small"),
)

testset = generator.generate_with_langchain_docs(docs, testset_size=20)
print(testset.to_pandas().head())

go deeper

for a junior

Be able to name the two calls in order and say what each takes: from_langchain builds the generator from a model plus embeddings, generate_with_langchain_docs runs it over loaded documents with a testset_size.

for a middle

Explain the three internal phases — graph build, transforms, query synthesis — and which arguments let you override each, including transforms_llm for using a cheaper model during enrichment.

for a senior

Show that you know generation is non-deterministic and that a shipped regression suite is a curated, frozen file, not a call you re-run in CI. Talk about where the corpus-size cost sits versus the testset-size cost.

for a principal

Own the decision of whether synthetic bootstrapping is the right starting point at all versus paying for human-authored questions, and set the policy for how a generated set graduates into a gating suite.

## The problem this solves Before you can measure a RAG pipeline you need a set of questions with known-good answers and known-good contexts. Hand-writing a hundred of those over your own corpus is days of work, and it is the single most common reason a team never gets around to evaluating anything. Ragas testset generation bootstraps that set from the documents you already have. ## The two calls Ragas separates *constructing* the generator from *running* it. **Construction.** `TestsetGenerator.from_langchain(llm, embedding_model)` is a classmethod convenience. You hand it a LangChain chat model and a LangChain embeddings object; it wraps them in ragas' own evaluator wrappers so the rest of the library can call them uniformly, and returns a generator with an empty knowledge graph. You can also construct `TestsetGenerator(llm=..., embedding_model=..., knowledge_graph=kg)` directly when you already have wrapped models or a prebuilt graph. **Execution.** `generate_with_langchain_docs(documents, testset_size=N)` takes a list of LangChain `Document` objects — whatever your loader produced — and returns a `Testset` with roughly `N` samples. Optional arguments let you override the pipeline: `transforms` (the knowledge-graph enrichment steps), `transforms_llm` and `transforms_embedding_model` (so graph building can use a cheaper model than question writing), and `query_distribution` (the mix of question types). There is a sibling `generate_with_llamaindex_docs` for LlamaIndex documents, and a bare `generate(testset_size=...)` that skips document ingestion and synthesizes directly from a knowledge graph you already built. ## What happens inside the second call Three phases, and it helps to know them because that is where the time and the money go: 1. **Graph construction.** Each input document becomes a node in a `KnowledgeGraph`. 2. **Transforms.** A default transform pipeline runs over the graph: splitters break documents into smaller nodes, extractors write properties onto nodes (headlines, summaries, keyphrases, named entities, embeddings), and relationship builders draw edges between nodes that are similar or share entities. Most of these steps are LLM or embedding calls, per node. 3. **Query synthesis.** Query synthesizers walk the enriched graph and write questions. Single-hop synthesizers use one node; multi-hop synthesizers follow a relationship between two nodes so the question can only be answered by combining them. Each generated sample carries the question, the contexts it was built from, and a reference answer. ## The output object `Testset.to_pandas()` gives you a DataFrame you can eyeball — this is not optional in practice, because generated questions vary in quality and you will want to delete some. The `synthesizer_name` column tells you which synthesizer produced each row, which is how you check that your requested easy/hard mix actually materialised. `to_evaluation_dataset()` converts the testset into the dataset shape a ragas evaluation run consumes, so the generated `reference` and `reference_contexts` can feed reference-requiring metrics. ## The naming trap A large amount of tutorial content and model recall contains `TestsetGenerator.from_langchain_docs(...)`. That method has never existed. It is a blend of the two real names, and writing it produces an `AttributeError` immediately. The real pair is `from_langchain` (construct) and `generate_with_langchain_docs` (run). Older ragas releases, before the knowledge-graph rewrite, exposed a different generator API entirely with docstore and distribution arguments — if you are reading a tutorial that mentions `simple`, `reasoning` and `multi_context` ratios, it predates the current line and none of the code will transfer. ## Practical notes `testset_size` is the number of *samples out*, not documents in — you can generate 20 questions from 5,000 documents or from 5. The corpus size drives the transform cost; `testset_size` drives the synthesis cost. Generation is non-deterministic: the same corpus and the same size will not give you the same questions twice, so if you want a stable regression suite you generate once, curate the rows by hand, and freeze the file.

  • What do you get back from generate_with_langchain_docs, and how do you feed it to an evaluation?
    A `Testset` object. `to_pandas()` renders it as a DataFrame with the question (`user_input`), the contexts it was built from (`reference_contexts`), the generated ground-truth answer (`reference`) and `synthesizer_name`. `to_evaluation_dataset()` converts it to the dataset shape a ragas evaluation consumes, so reference-requiring metrics have ground truth available.
  • If you already have a prepared knowledge graph, do you still call generate_with_langchain_docs?
    No. Pass the graph to the constructor as `knowledge_graph=` and call `generate(testset_size=...)` instead. `generate_with_langchain_docs` exists to do the ingestion and transform work for you; once that work is done and saved, going through it again just repeats the expensive part.
  • Is testset_size a hard guarantee on the number of rows returned?
    Treat it as a target, not a contract. Synthesis can drop samples — a synthesizer may fail to produce a usable question from the nodes it drew, and multi-hop synthesizers can find no eligible node pair at all if the graph has few relationships. Always check `len(testset.to_pandas())` rather than assuming.

saying these in an interview costs you the question

  • Calling TestsetGenerator.from_langchain_docs — that method does not exist
  • Thinking testset_size controls how many documents are ingested
  • Assuming two runs over the same corpus produce identical questions
  • Believing generation only embeds documents and makes no LLM calls
  • Quoting the pre-knowledge-graph API with simple/reasoning/multi_context ratios

context