skip to content

In Ragas testset generation, what do knowledge-graph transforms do?

level: middleimportance: should knowfreq 42%

answer

  1. Graph is enriched before any question exists
  2. Three families, one ordered list
  3. Extractors write node properties
  4. Builders draw the edges
  5. No edges means no multi-hop

basics

~20 s

Transforms enrich the KnowledgeGraph before any question is written. Splitters break documents into smaller nodes, extractors attach properties like headlines, summaries, keyphrases and entities, and relationship builders draw edges between related nodes so multi-hop synthesizers have something to traverse.

solid answer

~50 s

Ragas does not synthesize questions straight from raw text; it first builds a `KnowledgeGraph` whose nodes are documents and chunks, then runs a transform pipeline over it. Three kinds of transform matter. **Splitters** such as `HeadlineSplitter` turn one document node into several child nodes. **Extractors** such as `HeadlinesExtractor`, `SummaryExtractor`, `KeyphrasesExtractor` and `NERExtractor` write properties onto each node — these are mostly LLM calls, one or more per node, which is where the cost lives. **Relationship builders** such as `CosineSimilarityBuilder` and `OverlapScoreBuilder` create edges between nodes that are semantically close or share entities. `default_transforms(documents, llm, embedding_model)` assembles a sensible pipeline, and `apply_transforms(kg, transforms)` runs it, mutating the graph in place. The payoff is that synthesizers then have structure to work with: summaries give abstract questions something to generalise over, and edges give multi-hop questions a path to follow.

code

python · 26 lines
python
from ragas.testset.graph import KnowledgeGraph, Node, NodeType
from ragas.testset.transforms import (
    HeadlinesExtractor,
    HeadlineSplitter,
    SummaryExtractor,
    apply_transforms,
)

kg = KnowledgeGraph()
for doc in docs:
    kg.nodes.append(
        Node(
            type=NodeType.DOCUMENT,
            properties={
                "page_content": doc.page_content,
                "document_metadata": doc.metadata,
            },
        )
    )

transforms = [
    HeadlinesExtractor(llm=transform_llm),
    HeadlineSplitter(),
    SummaryExtractor(llm=transform_llm),
]
apply_transforms(kg, transforms)

go deeper

for a junior

Know that Ragas builds a knowledge graph over your documents first, and that transforms are the steps that enrich it before any question is written.

for a middle

Name the three families — splitters, extractors, relationship builders — give a real class from each, and explain that extractors are per-node LLM calls while builders create the edges multi-hop questions need.

for a senior

Show the cost arithmetic (nodes times extractor passes), and demonstrate the diagnostic instinct: an all-easy testset points at a graph with no relationships, not at bad weights. Talk about trimming the pipeline to what your target question mix actually needs.

for a principal

Own the policy of which enrichment the organisation pays for by default — how much per-corpus preprocessing spend is justified by the difficulty of the questions it unlocks, and whether transform output is cached as a shared asset.

## Why a graph at all If you asked an LLM to "write a question about this chunk", you would get shallow, single-fact questions that a keyword search could answer. That is not a useful test of a RAG pipeline. Ragas' answer is to build an intermediate representation — a knowledge graph — rich enough that a synthesizer can ask genuinely hard things: questions that need a summary-level understanding of a section, or questions that can only be answered by combining two documents that mention the same entity. ## The graph A `KnowledgeGraph` holds `Node` objects and `Relationship` objects. Each node has a `type` (`NodeType.DOCUMENT` for a whole source document, `NodeType.CHUNK` for a piece of one) and a `properties` dict. Initially the only property of interest is the page content. Everything else on the node is put there by a transform. ## The three transform families **Splitters.** `HeadlineSplitter` takes a document node whose headlines have already been extracted and emits child chunk nodes, one per section. Note the ordering dependency: the splitter consumes a property that an extractor must have written first. This is why transform pipelines are ordered lists, not sets, and why hand-assembling one incorrectly produces an empty or near-empty graph rather than an error you can read. **Extractors.** These write node properties by calling the LLM (or the embedding model) once per node: - `HeadlinesExtractor` — the heading structure of a document. - `SummaryExtractor` — a short abstract of the node's content. - `KeyphrasesExtractor` / `ThemesExtractor` — salient phrases and themes. - `NERExtractor` — named entities mentioned in the node. - `EmbeddingExtractor` — a vector for the node, used by similarity-based relationship builders. - `TitleExtractor` — a title for the node. Each of these is an independent pass over every node. That is the arithmetic that surprises people: a 2,000-node graph with four LLM extractors is 8,000 model calls before a single question exists. **Relationship builders.** These are the ones people forget, and the omission is silent. `CosineSimilarityBuilder` compares node embeddings and creates an edge where similarity clears a threshold; `SummaryCosineSimilarityBuilder` does the same over summary embeddings rather than raw content; `OverlapScoreBuilder` creates edges between nodes that share extracted entities. Without these, the graph is a bag of disconnected nodes. ## Why the edges decide what questions you can get Single-hop synthesizers need one node. Multi-hop synthesizers need a *pair* of connected nodes — the whole point is that neither node alone answers the question. If your transform pipeline skipped relationship building, the multi-hop synthesizers have nothing to draw from, and you get a testset that is quietly all-easy no matter what weights you asked for. The `synthesizer_name` column in the output is the check: if you asked for a third multi-hop and see none, look at the graph's relationship count first. ## Running them `default_transforms(documents, llm, embedding_model)` inspects the corpus and returns an ordered pipeline appropriate for it — this is what `generate_with_langchain_docs` uses when you do not pass `transforms`. `apply_transforms(kg, transforms)` executes the pipeline against the graph, mutating it in place and returning nothing; you keep using the same `kg` object afterwards. Transforms that do not depend on each other can be wrapped so they run concurrently rather than as sequential passes. ## Tuning it The two levers worth knowing. First, **model choice for transforms**: `generate_with_langchain_docs` accepts `transforms_llm` and `transforms_embedding_model` separately from the generator's own models, so summarisation and entity extraction can run on a small cheap model while question writing uses a stronger one. Second, **pipeline trimming**: if you only want single-hop specific questions, you do not need summaries or entity overlap edges, and dropping those extractors can remove most of the cost. Conversely, if multi-hop abstract questions are the point of the exercise, the summary extractor and the summary-similarity builder are exactly what you must keep. ## What this leaf does not decide How the source documents were chunked and which embedding model represents them are upstream choices made when you loaded the corpus; the transforms operate on whatever nodes and vectors those choices produced.

  • What happens if you build a graph but skip the relationship builders entirely?
    Single-hop synthesis still works, but multi-hop synthesizers have no connected node pairs to draw on, so they contribute few or no samples. The run does not fail loudly — you just get a testset that is easier than the one you asked for. Check the `synthesizer_name` distribution in `to_pandas()` to catch it.
  • Why can you pass transforms_llm separately from the generator's own llm?
    Because the two jobs have different quality requirements. Summarising a chunk or pulling out named entities is routine work a small model does acceptably, and it runs once per node over the whole corpus. Writing the actual questions runs `testset_size` times and benefits from a stronger model. Splitting the models moves the bulk of the token spend onto the cheap one.
  • Does apply_transforms return a new graph or modify the one you pass in?
    It mutates the graph in place. You pass `apply_transforms(kg, transforms)` and then keep using the same `kg` object — the added node properties and relationships are on it. This matters when you intend to save the graph afterwards: save after applying, not before.

saying these in an interview costs you the question

  • Thinking transforms generate the questions themselves
  • Treating the transform list as unordered when splitters depend on extractors
  • Forgetting relationship builders and then blaming the synthesizer weights
  • Assuming transforms are pure embedding work with no LLM cost
  • Expecting apply_transforms to return a new graph

context