In Ragas, how do you control the single-hop vs multi-hop query mix?
answer
- A list of weighted synthesizers
- Weights are proportions summing to one
- Easy questions do not discriminate
- Hard ones need graph edges
- Check synthesizer_name in the output
basics
~10 sPass query_distribution: a list of (synthesizer, weight) tuples whose weights sum to 1, using SingleHopSpecificQuerySynthesizer, MultiHopAbstractQuerySynthesizer and MultiHopSpecificQuerySynthesizer. default_query_distribution(llm) supplies a starting mix you can reweight.
solid answer
~50 sThe `query_distribution` argument decides what kind of questions come out. It is a list of `(synthesizer, weight)` tuples where the weights sum to 1, and ragas draws each sample from that distribution. The three synthesizers in the current line are `SingleHopSpecificQuerySynthesizer` (one node, one concrete fact), `MultiHopSpecificQuerySynthesizer` (two connected nodes joined by a shared entity) and `MultiHopAbstractQuerySynthesizer` (two connected nodes joined at a summary/theme level, producing broader questions). `default_query_distribution(llm)` gives you all three with a default weighting that leans on single-hop; you pass your own list to shift it. The reason this is the tunable that matters: a testset that is all single-hop specific questions is passed by almost any retriever, so it cannot discriminate between a good pipeline and a mediocre one. Multi-hop questions are where retrieval actually gets tested — but they only materialise if the knowledge graph has relationships to traverse.
code
python · 18 linesfrom ragas.testset.synthesizers import (
default_query_distribution,
SingleHopSpecificQuerySynthesizer,
MultiHopAbstractQuerySynthesizer,
MultiHopSpecificQuerySynthesizer,
)
distribution = [
(SingleHopSpecificQuerySynthesizer(llm=generator_llm), 0.3),
(MultiHopSpecificQuerySynthesizer(llm=generator_llm), 0.4),
(MultiHopAbstractQuerySynthesizer(llm=generator_llm), 0.3),
]
testset = generator.generate_with_langchain_docs(
docs, testset_size=30, query_distribution=distribution
)
print(testset.to_pandas()["synthesizer_name"].value_counts())go deeper
Know that Ragas can generate both easy single-hop and harder multi-hop questions, and that the mix is set by a query_distribution argument rather than being fixed.
Name the three synthesizer classes, describe the (synthesizer, weight) tuple list summing to 1, and explain what makes a multi-hop question harder — it needs two connected nodes, not one.
Demonstrate the diagnostic loop: verify the realised mix from synthesizer_name, and trace a missing multi-hop share back to an under-connected graph rather than to the weights. Argue for a mix that discriminates in both directions.
Own the difficulty policy for the organisation's eval suites — how hard a suite has to be to detect the regressions you actually care about, and the tradeoff between discrimination and the generation spend that harder questions incur.
## What a query synthesizer is After the knowledge graph is built and enriched, ragas needs to decide what kind of question to write from it. A query synthesizer is the component that does that: it selects node(s) from the graph, applies a prompt strategy, and emits a sample containing the question, the contexts it drew on, and a reference answer. Different synthesizer classes select differently and prompt differently, and that is the whole source of difficulty variation in a generated testset. ## The three synthesizers **`SingleHopSpecificQuerySynthesizer`** picks a single node and writes a question whose answer sits in that one node — typically anchored on a concrete extracted item such as a keyphrase or an entity. These are the easy questions. A retriever only has to find one chunk. **`MultiHopSpecificQuerySynthesizer`** picks two nodes connected by an entity-overlap relationship and writes a question that requires both. Think "what does document A say about X that document B contradicts" — concrete, but the answer is not in either chunk alone. **`MultiHopAbstractQuerySynthesizer`** also uses two connected nodes, but works at the summary/theme level rather than the entity level, producing broader, more conceptual questions that generalise across the pair. ## The distribution `query_distribution` is a list of `(synthesizer_instance, weight)` tuples. The weights are proportions and must sum to 1. Each synthesizer is instantiated with the LLM it should use, so you can in principle even give different synthesizers different models. `default_query_distribution(llm)` builds the standard three-way mix; calling it and then editing the weights is usually easier than assembling the list from scratch, because it also gets the constructor arguments right for you. You pass it either to `generate_with_langchain_docs(docs, testset_size=N, query_distribution=dist)` or to `generate(testset_size=N, query_distribution=dist)`. ## Why this is the tunable that matters A testset made entirely of single-hop specific questions is a weak instrument. Almost any embedding-based retriever finds the one chunk that contains the answer, so every candidate pipeline scores near the top and the test tells you nothing about which is better. Difficulty is what creates discrimination. Shifting weight toward the multi-hop synthesizers gives you questions where retrieval has to assemble evidence, which is exactly the failure mode real users hit and naive retrieval misses. The opposite error is also real. A testset that is overwhelmingly multi-hop abstract can be *unfairly* hard: it may demand a synthesis that your product was never designed to do, so scores are uniformly low and again fail to discriminate. A mix — enough easy questions that a regression in the basics is visible, enough hard ones that improvements register — is the practical target. ## The silent failure Weights are a request, not a guarantee. Multi-hop synthesizers need pairs of *connected* nodes; if the transform pipeline that ran over your graph did not include relationship builders, or its similarity threshold was strict enough that few edges were drawn, the multi-hop synthesizers will find nothing eligible and contribute few or no samples. You asked for 50% hard questions and got 5%, with no error raised. The diagnostic is cheap and you should make it routine: after generation, group the `to_pandas()` output by `synthesizer_name` and compare the realised proportions against the weights you passed. If they do not match, inspect the graph's relationship count before touching the weights again — reweighting a distribution whose synthesizers have no input does nothing. ## Cost implications Multi-hop synthesis is not free relative to single-hop: it draws more context into the prompt and asks the model to do more. Increasing the multi-hop share increases the per-sample generation cost. This is usually the smaller half of the bill compared with graph transforms over a large corpus, but it scales with `testset_size` rather than corpus size, so for a small corpus and a large testset it can dominate. ## Version note These three class names belong to the knowledge-graph generation line. Very early ragas releases had a different set of synthesizer names and a ratio-based API; code written against those does not port, and the mix concept is the only thing that carries over.
- You asked for 40% multi-hop and the output has almost none. What do you check first?The graph, not the weights. Multi-hop synthesizers need pairs of connected nodes, so if the transform pipeline had no relationship builder — or its similarity threshold drew almost no edges — there is nothing eligible to sample from and the request is silently unfulfillable. Confirm by grouping the generated rows by `synthesizer_name`, then inspect the graph's relationship count.
- Why not just generate 100% multi-hop questions since they are the hard ones?Because a suite that is uniformly hard stops discriminating too. If every candidate pipeline scores low, you cannot tell an improvement from noise, and you lose the ability to detect a regression in the basics — a retriever that suddenly fails simple lookups would still look the same on an all-hard suite. You want both floors and ceilings represented.
- What is the difference between the specific and abstract multi-hop synthesizers?Both draw two connected nodes, but they connect them differently. The specific one leans on concrete shared items such as overlapping entities, producing pointed factual questions spanning two chunks. The abstract one works at the summary or theme level, producing broader conceptual questions that generalise across the pair. Abstract questions stress synthesis; specific ones stress evidence assembly.
- Do the weights have to sum to exactly 1?Treat that as the contract — the list is a probability distribution over synthesizers, and ragas' own default distribution is built that way. Passing weights that do not sum to 1 is not something to rely on behaving sensibly; normalise them yourself before passing them in.
saying these in an interview costs you the question
- Assuming the requested weights are always realised in the output
- Thinking multi-hop just means longer questions rather than two connected nodes
- Generating an all-easy suite and concluding the retriever is excellent
- Believing weights can fix a graph that has no relationships
- Quoting synthesizer names from the pre-knowledge-graph ragas API