In DeepEval's Synthesizer, what do evolutions do to a generated input?
answer
- first drafts are lookup questions
- rewrite passes, not one-shot
- one type broadens instead of deepening
- each pass is another model call
- too many passes leave the context behind
basics
~20 sEvolutions are rewrite passes that make a freshly generated question harder — adding reasoning steps, forcing several contexts to be combined, adding constraints or hypotheticals. Each pass is another LLM call, and too many push the question away from its source context.
solid answer
~50 sAfter the Synthesizer writes a first-draft input from a context, it rewrites it through *evolutions* — a technique borrowed from Evol-Instruct. `EvolutionConfig` controls this with `num_evolutions` (how many rewrite passes each input goes through) and `evolutions`, a dict mapping `Evolution` members to sampling weights. The members are `REASONING`, `MULTICONTEXT`, `CONCRETIZING`, `CONSTRAINED`, `COMPARATIVE`, `HYPOTHETICAL` and `IN_BREADTH`. The point is that a first-draft synthetic question is usually a lookup — its answer is one sentence of the context restated — and a suite of those cannot distinguish a good RAG system from a mediocre one. Evolving toward `REASONING` and `MULTICONTEXT` produces questions that need synthesis across chunks. The cost is linear: each evolution is an extra call per golden. The risk is over-evolution, where the rewritten question drifts far enough from the context that the stored reference answer no longer follows from it, and your metric then penalizes a correct application.
code
python · 20 linesfrom deepeval.synthesizer import Synthesizer
from deepeval.synthesizer.config import Evolution, EvolutionConfig
synthesizer = Synthesizer(
model="gpt-4o-mini",
evolution_config=EvolutionConfig(
num_evolutions=2,
evolutions={
Evolution.MULTICONTEXT: 0.5,
Evolution.REASONING: 0.3,
Evolution.CONSTRAINED: 0.1,
Evolution.IN_BREADTH: 0.1,
},
),
)
goldens = synthesizer.generate_goldens_from_docs(
document_paths=["handbook.pdf"],
include_expected_output=True,
)go deeper
Know that the synthesizer rewrites its first-draft questions to make them harder, and that this is configurable. Being able to say 'evolutions add reasoning or force multiple chunks' is enough here.
Name EvolutionConfig, num_evolutions and the evolutions weight dict, and describe several Evolution members including what MULTICONTEXT and IN_BREADTH each change. Explain why unevolved questions are mostly lookups.
Show the tuning loop: check the score distribution, weight toward the failure mode you care about, and sample-read evolved questions against their contexts to catch over-evolution before it poisons the suite.
Own the difficulty target. Decide what separation between a good build and a deliberately degraded one makes the suite worth gating on, and budget the generation spend that difficulty implies.
## Why a first draft is not good enough Give a model a paragraph and ask it for a question, and you reliably get something whose answer is a sentence of that paragraph, lightly reworded. Score a RAG application on a hundred of those and almost everything passes — including a broken retriever, as long as it happens to return the right chunk, and including a weak generator, because the task is extraction rather than reasoning. The suite has no discriminating power. Evolutions exist to fix that. Each evolution is a prompted rewrite of the input that makes it demand more of the system under test. ## The evolution types The `Evolution` enum in `deepeval.synthesizer.config` provides: - **REASONING** — rewrite so the answer requires inference rather than lookup. - **MULTICONTEXT** — rewrite so answering requires information from more than one context, which stresses retrieval breadth directly. - **CONCRETIZING** — replace vague terms with specific ones, so the question is unambiguous and precise. - **CONSTRAINED** — add a condition or restriction the answer must respect. - **COMPARATIVE** — turn it into a comparison between two things in the corpus. - **HYPOTHETICAL** — pose a what-if scenario built on the facts. - **IN_BREADTH** — broaden the topic, producing a related but different question rather than a harder version of the same one. The first six deepen a question; `IN_BREADTH` diversifies the set. That distinction matters when you are choosing weights: if your complaint is "my suite is too easy", you want REASONING and MULTICONTEXT. If your complaint is "my suite asks the same thing forty times", you want IN_BREADTH. ## Configuring it `EvolutionConfig` takes `num_evolutions` — how many rewrite passes each input receives — and `evolutions`, a dict of `Evolution` member to weight, from which the pass type is sampled. Pass the config to the `Synthesizer` constructor as `evolution_config`. Weights need not be uniform: a RAG suite that keeps failing on cross-document questions is well served by loading MULTICONTEXT heavily. ## The cost arithmetic Evolutions are the most direct cost lever in the whole pipeline, because they multiply. `num_evolutions=3` over 500 goldens is 1,500 rewrite calls on top of generation, filtering and reference-answer calls. This is one of the places where a cheap generation model pays for itself: the rewrite is a mechanical transformation, not a judgment call, so it rarely needs a frontier model — whereas the critic doing quality filtering does. ## Over-evolution: the failure mode to name Every rewrite pass moves the question further from the context that grounds it. After enough passes you get questions that: - **Cannot be answered from the corpus at all**, because the hypothetical or the added constraint introduced facts nobody wrote down. The application correctly says "I don't know", and your suite records a failure. - **Have a stored reference answer that no longer follows.** If the reference was generated before or alongside the rewrite, an aggressively evolved question can end up paired with an answer to a subtly different question. Every metric that compares against that reference is now measuring noise. - **Are ambiguous**, because layered constraints made the intended reading unclear. Two reasonable answers exist and the metric arbitrarily rewards one. This is why the filtration stage matters and why a human sample-read after generation is not optional. A practical habit: pull twenty evolved goldens at random, read the question next to its `context`, and confirm you could answer it yourself from that context alone. If you cannot, `num_evolutions` is too high for this corpus. ## How to tune it Start low — one or two evolutions — and check the resulting score distribution. If your current application scores near the ceiling on everything, the suite is too easy and evolution weight toward REASONING and MULTICONTEXT is the lever. If scores are near the floor and manual inspection shows the questions are unanswerable rather than hard, you have over-evolved. The target is a suite where a good build clearly separates from a deliberately degraded one — that separation, not the absolute number, is what tells you the dataset is doing its job. One last boundary worth stating in an interview: evolution shapes the *input*. It does not make your metric more reliable, and it does not fix a reference answer nobody reviewed. It makes the questions harder, which is a necessary but not sufficient condition for a suite that catches regressions.
- Which evolution types would you weight up for a RAG system that fails on cross-document questions?MULTICONTEXT first — it explicitly rewrites the input so answering requires information from more than one context, which is precisely the retrieval breadth that is failing. REASONING pairs well with it, because combining chunks usually also requires inference rather than extraction. Keep some weight on the others so the suite does not become monotonous, and verify by sampling that the evolved questions really do need two chunks.
- How does IN_BREADTH differ from the other evolution types?The others deepen the same question — more reasoning, more constraints, more contexts. IN_BREADTH broadens the topic instead, producing a related but different question. So it is the lever for diversity rather than difficulty. If your generated suite asks forty variations of the same lookup, IN_BREADTH helps; if it is simply too easy, REASONING and MULTICONTEXT are what you want.
- You raised num_evolutions to 5 and your scores collapsed. How do you tell whether the app regressed or the dataset did?Read a random sample of the evolved goldens next to their context and try to answer each yourself from that context alone. If you cannot, the questions became unanswerable and the dataset is at fault. Cross-check by replaying the previous, less-evolved dataset against the same build: if the old suite still passes, nothing regressed in the application and the new goldens need filtering or fewer passes.
saying these in an interview costs you the question
- Thinks evolutions improve the metric's accuracy rather than the question's difficulty
- Sets num_evolutions high on the assumption that harder is always better
- Cannot name any evolution type beyond 'makes it harder'
- Ignores that each evolution multiplies generation cost
- Never rereads evolved questions against their source context