skip to content

How do you keep DeepEval's synthetic goldens from being too easy?

level: seniorimportance: should knowfreq 55%

answer

  1. scoring in the nineties is a warning sign
  2. questions written from the chunk cannot miss
  3. real users do not write exam questions
  4. filter with a stronger model than you generate with
  5. break the app on purpose and re-run

basics

~20 s

Generate from the chunks your production retriever really returns, weight evolutions toward reasoning and multi-context, use StylingConfig so inputs read like real user messages, set a strong critic model in FiltrationConfig, and sample-read the output before it becomes the standard.

solid answer

~50 s

Default synthetic goldens are extraction questions: a model saw one context and wrote a question that context answers in a sentence. A suite of those passes even on a degraded build, so it catches nothing. Four levers actually move it. First, source: feed `generate_goldens_from_contexts` the chunks your own retriever returns, so the questions test the system you ship rather than the synthesizer's re-chunking. Second, difficulty: weight `EvolutionConfig` toward `Evolution.REASONING` and `Evolution.MULTICONTEXT`. Third, register: `StylingConfig` lets you describe the `task`, the `scenario`, and the `input_format`, so generated inputs read like terse, typo-ridden, multi-intent real user messages instead of well-formed exam questions. Fourth, filtering plus human eyes: `FiltrationConfig` takes a `critic_model` and a minimum quality bar with retries, but nothing replaces reading a sample against its context. The empirical check is discriminating power — run the suite against a deliberately degraded build and confirm the scores actually drop.

code

python · 27 lines
python
from deepeval.synthesizer import Synthesizer
from deepeval.synthesizer.config import (
    Evolution,
    EvolutionConfig,
    FiltrationConfig,
    StylingConfig,
)

synthesizer = Synthesizer(
    model="gpt-4o-mini",
    filtration_config=FiltrationConfig(critic_model="gpt-4o"),
    evolution_config=EvolutionConfig(
        num_evolutions=2,
        evolutions={Evolution.REASONING: 0.5, Evolution.MULTICONTEXT: 0.5},
    ),
    styling_config=StylingConfig(
        task="Answering internal policy questions for support agents",
        scenario="Agents checking refund and warranty rules mid-call",
        input_format="Short lowercase message, often misspelled, sometimes two questions at once",
        expected_output_format="Two sentences naming the policy it comes from",
    ),
)

goldens = synthesizer.generate_goldens_from_contexts(
    contexts=[["Refunds are accepted within 30 days of delivery."]],
    include_expected_output=True,
)

go deeper

for a junior

Know that synthetic questions tend to be easy because they are written from the same chunk that answers them, and that someone should read a sample before the dataset is trusted.

for a middle

Name the concrete knobs — EvolutionConfig weights, StylingConfig's task and input_format, FiltrationConfig's critic_model — and explain what each changes about the generated goldens.

for a senior

Demonstrate the loop: generate from your own retriever's chunks, tune difficulty and register, review a sample by hand, then prove discriminating power by scoring a deliberately degraded build alongside the real one.

for a principal

Own what the suite is allowed to certify. Decide the required separation between good and degraded builds, who reviews generated references, and how much of the suite must stay hand-curated from real incidents.

## The failure this question is about You run `generate_goldens_from_docs` over the corpus, get 300 goldens, run your application, and score in the nineties on everything. That is not good news. It usually means the questions are extraction tasks whose answers sit verbatim in a single chunk, phrased in the same vocabulary as the source. Your retriever cannot miss, because the question was written from the chunk. Your generator cannot fail, because the task is restatement. The suite will report success right through a regression. Making synthetic data useful is therefore a deliberate act, not a default. ## Lever 1 — where the contexts come from The synthesizer's own splitting is a convenience, not a model of your system. If production retrieval returns 400-token chunks with heavy overlap, and the synthesizer generated from 1,024-token contexts, the goldens are anchored to text your retriever never surfaces as one unit. Passing your retriever's actual chunk groups into `generate_goldens_from_contexts` closes that gap. This is the single highest-leverage change for a RAG suite, and it costs nothing extra at generation time. ## Lever 2 — difficulty via evolutions Weighting `EvolutionConfig` toward `Evolution.REASONING` and `Evolution.MULTICONTEXT` turns lookups into synthesis. Raise `num_evolutions` cautiously: past two or three passes on most corpora, questions start leaving their grounding behind, and an unanswerable question is a false failure, not a hard test. ## Lever 3 — register, via StylingConfig This one is routinely skipped and it matters more than people expect. LLM-generated questions are well-formed, complete sentences with full context and correct spelling. Real users write "refund window??" and "can i cancel + also change address". Your application may handle the first perfectly and fall apart on the second — and a suite made only of the first will never tell you. `StylingConfig` takes plain-English descriptions: `task` (what the application does), `scenario` (who is asking and why), `input_format` (how the input should look), `expected_output_format` (how a good answer should look). Describing the input as "a short, lowercase, sometimes misspelled message typed by a support agent mid-call" changes the generated distribution substantially. ## Lever 4 — filtering and human review `FiltrationConfig` takes a `critic_model` and a quality bar, with retries when a generated input scores below it. Use a stronger model here than for generation: filtering is a judgment task and generation is a volume task, so the money is better spent on the critic. But filtering is a model checking a model. The irreducible step is a human sample-read: pull twenty goldens, read each question beside its `context`, and answer three things — could I answer this from this context? Is this a question a real user would ask? Would a mediocre system get it wrong? A golden that fails the third question is filler. The same scrutiny applies harder to generated `expected_output`. It was written by a model from the same context, so it inherits that model's blind spots — and once it is stored, it becomes the standard your application is judged against. Review references before you trust any metric that compares against them, or keep the goldens reference-free and rely on metrics that do not need one. ## The check that settles it: discriminating power Whatever you do to the generation config, the dataset is only worth gating on if it can tell good from bad. Build a deliberately degraded variant of your application — return three random chunks instead of the top-k, or swap in a much weaker generator — and run both against the suite. If the scores barely move, the suite is measuring nothing regardless of how sophisticated the generation config was. If the degraded build drops clearly and the good build holds, you have a dataset worth keeping. Run the same check when you regenerate. A new synthetic batch is a new measuring instrument, and instruments need calibrating against a known-bad reference before you trust their readings. ## Mixing in the real thing The strongest suites are not purely synthetic. Seed them with real inputs — questions from support tickets, logged production queries, the specific cases behind past incidents — and let synthetic generation provide breadth around that spine. Every real failure you have shipped deserves a permanent golden, and that one is curated by hand, forever, no matter what the synthesizer produces.

  • What exactly does StylingConfig change about the generated goldens?
    It steers the register and shape of what gets produced. You describe the task the application performs, the scenario the user is in, the input_format you want the synthetic inputs to look like, and the expected_output_format for references. Describing inputs as terse, lowercase, occasionally misspelled agent messages yields a very different distribution from the default well-formed questions — and that distribution is closer to what your application actually receives in production.
  • How do you decide which model to put in FiltrationConfig's critic_model?
    Filtering is a judgment task on a small number of decisions per golden; generation is a volume task. So the economics favour a cheap model for generation and the strongest one you can justify as the critic. If the critic is weaker than the generator, it will wave through exactly the flawed inputs the generator is prone to producing, which is the worst of both worlds — you pay for filtering and get none.
  • Your synthetic suite scores 0.94 average. Why might that be bad news?
    Because a suite everything passes cannot detect anything. It usually means the questions are extraction tasks written from single chunks in the corpus's own vocabulary, so retrieval cannot miss and generation only has to restate. Confirm by running a deliberately degraded build — random chunks instead of top-k, or a much weaker generator — against the same suite. If the score barely moves, the dataset has no discriminating power and needs harder, more realistic inputs.
  • Should the whole eval dataset be synthetic?
    No. Keep a hand-curated spine of real inputs — logged production queries, support tickets, and a permanent golden for every failure you have actually shipped — and use synthetic generation for breadth around it. Synthetic data is good at volume and coverage of a corpus; it is bad at anticipating the specific weird thing a real user did last Tuesday, which is exactly the case you most need to keep passing.

saying these in an interview costs you the question

  • Treats a high average score on a synthetic suite as proof of quality
  • Trusts model-generated expected_output without any human review
  • Uses the same cheap model for generation and for filtering
  • Never checks the suite against a deliberately degraded build
  • Replaces all curated goldens with generated ones

context