What does DeepEval's Synthesizer.generate_goldens_from_docs actually do?
answer
- it is a pipeline, not one call
- chunks become contexts become questions
- count the model calls per golden
- one method skips the document stage
- context ends up on the golden
basics
~20 sIt reads the documents you point it at, splits and groups them into contexts, then asks an LLM to write synthetic inputs grounded in each context, rewrites them to be harder, and returns Golden objects — optionally with a generated expected_output.
solid answer
~50 s`Synthesizer.generate_goldens_from_docs(document_paths=[...])` runs a multi-stage pipeline. It loads each file, splits it, and assembles groups of related chunks into *contexts*, filtering out low-quality ones with a critic model — those knobs live on `ContextConstructionConfig` (`chunk_size`, `chunk_overlap`, `max_contexts_per_document`). For each surviving context it prompts a model to write synthetic inputs that the context can answer, then applies evolutions to make them harder, and with `include_expected_output=True` generates a reference answer too. The return value is a list of `Golden`s whose `context` field is the source chunks, which is exactly what reference-requiring metrics need. The cost is the thing to say out loud: every stage is an LLM call, so a few hundred goldens is thousands of calls. If you already have your own chunks, `generate_goldens_from_contexts` skips the document stage entirely; `generate_goldens_from_scratch` skips documents *and* contexts and invents inputs from a described task.
code
python · 18 linesfrom deepeval.synthesizer import Synthesizer
from deepeval.synthesizer.config import ContextConstructionConfig
synthesizer = Synthesizer(model="gpt-4o-mini")
goldens = synthesizer.generate_goldens_from_docs(
document_paths=["handbook.pdf", "faq.txt"],
include_expected_output=True,
max_goldens_per_context=2,
context_construction_config=ContextConstructionConfig(
critic_model="gpt-4o",
max_contexts_per_document=3,
chunk_size=1024,
chunk_overlap=0,
),
)
synthesizer.save_as(file_type="json", directory="./synthetic_data")go deeper
Know the method name, that you pass document_paths, and that what comes back is a list of Goldens rather than finished test cases. Being able to name the three generation entry points is enough at this level.
Walk the pipeline: load, chunk into contexts, filter, generate inputs, evolve, filter, optionally write expected_output. Name ContextConstructionConfig and include_expected_output and say what each changes.
Lead with cost and fidelity. Do the call arithmetic out loud, split cheap generator from stronger critic, and explain why generating from your own retriever's chunks beats letting the synthesizer re-chunk the corpus.
Own generation as a budgeted, reviewed event rather than a build step. Decide how often a corpus change justifies regeneration, who signs off on generated references, and how you keep results comparable across regenerations.
## The three entry points The `Synthesizer` has three generation methods and they differ in how much of the pipeline they run: - `generate_goldens_from_docs(document_paths=[...])` — start from files on disk. Runs every stage. - `generate_goldens_from_contexts(contexts=[[...], [...]])` — you supply the grouped chunks yourself; the document stage is skipped. - `generate_goldens_from_scratch(num_goldens=...)` — no documents at all. Inputs are invented from a described task, so the goldens carry no `context`. Knowing which one to reach for is most of the interview answer, because they answer different questions. From-docs is for grounding a suite in a corpus you own. From-contexts is for making the synthetic data match the chunks your *production* retriever actually returns. From-scratch is for exercising an application that has no corpus at all — a classifier, a rewriter, a general assistant. ## What from_docs does, stage by stage **1. Load.** Each path in `document_paths` is read; the loader handles common document formats. This is a plain file read, not a crawl — the Synthesizer will not follow links or pull from your vector store. **2. Split and group into contexts.** The text is split, embedded, and related pieces are grouped into a *context*: a small list of chunks that together support answering something. `ContextConstructionConfig` exposes the mechanical knobs — `chunk_size`, `chunk_overlap`, `max_contexts_per_document`, `critic_model`, and quality thresholds. Note the boundary: these parameters control how the *synthesizer* reads your corpus for generation purposes; how you chunk documents for your production retriever is a separate design decision with its own tradeoffs. **3. Filter contexts.** A critic model scores candidate contexts and weak ones are dropped, so you do not generate questions from a page of boilerplate, a table of contents, or a legal footer. **4. Generate inputs.** For each surviving context, a model writes synthetic inputs that the context can answer. `max_goldens_per_context` bounds how many per context. **5. Evolve.** Each input is rewritten one or more times to raise difficulty — adding reasoning steps, forcing multiple contexts to be combined, adding constraints. This stage is configurable and is where a shallow lookup question becomes something worth testing. **6. Filter inputs.** A critic model scores the generated inputs and low-quality ones are regenerated rather than shipped. **7. Optionally generate expected_output.** With `include_expected_output=True`, a model also writes the reference answer from the context. Without it, goldens carry only `input` and `context`. The result is a list of `Golden` objects, also available as `synthesizer.synthetic_goldens`, exportable with `synthesizer.to_pandas()` or `synthesizer.save_as(file_type="json", directory=...)`. ## What it costs This is the question a senior interviewer is really asking. Count the calls: context filtering, input generation, each evolution, input filtering with retries, and optionally the expected output. A single golden can easily cost five to ten model calls, plus embeddings for the chunking stage. Generating a thousand goldens is not a thousand calls — it is closer to five or ten thousand, and if you point the synthesizer at a frontier model for every stage the bill is real. The usual shape is a cheap model for generation and a stronger `critic_model` for the filtering stages, because filtering is where judgment matters and generation is where volume lives. Generation is also a one-off: goldens are static, so you pay once and replay the dataset forever. Treat regeneration as a deliberate act, not something CI does on every run. ## Why from_contexts is often the better call When you generate from docs, the synthesizer chunks the corpus its own way. Your production retriever chunks it a different way. The resulting goldens are grounded in text your retriever may never return as a unit, so the evaluation measures a system slightly different from the one you ship. Feeding your own retriever's chunks into `generate_goldens_from_contexts` removes that mismatch, at the cost of doing the chunking yourself. ## Failure modes Documents with heavy tables, scanned images or code produce weak contexts and therefore weak questions. A corpus with duplicated boilerplate produces duplicated goldens. And an unfiltered run over a large corpus will happily generate hundreds of near-identical lookup questions that all pass, which feels like coverage and is not. Cap the run, read a sample by hand, and keep only what discriminates.
- When would you use generate_goldens_from_contexts instead of generate_goldens_from_docs?When you want the synthetic questions grounded in exactly the chunks your production retriever returns. Generating from docs makes the synthesizer chunk the corpus its own way, so the goldens may be anchored to text your retriever never surfaces as a unit — the evaluation then measures a slightly different system than the one you ship. Passing your own chunk groups removes that mismatch, at the cost of building them yourself.
- Your from_docs run produced 400 goldens and cost far more than you expected. Where did the money go?Every stage is an LLM call: context filtering, input generation, one call per evolution, quality filtering with retries, and the expected output. That is commonly five to ten calls per golden plus embeddings, so 400 goldens is a few thousand calls. The usual fix is a cheap model for generation with a stronger critic model reserved for filtering, plus capping contexts per document and goldens per context.
- What does include_expected_output=True change about the resulting goldens?It adds a generated reference answer to each golden, written by a model from the same context, which unlocks metrics that compare against a reference. The catch is provenance: that reference was written by an LLM, not vouched for by a human, so treat it as a draft to review before it becomes the standard your application is judged against. Without the flag, goldens carry input and context only.
- Would you regenerate the synthetic dataset on every CI run?No. Goldens are static assets; regenerating them means every run scores against a different suite, so a score change tells you nothing about whether the application changed. Generate deliberately, review, store the result, and replay it. Regeneration is an explicit event — new corpus, new product surface — and it should be reviewed like any other change to the test suite.
saying these in an interview costs you the question
- Thinks it is a single prompt rather than a multi-stage pipeline
- Assumes generation is cheap because it is 'just one call'
- Regenerates the synthetic dataset on every CI run
- Believes it can pull documents from a vector store or URL
- Does not know from_contexts and from_scratch exist