Should DeepEval goldens live in your repo or in Confident AI?
answer
- reproducibility versus collaboration
- who is allowed to change the suite
- a score moved and git explains nothing
- snapshot the shared thing, run the snapshot
- what leaves your network when you push
basics
~20 sIt is a tradeoff between reproducibility and collaboration. Files in the repo are pinned by commit and reviewable in a pull request; a hosted dataset pulled with dataset.pull(alias=...) lets domain experts curate it, but the suite can change under CI without any code change.
solid answer
~50 sDeepEval supports both: `EvaluationDataset` can be built from local files, and `dataset.push(alias=...)` / `dataset.pull(alias=...)` sync it with Confident AI using `CONFIDENT_API_KEY`. Repo-resident goldens give you the properties CI wants — pinned by commit, diffed in review, no network call, no credential, works air-gapped. Hosted goldens give you the properties the *business* wants — a support lead can add cases without a pull request, annotations and review live next to the data, and there is one alias everyone means. The failure mode of hosted-only is silent drift: someone edits the dataset, CI pulls at run time, the score moves, and `git log` explains nothing. My usual answer is a hybrid — curate in the platform, then snapshot to a versioned file that CI actually runs, and treat updating that snapshot as a reviewed change. There is also a data question: pulling means your inputs, references and source-document text sit in a third party's store, which some corpora simply cannot do.
code
python · 11 linesimport os
from deepeval.dataset import EvaluationDataset
# requires CONFIDENT_API_KEY in the environment, or a prior `deepeval login`
assert os.environ.get("CONFIDENT_API_KEY")
dataset = EvaluationDataset()
dataset.pull(alias="support-regression")
print(f"pulled {len(dataset.goldens)} goldens")go deeper
Know that DeepEval datasets can be kept as local files or synced with dataset.push and dataset.pull using an alias and an API key, and that CI needs to know which one it is running.
Contrast the two: a repo file is pinned by commit and reviewed in a diff, a pulled dataset can change between runs. Be able to describe the credential and network dependency that pulling introduces.
Design the flow. Explain live-pull for nightly jobs versus a pinned snapshot for per-PR gating, and how you would diagnose a score move when the dataset itself is a moving part.
Own the invariant — every reported number traceable to a retrievable dataset version — and the residency question of what leaves your network on push. State the conditions under which you would switch models.
## The two options concretely **Repo-resident.** Goldens live as JSON or CSV in the repository next to the code. You construct an `EvaluationDataset`, populate its goldens, and run. The dataset version is whatever the checked-out commit says it is. **Platform-resident.** You call `dataset.push(alias="support-regression")` once, and thereafter `dataset.pull(alias="support-regression")` fetches the current contents at run time, authenticated by `CONFIDENT_API_KEY` (or a prior `deepeval login`). The dataset is edited in the platform UI, by anyone with access. ## What each buys you Repo-resident wins on everything CI cares about: - **Reproducibility.** A commit fully determines the suite. Re-running last month's build gives last month's dataset. - **Review.** A change to the goldens shows up as a diff in a pull request, with an author and a rationale. - **Isolation.** No network call, no secret in the CI environment, no third-party outage in your build path. - **Bisect.** When a score moved, you can attribute it to a code change or a dataset change by looking at history. Platform-resident wins on everything the humans care about: - **Non-engineer access.** The person who knows which support answers are wrong is usually not the person who can open a pull request. - **Annotation workflow.** Reviewing, commenting on and approving goldens is a UI problem, and a text file is a bad UI. - **Single source of truth.** One alias, referenced from every service, instead of four slightly divergent copies. - **Scale.** Ten thousand goldens with metadata is unpleasant as a reviewable diff. ## The drift problem The sharp edge of platform-resident is that your test suite becomes a moving input to a process whose whole purpose is detecting change. CI pulls at run time; someone adds thirty hard goldens on Tuesday; Wednesday's score drops eight points; an engineer spends the afternoon bisecting application code that never changed. Nothing in the repository records what happened. This is not an argument against the platform — it is an argument for the dataset having a version you can name. However you store goldens, you want to be able to answer "which dataset produced this number?" without asking a person. ## The hybrid that usually wins Curate in the platform, snapshot to the repo, run the snapshot: 1. Domain experts add and review goldens in Confident AI. 2. A scheduled or manual job pulls the alias and writes it to a versioned file in the repo. 3. That file update lands as a reviewed pull request — a human sees which goldens arrived. 4. CI runs the file, offline and deterministic. You keep the collaboration surface and CI keeps its pinning. The cost is a lag between someone adding a golden and CI seeing it, which is almost always acceptable — an eval suite that changes hourly is not a suite. A reasonable variant: pull live in a nightly or pre-release job where a moving dataset is *desirable* (you want the newest cases), and run the pinned snapshot on every pull request where determinism matters. ## Data residency Often the deciding factor, and it deserves to be raised early. Pushing a dataset uploads inputs, reference answers, and — for a corpus-derived synthetic set — substantial verbatim excerpts of your source documents in the `context` field. If those documents are customer records, medical text, or anything under a residency or confidentiality obligation, the hosted option is off the table for that dataset regardless of its ergonomics. You can sometimes split: hosted for a synthetic, non-sensitive suite; repo-resident, access-controlled for the sensitive one. ## How to argue it in an interview Do not pick a side flatly. Name the axes — reproducibility, review, collaboration, residency, operational coupling — and then say what you would do and under what conditions you would switch. The strongest version also states the invariant you refuse to give up: every reported evaluation number must be traceable to a specific, retrievable dataset version, whatever the storage mechanism. Any design that satisfies that invariant is arguable; any design that does not is broken no matter how convenient it feels.
- What breaks if CI pulls the dataset live on every pull request?Three things. Determinism: two runs of the same commit can score differently because the dataset changed between them. Attribution: a score move can no longer be blamed on the diff under review. Availability: your build now fails when the platform or the network does, and it needs an API credential in the CI environment. None of these is fatal for a nightly job; all of them are painful on a per-PR gate.
- You go hybrid and snapshot the dataset into the repo. What is the cost?Lag and a second workflow. A golden added by a domain expert does not affect CI until someone refreshes the snapshot and gets that pull request merged, so there is a window where the shared dataset and the tested dataset disagree. You also need to keep the refresh job honest — an unattended snapshot that auto-merges reintroduces exactly the silent drift you were avoiding, so the review step is the part that must not be optimized away.
- What actually leaves your network when you call dataset.push?The goldens themselves: inputs, expected outputs, metadata, and the context field — which for a corpus-derived synthetic set contains verbatim excerpts of the source documents. That last one surprises people. If the corpus is customer data, medical text, or anything with a residency obligation, pushing it is a data transfer decision, not a tooling convenience, and it needs the same approval any other export would.
- How do you make an evaluation number traceable to a dataset version?Record the dataset identity alongside the result: the commit for a repo-resident file, or the alias plus a pulled-at timestamp and a content hash for a hosted one. Store it with the run, not in someone's memory. The invariant is that anyone reading a score months later can retrieve the exact goldens that produced it — without that, comparing two numbers is guesswork, and a regression report cannot be defended.
saying these in an interview costs you the question
- Declares one storage model correct without naming a tradeoff
- Pulls the dataset live on every pull request and calls it reproducible
- Ignores that pushed goldens carry verbatim source-document text
- Cannot say how a past score maps to a specific dataset version
- Assumes domain experts will file pull requests to add test cases