How would you keep a DeepEval suite's judge cost and runtime viable on every pull request?
answer
- cases times metrics times calls
- scope is the biggest lever
- parallelism ends at the provider's limit
- the cache key includes the output text
- log what the run was produced under
basics
~20 sDecide what the merge path must prove, then run only that: a small marked subset per pull request with parallel processes and the result cache on, the full sweep on a schedule, cheaper judges for breadth and expensive ones for the gate, and hyperparameters logged so runs stay comparable.
solid answer
~60 sTreat it as a budget allocation, not a tuning exercise. Every case times every metric is a judge call, so a 500-case suite with three metrics is 1,500 model calls per run — multiply by pull requests per day and the number is a real line item. The levers DeepEval gives you are: `-n` to spend wall-clock time in parallel (bounded by the provider's rate limits, not by your cores); `-m` to run a small representative subset on pull requests and the full set nightly; `-c` to reuse cached metric results, which pays off only when the generated outputs are unchanged, since the cache key includes the output text; a cheaper judge model on the broad sweep with a stronger one on the gating subset; and `@deepeval.log_hyperparameters` or `evaluate(hyperparameters=...)` so every run records the model, prompt, and settings it was produced under and two runs remain comparable. What you gate on is the real decision: proving no regression on twenty high-signal cases beats proving nothing on five hundred nobody waits for.
code
python · 10 linesimport deepeval
@deepeval.log_hyperparameters
def hyperparameters():
return {
"model": "gpt-4o-mini-2024-07-18",
"prompt_version": "v7",
"chunk_size": 500,
}go deeper
Know that every metric is a paid model call, so the size of an eval suite is a cost decision. Running the whole dataset on every commit is not free.
Explain the concrete knobs — process count, marked subsets, the result cache and what invalidates it — and why the cache rarely helps when outputs are regenerated each run.
Show that you size parallelism against the provider's rate limit, keep the gating tier small and strict, and log hyperparameters and run identifiers so a score delta can be attributed to a change.
Own the framing: separate the merge-path gate from the quality-tracking sweep, give each its own scope, judge, thresholds and owner, and justify the spend against what each tier actually prevents.
## Frame the number first Cost here is not incidental. Each metric is an LLM-as-judge call, often more than one per metric, and metrics multiply by cases. Before optimising, write the arithmetic down: cases times metrics times calls-per-metric times runs-per-day times price-per-call, plus the wall-clock time a developer waits. Teams that skip this step end up with a suite that is quietly expensive and, worse, one that is slow enough that people stop waiting for it and start merging around it. ## Lever 1: what runs on the merge path The highest-leverage decision is scope, not configuration. A pull-request gate exists to catch regressions on things you have decided must not break. That is usually a few dozen carefully chosen cases, not the whole corpus. Mark them and select with the runner's `-m` flag; run the full dataset on a schedule, where latency is nobody's problem and a failure opens an investigation rather than blocking a merge. A related choice is metric count per case. Three metrics on every case triples the bill; often one gating metric per case plus a broader multi-metric sweep is the better allocation. ## Lever 2: parallelism, bounded by the provider `deepeval test run -n N` distributes pytest nodes across processes, and since eval tests are almost pure network wait, throughput scales well — up to the point where the judge provider rate-limits you and metrics start erroring. The ceiling is the provider's limit divided by the calls each worker makes, not the machine's core count. Raise `-n` until errors appear, then back off; if you are also calling `evaluate()` with its own concurrency inside those workers, the two multiply and you will find the limit sooner than you expect. ## Lever 3: caching, and what actually invalidates it DeepEval writes cached metric results to a JSON file under a hidden directory (the folder is overridable via `DEEPEVAL_CACHE_FOLDER`), and `-c` opts into reading it. Understanding the key is what makes caching useful rather than mysterious: it is built from the test case's content — input, actual output, expected output, context, retrieval context — plus the logged hyperparameters, and a cached metric result is reused only when the metric's configuration matches too (threshold, evaluation model, strict mode, criteria, include-reason, and so on). The consequence is the important part. If your test generates a fresh output from the model each run, the output text differs and the cache misses — caching does not make a generation-plus-evaluation suite cheap. It pays when you re-run over *recorded* outputs: iterating on metric wiring, re-running after an unrelated test fix, or evaluating a fixed snapshot. Cache also has to persist between pipeline runs to help at all, which means restoring the directory from the runner's cache store, and it is deliberately off during repeats. ## Lever 4: judge choice per tier Judge model choice is a cost decision made per tier, not per project. A smaller, cheaper model is often adequate for a broad nightly sweep whose job is to surface a trend; a stronger judge is worth the money on the twenty cases that gate a merge. Whatever you choose, pin the version, because a judge that moves on its own makes every historical score incomparable — and re-baselining a suite is more expensive than any of these savings. ## Lever 5: keeping runs comparable A cost programme is only meaningful if you can tell whether the numbers moved for a good reason. DeepEval records hyperparameters with a run: in 4.1.9 you decorate a zero-argument function returning a dict with `@deepeval.log_hyperparameters`, or pass `hyperparameters=` to `evaluate()`. Recording the application model, the prompt, chunking settings and anything else you might change means a score delta can be attributed rather than guessed at. Pair that with a run identifier (`-id`) per commit so a build maps to a reviewable run. ## The judgment call All of these are secondary to one question the interviewer is really asking: what is the eval suite's job on the merge path? If the answer is "prove this change did not break the five behaviours we promise customers", the suite is small, fast, strictly gated, and cheap. If the answer is "track quality over time", it belongs on a schedule with a report and an owner, not in front of a merge button. Most expensive, flaky eval suites are the result of never having separated those two jobs. ## What not to do Do not shrink the gate by lowering thresholds — you keep the cost and lose the signal. Do not skip the suite on "small" changes by path filter alone; prompt and retrieval changes rarely announce themselves in the diff. And do not let cost pressure push you into evaluating with the same model that generated the output without acknowledging what that does to the score's meaning.
- Why does enabling the DeepEval cache often save nothing in a real pull-request pipeline?Two reasons. The cache key includes the test case's actual output, so a suite that regenerates the output each run misses on every case. And the cache file lives in a local directory, so unless the pipeline restores it between runs there is nothing to hit. Caching pays for repeated evaluation of fixed, recorded outputs.
- How do you decide the process count for the runner?By the judge provider's rate limit, not the machine. Estimate the calls per worker per second, divide the account's limit by that, and set `-n` below the result; then watch for metric errors, which are the observable symptom of exceeding it. Remember any concurrency inside a test multiplies with the process count.
- How does DeepEval record what a run was produced under?Through hyperparameters. In 4.1.9 you decorate a zero-argument function that returns a dict with `@deepeval.log_hyperparameters`, and the returned values are attached to the test run; `evaluate()` takes the same thing as a `hyperparameters=` argument. Values may include Prompt objects, so the exact prompt version travels with the run and score deltas become attributable.
- Is running a cheaper judge on the nightly sweep and a stronger one on the gate defensible?Yes, provided you never compare their scores directly. Each tier has its own baseline and its own thresholds; a trend from the cheap judge is a signal to investigate, not a number to place beside the gate's. Pin both versions, and re-baseline the affected tier whenever either changes.
saying these in an interview costs you the question
- Runs the full dataset on every pull request by default
- Assumes caching makes a generate-and-evaluate suite cheap
- Sets the process count from CPU cores rather than rate limits
- Cuts cost by lowering thresholds
- Compares scores from two different judge models directly