What does num_repetitions do in LangSmith's evaluate(), and what does it cost?
answer
- each example run N times
- one experiment, not N experiments
- cost multiplies, judges included
- spread, not a better average
- small subset, high N
basics
~20 snum_repetitions runs every example N times inside one experiment instead of once, so you can see the spread of a non-deterministic system rather than a single sample. It multiplies both target calls and evaluator calls by N, so cost and wall clock scale linearly with it.
solid answer
~50 s`evaluate(target, data=..., num_repetitions=3)` executes **each example three times within a single experiment**, rather than producing three experiments. The reason to want that is that an LLM application is not deterministic — even at temperature 0, a hosted model can return different text across runs — so one execution per example gives you one sample of a distribution and no idea how wide it is. Repetitions let you see whether an example is reliably good, reliably bad, or flapping, which is usually more actionable than the aggregate. The cost is unambiguous and linear: N repetitions means N times the target calls **and** N times the evaluator calls, so a 500-example run with two LLM judges goes from 1,500 model calls to 6,000 at `num_repetitions=4`, with wall clock rising in step unless you also raise concurrency. That is why teams typically keep repetitions at 1 for the frequent loop and turn them up on a small, high-signal subset when they specifically want to measure variance.
code
python · 14 linesfrom langsmith import evaluate
def target(inputs: dict) -> dict:
return {"answer": "stub"}
evaluate(
target,
data="support-qa-smoke",
num_repetitions=5,
max_concurrency=4,
experiment_prefix="stability-check-prompt-v7",
)go deeper
Know that num_repetitions runs every example more than once inside a single experiment, and that each extra repetition costs another full pass of model calls.
Explain why it exists — one execution samples a stochastic system once — and state the cost multiplier precisely, including the per-evaluator calls that people forget.
Show where you would spend it: a small subset at high N to investigate stability, and the diagnostic that identical outputs with moving scores mean the judge is the unstable component.
Own the policy tradeoff — repetitions buy variance information at linear cost, and a flapping suite is a signal to make the product more deterministic rather than to buy more samples.
## The mechanic `num_repetitions` is an argument on `evaluate` and `aevaluate`. Set it to N and every example in `data` is executed N times, all inside **one** experiment. The runs are separate — each is traced individually with its own output, latency and scores — but they belong to the same experiment record, so the experiment's aggregate is computed across all of them. This is deliberately different from running `evaluate` N times, which would give you N experiments to eyeball side by side. One experiment with repetitions keeps the repeats together as one unit of measurement. ## Why anyone bothers A single execution per example tells you what happened once. LLM applications are stochastic: sampling temperature above zero obviously, but also load-dependent routing, provider-side model updates, and any nondeterminism in retrieval. At `num_repetitions=1`, an example that passes 60% of the time and one that passes 100% of the time look identical in a green run — and the flaky one is the one that will page you. Repetitions convert that invisible property into a visible one. Three executions of the same input let you see agreement: three passes, or two passes and a fail. Aggregated over the dataset, this is the difference between "the suite scored 0.82" and "the suite scored 0.82 and here are the eleven examples that were not stable across repeats". The unstable rows are usually the interesting engineering work — an under-specified prompt, an ambiguous reference, a retrieval step whose ranking flips. ## What it costs Linear, in every dimension, and the multiplication catches people out because it applies to the evaluators too: - target calls: `examples × N` - LLM evaluator calls: `examples × evaluators × N` - wall clock: roughly `× N` at fixed concurrency - storage and trace volume: `× N` A 500-example dataset with two judge metrics is 1,500 model calls at N=1 and 6,000 at N=4. If the judge is an expensive model, the judge line dominates the bill, which is an argument for a cheaper judge before it is an argument against repetitions. ## Where to spend it The usual policy is: repetitions off for the frequent, fast loop, and on for a small deliberate subset. Concretely, a few dozen examples at N=5 costs less than a few hundred at N=1 and tells you something the big cheap run cannot. When you are specifically investigating stability — a new sampling setting, a model swap, a prompt that reviewers say "sometimes" misbehaves — that is the run to make. It is also a diagnostic tool for the eval suite itself. If a suite goes red intermittently and you cannot tell whether the application or the judge is unstable, repetitions separate them: an example whose target output is identical across repeats but whose score moves has an unstable *judge*, not an unstable application. That is a genuinely useful thing to be able to demonstrate. ## What it does not give you Repetitions measure spread. They do not, by themselves, tell you whether the gap between two experiments is real — that is a statistical question about sample size and effect size, and it lives with evaluation methodology rather than with the tool. What the tool gives you is the raw material: repeated observations per example, each individually traceable. They also do not fix a fundamentally unstable system. If half your examples flap, the answer is to make the application more deterministic — constrain the output format, pin the model version, tighten the prompt — not to average over more samples until the number looks calm. Averaging a flapping system produces a stable number describing an unstable product. ## Interview shape A good answer names the mechanic (N executions per example, one experiment), names the cost multiplier including evaluators, and says where you would actually spend it: small subset, high N, when the question is stability. Mentioning the judge-versus-application diagnostic is the detail that reads as having actually used it.
- How does num_repetitions=3 differ from calling evaluate() three times?Three calls produce three separate experiments you must line up by hand; num_repetitions produces one experiment containing three runs per example, aggregated together as a single measurement. The repeats also share the same dataset snapshot and configuration, so nothing can drift between them, which is what makes them a fair look at variance rather than three loosely-related runs.
- Your target output is byte-identical across repeats but the score changes. What does that tell you?The instability is in the evaluator, not the application. An LLM-based judge is itself a model call and will disagree with itself on borderline cases. Repetitions separate the two sources cleanly, and the fix is on the judge side — a tighter rubric, a lower-variance judge model, or a metric that does not need a model at all.
- Half your examples flap across repetitions. Is turning repetitions up the answer?No — that only produces a calm number describing an unstable product. Averaging more samples hides the variance from your dashboard while users still meet it. Reduce the variance at the source: constrain the output format, pin the model version, tighten an ambiguous prompt, or fix a retrieval step whose ranking flips, and use repetitions to confirm the fix landed.
saying these in an interview costs you the question
- Thinking repetitions create separate experiments
- Forgetting that evaluator calls multiply too
- Assuming temperature 0 makes repeats pointless
- Turning repetitions up to smooth away real instability
- Believing repetitions prove a difference is significant