How does LlamaIndex's BatchEvalRunner run evaluators, and where does it bite at scale?
answer
- a dict in, a dict out
- concurrency has a fixed cap
- queries times evaluators equals calls
- two entry points: with or without generation
- positional alignment of extra kwargs
basics
~20 sBatchEvalRunner takes a dict of named evaluators and runs them concurrently over many queries with asyncio, capped by its workers setting. It returns a dict keyed by evaluator name holding one EvaluationResult per row. The cost is judge LLM calls: queries times evaluators.
solid answer
~40 s`BatchEvalRunner` in `llama_index.core.evaluation` exists so you do not hand-roll an async loop over an eval set. You construct it with a dict like `{"faithfulness": FaithfulnessEvaluator(...), "relevancy": RelevancyEvaluator(...)}`, plus `workers` (concurrency cap, default 8) and `show_progress`. Then either `aevaluate_queries(query_engine, queries=[...])`, which runs the pipeline and evaluates the results in one pass, or `aevaluate_responses(queries=[...], responses=[...])` when you already have responses and only want to score them. Results come back as a dict of lists keyed by the names you chose, aligned positionally with the input queries, so aggregating is a comprehension over one key. The scaling pain is arithmetic: N queries times M evaluators means N×M judge calls, and `aevaluate_queries` adds N pipeline runs on top. Raise `workers` and you hit provider rate limits; leave it low and a 500-row set takes a long time.
code
python · 30 linesimport asyncio
from llama_index.core import Document, VectorStoreIndex
from llama_index.core.evaluation import (
BatchEvalRunner,
FaithfulnessEvaluator,
RelevancyEvaluator,
)
from llama_index.llms.openai import OpenAI
judge = OpenAI(model="gpt-4o", temperature=0)
index = VectorStoreIndex.from_documents(
[Document(text="Refunds are accepted within 30 days of purchase.")]
)
query_engine = index.as_query_engine()
queries = ["What is the refund window?", "Can I return an opened item?"]
runner = BatchEvalRunner(
{
"faithfulness": FaithfulnessEvaluator(llm=judge),
"relevancy": RelevancyEvaluator(llm=judge),
},
workers=4,
show_progress=True,
)
results = asyncio.run(runner.aevaluate_queries(query_engine, queries=queries))
for name, rows in results.items():
rate = sum(1 for r in rows if r.passing) / len(rows)
print(name, round(rate, 2))go deeper
Know that it takes a dict of named evaluators and a list of queries, runs them concurrently, and hands back results grouped by the evaluator name you supplied.
Explain the difference between evaluating queries end-to-end and evaluating responses you already have, what workers controls, and how the result dict lines up with the input queries.
Talk about running it in CI: call-count budgeting, tuning workers against provider rate limits, keeping per-row results for triage, and not blocking a release on noise.
Own the economics and the design — how large the eval set must be for a delta to mean something, what a full evaluation costs per release candidate, and whether the sample reflects real traffic.
## The problem it solves Evaluating one response is a single call. Evaluating a regression set is hundreds of independent, IO-bound LLM calls that should overlap. Written by hand that becomes an `asyncio.gather` with a semaphore, per-evaluator bookkeeping, and result alignment. `BatchEvalRunner` is that loop, packaged. ## Construction You pass a dict mapping **your own** names to evaluator instances. The names are arbitrary labels and they become the keys of the result dict, so pick names you want to read in the aggregate report. `workers` bounds how many evaluations run concurrently — it is the knob that trades wall-clock time against provider rate limits. `show_progress=True` prints a progress bar, which matters more than it sounds when a run takes twenty minutes and you need to know it is alive. ## The two entry points `aevaluate_queries(query_engine, queries=[...])` runs the query engine over each query and then evaluates each resulting response. This is the end-to-end path: it measures the pipeline as configured today, and it is what you use when comparing configuration A against configuration B. `aevaluate_responses(queries=[...], responses=[...])` skips generation and scores responses you already hold. Use it when you want to score the same fixed set of answers under several judges, when replaying stored production traces, or simply to avoid paying for generation twice while you iterate on evaluator prompts. Both are coroutines, so they are awaited — `asyncio.run(...)` from a script, or awaited directly inside an existing loop. In notebooks, which already run an event loop, the usual workaround is a library that permits nested loops; forgetting this produces the classic "this event loop is already running" error rather than anything about evaluation. Evaluator-specific keyword arguments (a list of `reference` answers for a correctness evaluator, for example) are passed through as extra keyword arguments and must be the same length as `queries`, positionally aligned. Misalignment is silent and poisonous: every row gets graded against someone else's reference, and the scores look plausibly bad rather than obviously broken. ## Reading results The return value is `Dict[str, List[EvaluationResult]]` — one list per evaluator name, in query order. Aggregation is then trivial: a pass rate is the fraction of results whose `passing` is true; a mean score averages `score`. Keep the per-row results, not just the aggregate — when a pass rate drops from 0.88 to 0.79 the only useful next step is reading the `feedback` on the rows that flipped. ## Where it bites **Cost and rate limits.** The call count is queries × evaluators for scoring, plus queries for generation if you used `aevaluate_queries`, plus embedding calls for retrieval. Three evaluators over 300 questions is 900 judge calls per run — and you run it on every candidate configuration. Teams discover this when a "quick eval" invoices like a training job. **Concurrency versus 429s.** `workers` is a fixed cap, not an adaptive backoff. Set it above what your provider tier allows and the run degrades into retries or errors. The practical approach is to raise it until throughput stops improving, then leave headroom for whatever else shares the API key. **Judge nondeterminism.** Even at temperature 0, LLM judges are not perfectly stable, and a batch aggregates that noise rather than removing it. A one- or two-point move on a 50-question set is not a signal. Either grow the set or repeat the run before you act on a delta. **Sequencing effects.** `aevaluate_queries` hits your real pipeline concurrently, which means your vector store, embedding endpoint and generator all see a burst of load. If your production infrastructure is what you are evaluating against, an aggressive `workers` value can distort latency measurements taken during the same run. ## What it is not It is an execution harness, not a benchmark design. It will not build an eval set, will not hold your judge model constant across runs, and will not tell you whether your sample is representative. Those decisions sit outside the class, and they are what actually determine whether a batch result means anything.
- When would you prefer aevaluate_responses over aevaluate_queries?Whenever generation is not what you are varying. If you want to score one fixed set of answers under several judges or evaluator prompts, `aevaluate_responses` avoids re-running the pipeline each time — you pay for generation once. It is also the path for replaying stored production traces, where the answers already exist and re-running the pipeline would measure today's index rather than what the user actually saw.
- Your eval run starts returning rate-limit errors halfway through. What do you change?Lower `workers` first — it is a fixed concurrency cap with no adaptive backoff, so it will happily saturate your quota. Then consider splitting the run: evaluate responses you generated in an earlier pass so the run makes only judge calls, or shard the eval set across runs. Long term, a separate API key or tier for evaluation keeps a CI eval from stealing quota from production traffic.
- How do you pass reference answers for a correctness evaluator through the batch runner?As an extra keyword argument alongside `queries`, a list positionally aligned with them — element i is the reference for query i. The alignment is the risk: nothing checks that your references are in the same order as your questions, so a shuffled or filtered list silently grades every row against the wrong ground truth and produces scores that look merely bad rather than obviously broken.
saying these in an interview costs you the question
- Thinking it runs evaluators sequentially like a for loop
- Assuming higher workers is always faster
- Forgetting it is async and calling it like a sync function
- Comparing runs whose judge model or eval set changed
- Treating a small-set delta as a proven improvement