In Ragas, what do RunConfig's max_workers, timeout and max_retries control?
answer
- knobs for the executor, not the metrics
- concurrency, per-call deadline, backoff
- the default worker count is 16
- timeout bounds one call, not the run
- under throttling, the safe direction is down
basics
~20 sRunConfig tunes how a Ragas run executes: max_workers caps how many metric calls run concurrently (default 16), timeout bounds a single call (default 180 seconds), and max_retries is how many times a failed call is retried with backoff (default 10). Pass it as evaluate(run_config=RunConfig(...)).
solid answer
~50 s`evaluate()` builds one job per sample-and-metric pair and pushes them through a bounded executor, and `RunConfig` is the knob set for that executor. `max_workers` (default 16) caps in-flight judge calls, which is effectively your request rate against the provider. `timeout` (default 180 seconds) applies to a single call, not the run. `max_retries` (default 10) plus `max_wait` (default 60) govern the exponential backoff when a call raises — a provider 429 or 5xx is retried rather than failing the sample immediately. You pass it in as `evaluate(dataset, metrics, run_config=RunConfig(max_workers=4))`. The tuning instinct that matters: when a run is slow *because* the provider is throttling you, raising `max_workers` makes it worse — it converts throughput into retries, each of which is another billable judge call. Lower the concurrency to sit under the provider's limit and the wall-clock time usually drops.
code
python · 16 linesfrom ragas import evaluate
from ragas.run_config import RunConfig
run_config = RunConfig(
max_workers=4,
timeout=120,
max_retries=3,
max_wait=30,
)
result = evaluate(
dataset=dataset,
metrics=metrics,
llm=evaluator_llm,
run_config=run_config,
)go deeper
Know that Ragas has a RunConfig passed to evaluate(), and that it holds concurrency, a per-call timeout and retry settings rather than anything about the metrics themselves.
Explain the defaults and what each one bounds — workers cap in-flight judge calls, timeout bounds a single call, retries implement backoff — and that these are the settings you touch when a run is slow or flaky.
Show the operational judgment: diagnose throttling before turning knobs, size concurrency under the provider's limit, keep evaluation off the production key, and recognise that a tight timeout silently costs you scored rows.
Own the isolation and budget stance — evaluation traffic gets its own credentials and quota so a run cannot degrade the product, and retry plus concurrency policy is treated as a cost control, not just a reliability setting.
## Where RunConfig sits A Ragas evaluation is a large fan-out of independent LLM calls. Fifty samples scored on four LLM-based metrics is two hundred scoring jobs, and several of those metrics issue more than one model call per sample. `RunConfig` is the object that decides how aggressively that fan-out is pushed and how failures inside it are handled. It is passed per run: `evaluate(dataset=ds, metrics=metrics, run_config=RunConfig(...))`. ## The four settings that matter **max_workers (default 16).** The ceiling on concurrently executing jobs. In practice this is your concurrent request count against the judge provider. Higher finishes sooner *only while the provider keeps up*. **timeout (default 180 seconds).** The bound on one call, not on the run. A run of a thousand samples does not have a deadline; each individual scoring call does. Raise it when a slow self-hosted judge legitimately needs longer; lowering it is how you stop a wedged call from occupying a worker slot for three minutes. **max_retries (default 10).** How many times a raising call is retried before the job is abandoned. Combined with `max_wait` (default 60 seconds) this implements exponential backoff with a cap, so a transient 429 or 503 does not cost you the sample. **exception_types.** Which exceptions are treated as retryable at all. Widening it retries everything, including genuine bugs in a custom metric; narrowing it makes real failures surface fast instead of being retried ten times. ## The rate-limit trap The single most common production mistake is treating a slow run as a concurrency problem. The reasoning goes: the run takes forty minutes, so raise `max_workers` from 16 to 64. What happens instead is that the provider starts returning 429s, every 429 is retried with backoff, the retries themselves consume tokens and queue slots, and the run gets slower *and* more expensive. The throughput ceiling was never the executor — it was the provider's requests-per-minute and tokens-per-minute allowance. The correct move is to size `max_workers` to fit under the provider's limit, ideally with headroom for the rest of your organisation using the same key. If the judge model is a shared production key, an evaluation run can throttle live traffic; a separate key or a separate project for evaluation is the usual defence. ## Retries are not free Each retry is another billable judge call. A run configured with generous retries and aggressive concurrency against a rate-limited endpoint can bill several times what a well-sized run costs for exactly the same scores. Retry settings are as much a cost knob as a reliability knob, and this is the observation an interviewer is listening for. ## Timeout interacts with metric shape Some metrics decompose the work into several sequential model calls per sample — extract claims, then verify each. A per-call timeout of 180 seconds does not mean a sample completes in 180 seconds; it means each of the calls behind that sample gets up to 180. A long run with a slow judge should be reasoned about in calls, not samples. ## When a job runs out of road If every retry is exhausted, or the call times out, the job does not by default blow up the whole run. That sample's score is recorded as missing and the run continues, which is why a run configured with a too-small timeout can complete quickly and look successful while a large fraction of its rows never scored. Concurrency, timeout and retry settings therefore have a direct effect on data completeness, not just on speed. ## Sensible defaults for real situations - **Hosted commercial judge, shared key** — reduce `max_workers` well below the default and accept a longer run; the provider's limit is the real constraint. - **Self-hosted or local judge model** — concurrency is bounded by your own GPU or CPU; a high `max_workers` just queues internally and inflates per-call latency until timeouts start firing. Match workers to what the server can genuinely serve in parallel and raise the timeout. - **Debugging a new metric** — small worker count, small retry count, so failures surface immediately instead of being retried into a slow, quiet run. ## The one-line summary `RunConfig` is where the run's speed, its bill and its completeness are all decided at once, and the counterintuitive part is that the safe direction under throttling is *down*.
- Your evaluation run takes 40 minutes and the provider is returning 429s. Do you raise or lower max_workers?Lower it. The 429s mean the provider, not the executor, is the bottleneck; more workers just produce more rejected requests, each retried with backoff and each billable. Sizing concurrency to sit under the provider's requests-per-minute allowance usually reduces wall-clock time as well as cost. If the run must go faster after that, the answer is a separate key or a cheaper judge, not more workers.
- How does a low timeout change your results rather than just your runtime?A call that exceeds the timeout is retried and, if the retries are exhausted, that sample's score is recorded as missing rather than failing the run. So a timeout tuned too tight produces a run that finishes fast and prints a plausible average computed over only the samples that survived. Timeout is a data-completeness setting, not only a latency setting.
- Why is running an evaluation against your production judge API key risky?An evaluation fans out hundreds of concurrent calls and consumes the same requests-per-minute and tokens-per-minute allowance as live traffic. A large run can throttle the user-facing application, and retries amplify the pressure. Evaluation should use its own key, project, or provider account so a bad run degrades a report rather than the product.
saying these in an interview costs you the question
- Raising max_workers to fix slowness caused by provider throttling
- Thinking timeout applies to the whole run rather than one call
- Treating retries as free rather than as extra billable judge calls
- Running a large evaluation on the production judge API key
- Assuming exhausted retries abort the run loudly