skip to content

How do you stop a Ragas re-run from re-paying for identical judge calls?

level: seniorimportance: should knowfreq 28%

answer

  1. repeated runs re-issue identical calls
  2. the cache lives on the model, not the run
  3. the key covers the whole request
  4. warm runs look deterministic — they are replays
  5. the directory holds prompts verbatim

basics

~20 s

Attach a cache backend to the evaluator model — Ragas ships DiskCacheBackend in ragas.cache, passed as cache= when constructing the evaluator LLM wrapper. Identical judge requests then hit a local cache directory instead of the provider, so re-runs over unchanged data are fast and free.

solid answer

~60 s

Ragas caches at the model layer, not the run layer. You construct a backend — `DiskCacheBackend()` from `ragas.cache`, backed by a local directory (default `.cache`) — and pass it as `cache=` when wrapping your evaluator LLM. Every judge request is then keyed on the call, so a repeated identical request is served from disk. That makes three real situations cheap: adding a new metric to a dataset whose other metrics already scored, resuming after a run crashed partway, and re-running a demo or notebook. The caveats matter as much as the mechanism. The key covers the request, so changing the prompt, the model, or the sampling parameters is a miss — a cache never silently hides a change you made to those. It *does* make a re-run look perfectly reproducible by replaying prior judge outputs, which is convenient but is not evidence that the judge is stable. And the cache directory stores prompts and completions verbatim, so it inherits whatever data-sensitivity your evaluation content has and must not be committed to the repository.

code

python · 9 lines
python
from ragas.cache import DiskCacheBackend
from ragas.llms import LangchainLLMWrapper
from langchain_openai import ChatOpenAI

cacher = DiskCacheBackend()
evaluator_llm = LangchainLLMWrapper(
    ChatOpenAI(model="gpt-4o-mini"),
    cache=cacher,
)

go deeper

for a junior

Know that Ragas can cache judge calls through a disk-backed cache attached to the evaluator model, so re-running the same data does not pay for the same calls twice.

for a middle

Explain that caching sits on the model wrapper rather than on the run, that the key covers the whole request so a changed prompt or model is a miss, and which iteration scenarios it actually speeds up.

for a senior

Raise the tradeoffs unprompted: a warm run replays judge output rather than demonstrating stability, the cache directory holds prompts and completions verbatim, and CI only benefits if the cache is deliberately persisted and invalidated.

for a principal

Own the policy — caching on for local iteration, cold or clearly labelled runs for published numbers, and an explicit stance on where cached prompt content may live given the sensitivity of the evaluation data.

## The problem caching solves Evaluation is iterative in a way that repeatedly re-issues identical work. You score a dataset on two metrics, look at the results, decide you want a third metric, and re-run — and the naive re-run pays the full price of all three metrics again even though two of them would produce the same calls. The same waste appears when a run dies at 80% and you restart it, when a notebook is re-executed, and when a demo is run three times in a meeting. ## Where Ragas puts the cache Not on `evaluate()`. Ragas attaches caching to the **model wrapper**, which is the right place because that is the layer where a request and a response exist. `ragas.cache` provides `DiskCacheBackend`, which persists to a local directory (`.cache` by default). You construct it once and hand it to the evaluator LLM wrapper via `cache=`. From then on, a judge request that matches a previously seen one is answered from disk instead of the network. Because the cache lives on the model object, it applies to whatever that model is used for during the run — every metric that shares that evaluator benefits, with no per-metric configuration. ## What is and is not a cache hit The key is derived from the call, which means everything that defines the request participates: - **Same sample, same metric, same judge, same prompt** → hit. This is the re-run case, and the one you want. - **New metric added to an unchanged dataset** → the new metric's calls are misses, everything else hits. This is the iteration case. - **Prompt changed, judge model changed, temperature changed** → miss. Important: the cache cannot mask a change you actually made, which is what makes it safe to leave on. - **New or edited samples** → miss for those rows only. ## The three caveats to raise unprompted **Reproducibility is borrowed, not earned.** With a warm cache, two consecutive runs return byte-identical judge outputs and therefore identical scores. That looks like determinism, but it is replay. It tells you nothing about how much a fresh judge call would vary, and a team that only ever runs warm can be blindsided when a cold run on new data shows spread. If you need to know the judge's variance, you must run cold on purpose. **The cache holds your content.** Every cached entry contains the prompt — which in RAG evaluation embeds the retrieved contexts and the user question — and the judge's completion. If the evaluation data contains customer text, personal data, or anything under a retention policy, the cache directory is now a copy of it sitting on a developer laptop or a CI worker. It belongs in `.gitignore`, it should not be baked into an image, and in a regulated setting it may need to be off entirely or scoped to an ephemeral workspace. **It grows.** Disk caches over long-context RAG prompts get large quickly, and there is no automatic notion of relevance. Treat the directory as disposable state you can delete at any time, and delete it when the judge model or prompts change wholesale so you are not carrying dead entries. ## Cache in CI A cache in a continuous-integration job is only useful if it persists between runs, which means restoring it as a build cache keyed on something meaningful — the dataset revision and the judge configuration, say. Done well it makes a re-run of an unchanged suite nearly free. Done carelessly it either never hits (a fresh container every time) or persists across a change that should have invalidated it. And a cached suite that never issues a real judge call also stops exercising the provider integration, so a scheduled cold run remains worthwhile as a canary. ## Deciding whether to use it at all Use it while iterating locally, where the same calls repeat constantly and the savings are immediate. Be deliberate about it for the runs whose numbers you publish: those should either be cold, or clearly labelled as replayed, so nobody mistakes cache-driven stability for a property of the system being measured.

  • You changed the judge model but kept the cache. Are your next results stale?
    No — the model identity is part of the request, so calls to a different judge are cache misses and get made for real. The cache cannot silently serve you the old model's verdicts. What you do carry is a directory full of now-useless entries for the previous model, so it is worth clearing after a wholesale judge or prompt change.
  • Why is a warm-cache run a poor way to argue that your evaluation is stable?
    Because identical scores across two warm runs come from replaying stored judge completions, not from the judge behaving consistently. It measures the cache, not the model. To say anything about judge variance you have to run cold — ideally repeatedly on the same inputs — and observe the spread that fresh calls produce.
  • What data-handling concern does the cache directory introduce?
    It stores prompts and completions verbatim, and RAG evaluation prompts embed retrieved contexts and user questions. That makes the directory a copy of potentially sensitive content on a laptop or CI worker, outside whatever retention and access controls the source system has. It must be gitignored, kept out of images, and in regulated settings scoped to ephemeral storage or disabled.

saying these in an interview costs you the question

  • Expecting a cache to make an evaluation cheap after you changed the prompt
  • Presenting warm-cache stability as evidence the judge is deterministic
  • Committing or shipping the cache directory with the project
  • Assuming a fresh CI container benefits from caching without restoring it
  • Forgetting the cache stores customer content verbatim

context