skip to content

A nightly promptfoo red-team job finishes in seconds and reports the same failure count as last month, even though the team shipped a new model behind the same endpoint. How does promptfoo's response cache produce that result, and when is leaving the cache on still the right call?

level: middleimportance: must knowfreq 52%

answer

  1. cache key = request, not deployment
  2. seconds instead of minutes = replay
  3. target cache vs grader cache
  4. cold run for scheduled checks
  5. warm run for tuning assertions

basics

~20 s

promptfoo caches responses keyed by the request it sent. If the test cases and the provider settings did not change, a rerun replays stored answers instead of calling the new model, so the numbers still describe the old deployment. Keep the cache while iterating on graders; clear or disable it whenever the system under test changed.

solid answer

~50 s

The cache key is derived from what promptfoo sends: the provider configuration plus the rendered prompt and generation parameters. A model swap that happens **behind** an unchanged endpoint name moves nothing in that key, so every test case is a cache hit and no request reaches the new deployment. The suspicious signal is wall-clock time: a run that normally costs minutes of metered calls finishing in seconds means you measured your cache, not your target. Caching is genuinely useful in two places. First, while you iterate on the object that decides whether a response counts as a failure — you want the same model outputs re-graded, not re-sampled. Second, when reproducing an old run for a write-up. It is wrong for any scheduled run whose purpose is "is production still safe today", and for anything you will quote as a trend.

go deeper

for a junior

Knows promptfoo caches provider responses and that a rerun can be served from that cache instead of the live model.

for a middle

Explains that the key is built from the request, so a server-side model change is invisible to it, and names when warm caching is appropriate versus wrong.

for a senior

Separates target, grader and generation calls, sets cache mode per job type rather than globally, and uses duration plus provider usage to catch a replayed run before it becomes a trend point.

for a principal

Sets the org rule that any number quoted outside the team comes from a cold run with recorded provider usage, and budgets for that cost instead of letting cost quietly decide validity.

Treat promptfoo's response cache as part of the experiment design, not as a performance setting. It decides whether a run is a measurement at all. ### What the cache is and what it keys on promptfoo stores provider responses on disk between runs. The key is derived entirely from the **outgoing request**: the provider identifier and its options (model name, temperature, max tokens, any headers or body fields you set in the provider block), plus the fully rendered prompt after variable substitution. That hash maps to the stored completion. Caching is on by default for `promptfoo eval`; you switch it off for one run with `promptfoo eval --no-cache`, process-wide with the `PROMPTFOO_CACHE_ENABLED` environment variable, or for a single provider with a `cache: false` in that provider's config, and you empty the store with `promptfoo cache clear`. Every input to that key is client-side. Nothing in it observes what actually answered. ### Why that is exactly the wrong blind spot The changes a nightly red-team job exists to catch are almost all server-side: a vendor rolls a new model version behind a stable alias, your platform team swaps the base model behind an internal gateway, someone edits the system prompt in a hosted assistant configuration, a moderation filter is enabled in front of the endpoint. None of those touch your YAML, none change the rendered prompt, so every test case is a cache hit and not one request reaches the new deployment. The failure count is identical because it is literally the same numbers, recomputed from stored text. ### Three caches, not one A red-team run makes three kinds of model call, and all three can be served from cache: | call | what a hit means | |---|---| | target response | you graded an old deployment's answer | | grader / judge call (model-graded assertions such as `llm-rubric`) | you replayed the old verdict, so an assertion edit appears to do nothing | | generation call (`promptfoo redteam generate`) | you re-used yesterday's test cases | Saying "the cache" as if it were one object hides the useful configuration: when you are tuning an assertion you *want* target hits (identical outputs, so the grader is the only variable) and grader misses. ### What it costs, warm and cold A modest red-team configuration — a dozen plugins at ten tests each, with two or three strategies expanding them — lands in the range of several hundred test cases. Each case is one target call, and each model-graded assertion adds a judge call, so a cold run is commonly a four-figure number of metered calls, tens of minutes of wall clock against a rate-limited endpoint, and single-digit to low-tens-of-dollars per night depending on the models. Warm, the same run is seconds and effectively free. That gap is why caching gets switched on for the nightly job and then forgotten; cost quietly decides validity. ### Where the number misleads A fully warm run is the easy case: identical output, obviously suspicious once you look at the clock. The dangerous case is a **partial** hit. Add one plugin, or edit two test cases, and ninety-odd percent of the suite replays the old deployment while the rest hits the new one. The number moves a little — which reads exactly like genuine small drift — and no single system ever produced it. A second trap: caching freezes stochastic behaviour. A case that a target refuses seventy percent of the time gets one sampled answer stored forever, so an intermittent failure is reported as always-failing or never-failing, and its apparent stability is an artefact of the cache, not of the model. ### What you check before believing the result 1. Compare run duration and recorded provider usage against the last known-cold run; promptfoo reports token counts per result, and a replayed run shows near zero. 2. Confirm on the provider's own side that traffic arrived in that window. 3. Re-run a sample — twenty cases, not the whole suite — with `--no-cache` and see whether the failure count moves. Cheap and decisive. If it moves, the trend point was never a measurement of the current system: withdraw it rather than explain it. The standing rule that avoids all of this is per-job cache mode — scheduled and reported runs cold, developer and grader-tuning loops warm.

  • You want to change how a promptfoo grader judges a response and re-score last night's run. Should the cache be warm or cold?
    Warm for the target calls: you want the identical model outputs re-graded so the only variable is the grader. Grader calls themselves should be cold, or you will just replay the old verdicts.
  • What cheap signal in the run output tells you a promptfoo run was served from cache?
    Duration and provider usage. A cached run finishes far faster than the metered call volume would allow and shows near-zero tokens or requests against the provider.
  • Your provider endpoint name is stable but the vendor rolls model versions behind it. What do you change so cached results are not silently wrong?
    Make the deployment identity part of what you record and, where the provider exposes it, part of the request configuration; and run scheduled jobs cold so the question never depends on the key.

A warm cache is a photocopier where you expected a camera: it hands back a picture of whoever stood there last time, however different the person standing there now. A partial cache is worse, because it staples half of last month's photo onto half of this month's.

saying these in an interview costs you the question

  • Treating the cache purely as a speed setting with no effect on validity
  • Assuming a new model behind the same endpoint invalidates cached entries
  • Quoting a trend point from a run whose duration and provider usage were never checked
  • Turning caching off everywhere, then abandoning the nightly job when the bill arrives

context