skip to content

promptfoo

You will learn how promptfoo blends prompt evaluation with a red-team plugin catalog to generate adversarial tests and gate them in CI. Interviewers like it because it bridges dev-time eval and security testing, showing you can make red-teaming continuous.

on this pageshow

explore

questions

page 1 of 2

You ran promptfoo's red-team mode before and after editing your application's system prompt, and the failure rate moved. What has to have stayed identical between the two runs for that difference to be evidence about the prompt?

level: juniorimportance: must knowfreq 62%

answer

  1. one variable, frozen suite
  2. cases are generated, not fixed
  3. grader is a model too
  4. same target decoding settings
  5. regenerated = new denominator

basics

~20 s

The same generated adversarial test cases, the same promptfoo plugins and strategies that produced them, the same grader deciding pass or fail, and the same target model and decoding settings. Only the system prompt may differ. If the cases were generated again, the two runs attacked different inputs and the numbers are not comparable.

solid answer

~50 s

promptfoo's red-team mode does not carry a fixed list of attack cases: the plugins (what harm or failure is probed) and the strategies (how the probe is dressed up) **generate** the prompts. So two runs of the same configuration are not automatically the same suite. For a prompt A/B you need: the generated cases saved and replayed rather than produced fresh; the same grader, which is itself model-driven and is the thing that decides a response counted as a failure; and the same target settings on both sides — model, temperature, tools, context, output limits. Then the delta has exactly one candidate cause. If anything else moved, say so instead of reporting the delta. A number produced from two different populations of generated attacks mixes "the prompt got better" with "this batch of attacks was easier", and you cannot separate the two after the fact.

go deeper

for a junior

Should say the test cases and the grading must be the same and only the prompt should differ, and should know promptfoo's red-team cases are generated rather than shipped as a fixed list.

for a middle

Adds the target settings (temperature, output limits, tools, context) and notices that the grader is itself a model whose configuration is part of the control.

for a senior

Talks about running both configurations over one saved suite in a single run, recording the suite and grader alongside the number, and refusing to report a delta when more than one variable moved.

for a principal

Frames it as experimental hygiene for a recurring safety metric: what is versioned, who may change it, and how a comparison is invalidated and re-baselined when the suite has to change.

### What promptfoo's red-team mode actually produces promptfoo's red-team mode ships **no fixed corpus of attacks**. You declare intent in the `redteam` block of promptfoo's `promptfooconfig.yaml`: `redteam.purpose` describes what your application is for, `redteam.plugins` lists the harms and failure modes to probe (a PII-leak plugin, a harmful-content plugin, an out-of-scope/overreach plugin, and so on), `redteam.strategies` lists how each base case gets dressed up (an encoding wrapper, a multi-turn escalation, an attacker-model rewriting loop), and `redteam.numTests` sets how many base cases each plugin is given. promptfoo's `redteam generate` command then calls an attacker model and writes the concrete cases it invented into a file — `redteam.yaml` by default. promptfoo's `redteam eval` sends those saved cases to whatever is listed under the config's `providers` key and grades each response. promptfoo's `redteam run` does both in one command, and that convenience is exactly the trap in an A/B: run it twice and you have *generated* twice. ### Three of the moving parts are models, not fixtures 1. **The attack generator.** Sampling, not enumeration. Regenerate and you get different wordings, a different mix across plugin categories, and often a different total count — so the numerator *and* the denominator of your failure rate both moved. 2. **The grader.** Something must decide whether a given response counted as a failure. For open-ended harm that is a model reading the transcript against a rubric (promptfoo lets you pin it via `defaultTest.options.provider`; leave it unset and you inherit a platform default that can change under you). Change the grading model or the rubric and the *same* transcripts get relabelled. 3. **The target.** The `providers` entry is not just a model id: `temperature`, `max_tokens`, tool access, retrieved context and any system message all change how the same attack lands. Only one thing is allowed to differ between the two runs — your application's system prompt. Everything above is the control. ### What the run costs, so you know what a rerun buys Multiply it out. Twenty plugins at `numTests: 10` is 200 base cases; three strategies typically produce their own variant of each base case, so the saved suite is closer to 800. Each single-turn case is one target call plus one grader call; each multi-turn case is a target call *per turn* plus the attacker model's own call per turn, so a turn-capped conversational strategy costs five to ten times a single-turn case. A mid-sized suite is therefore a few thousand model calls, single-digit to low-double-digit dollars at commodity API prices, and ten to forty minutes of wall clock at the concurrency promptfoo defaults to. Generation itself is a separate bill on top. The reason to save `redteam.yaml` and replay it is not only rigour — it is that the generation step is the part you are paying for twice. ### Where the number misleads The specific wrong reading is: *"failure rate went 14% to 11% after the prompt edit, so the prompt helped."* If the second run regenerated, those two percentages have different denominators (say 802 cases and 764), different per-plugin composition, and disjoint case sets — nothing about them is paired. "Three points better" is then partly the prompt and partly which attacks the generator happened to invent this morning, in unknown proportion, and no post-hoc analysis can split them because the evidence was never collected. Two quieter versions of the same error. First, the aggregate is an average over a heterogeneous suite: one plugin drawing an easier batch moves the headline while nothing about your application changed. Second, promptfoo caches provider responses keyed by prompt plus provider; because you edited the system prompt, the new side is generated fresh while the unchanged side may be replaying cached responses from weeks ago — one side is a fresh sample, the other a frozen one, and a hosted model may have been silently updated in between. ### What you check before believing your own number ```text same saved cases? identical file, identical count and per-plugin mix same grader? grading model pinned by id + same rubric text same target settings? temperature, max tokens, tools, retrieved context cache? both sides fresh, or both sides replayed - not one each errors/timeouts? similar counts, and counted the same way on both sides only one thing moved? name it in the report, in one sentence ``` If more than one of those moved, the honest output is "incomparable, rerunning" — not a percentage with a caveat stapled to it. Where promptfoo can drive two providers over the same saved cases inside a single `eval`, prefer that: one run, one suite, one grading pass removes the entire class of "was it really the same suite?" doubt.

  • Why is saving the generated cases better than rerunning the same plugin and strategy configuration?
    The configuration only fixes what kind of attacks are produced, not which ones. Rerunning generation gives a fresh sample: different wordings, different mix, sometimes a different count. Replaying saved cases makes the two runs compare the same inputs.
  • You changed the prompt and the temperature at the same time. What can you report?
    Only that the pair scored differently on this suite. To attribute it you rerun from a common baseline moving one at a time; temperature especially changes variance, so the same prompt can score differently at a different setting.

Regenerating the adversarial cases between the two runs is like weighing yourself on a different scale in a different room and calling the difference weight loss. The reading is real; the comparison is not.

saying these in an interview costs you the question

  • Assuming two runs of the same promptfoo red-team configuration produce the same test cases.
  • Treating the grader as a fixed oracle rather than a model-driven component that is part of the control.
  • Reporting a percentage delta when the model, prompt and settings all moved, with the confounds only mentioned verbally.
  • Comparing failure counts across runs with different case counts without normalising.

context

open as a page

In a promptfoo eval configuration, what is the difference between a deterministic assertion (substring, regex, or a small script) and a model-graded llm-rubric assertion, and what does each one cost per test case?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A deterministic assertion compares the output against a rule you wrote, so it is free, instant, and gives the same verdict on every rerun. An llm-rubric assertion sends the output plus your written criterion to a grading model, costing one extra paid call per case and returning a verdict that can shift between runs.

open as a page

In promptfoo, you point the HTTP provider at your own chat service, which answers with a JSON envelope. What is the job of the response transform, and what do the assertions grade if you never declare one?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The transform picks the assistant's text out of the HTTP response body and hands that string to the graders. Without one, promptfoo grades whatever the raw body serialises to: the whole envelope, metadata included. Assertions then match against wrapper fields, so the report looks plausible but describes the envelope, not the reply.

open as a page

A promptfoo eval config lists 3 providers, 4 prompt variants and 25 test cases. How many target-model calls does one run make, and what should that number change about your plan?

level: juniorimportance: must knowfreq 45%

basics

~20 s

promptfoo crosses every prompt with every provider for every test case, so 3 x 4 x 25 is 300 calls per run, before any model-graded check adds its own. The matrix multiplies rather than adds, so each extra provider or prompt buys a whole column you pay for on every run.

open as a page

In promptfoo's red-team mode, what does adding a plugin to the red-team configuration actually change about the suite that gets generated?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Each promptfoo red-team plugin stands for one harm class. Adding one makes the generator write adversarial test cases aimed at eliciting that harm, and it supplies the grading criteria used to judge the replies. A plugin you leave out generates nothing, so that harm is never attempted and never shows up in the report.

open as a page

A promptfoo red-team run was configured with several attack plugins but no attack strategies enabled, and every generated case passed. What does that result support, and what should you refuse to claim from it?

level: juniorimportance: must knowfreq 66%

basics

~20 s

It supports one narrow claim: the target refused those harms when asked plainly, in a single turn, in the generator's own phrasing. It says nothing about the same asks under an adversarial framing or across several turns, which is where real failures live. Report it as passing the plain baseline, never as safe.

open as a page

In promptfoo's red-team mode you write a free-text description of the application under test. What is that description used for, and why does a one-line version weaken the results?

level: juniorimportance: must knowfreq 70%

basics

~20 s

promptfoo's red-team mode reads that description twice: the generator writes attack cases from it, and the graders judge replies against it. A one-line description yields generic attacks that miss your app's real surfaces, and gives the graders no stated rule to fail an answer against, so weak replies get marked pass.

open as a page

In a promptfoo red-team configuration, what does the per-plugin case count (numTests) actually control, and what does leaving it at a small default mean for a failure your app commits only rarely?

level: juniorimportance: must knowfreq 62%

basics

~20 s

It sets how many adversarial cases promptfoo generates for each enabled plugin, so total volume scales with plugins times that count. A small default gives each harm class only a handful of tries, so a failure the app commits rarely can easily never fire, and the run then reports a clean pass.

open as a page

A release swapped the backing model, rewrote the system prompt and added an input guardrail, and the promptfoo red-team failure rate for the app dropped afterwards. Why is "we got safer" not a supportable conclusion, and how would you rerun to attribute the change?

level: middleimportance: must knowfreq 68%

basics

~20 s

Three variables moved, so the drop has three possible causes and any of them could be masking a regression in another. Rerun the same saved adversarial suite from one baseline, adding one change at a time, with the grader and target settings fixed. Only then does a delta name a cause.

open as a page

promptfoo's red-team mode generates adversarial test cases from its plugin and strategy catalogue before running them. What changes if you write that generated suite to a file and commit it, versus letting every CI run generate a fresh one?

level: middleimportance: must knowfreq 46%

basics

~20 s

Generation is itself a model call, so a fresh run produces different test cases. Committing the generated suite makes it a fixed artifact: runs become comparable and diffable, and a failure is reproducible. Regenerating every run measures a new suite each time, so month-over-month numbers move for reasons nobody can attribute.

open as a page

A nightly promptfoo red-team job finishes in seconds and reports the same failure count as last month, even though the team shipped a new model behind the same endpoint. How does promptfoo's response cache produce that result, and when is leaving the cache on still the right call?

level: middleimportance: must knowfreq 52%

basics

~20 s

promptfoo caches responses keyed by the request it sent. If the test cases and the provider settings did not change, a rerun replays stored answers instead of calling the new model, so the numbers still describe the old deployment. Keep the cache while iterating on graders; clear or disable it whenever the system under test changed.

open as a page

In a promptfoo eval, why is asserting that the output contains a refusal phrase such as 'I cannot help with that' a weak safety check, and how does it fail in both directions?

level: middleimportance: must knowfreq 58%

basics

~20 s

It measures wording, not behaviour. A model can open with a polite refusal and then comply anyway, so the check passes an unsafe answer. It can also refuse in different words, or safely answer a benign case, and the check fails a fine output. You end up scoring phrasing.

open as a page

You are evaluating a multi-turn conversation through promptfoo's HTTP provider against a running chat app. Who can hold the conversation state — the tool or the app — and what has to be wired in each case?

level: middleimportance: must knowfreq 65%

basics

~20 s

Either side can. If the tool holds it, every request carries the full prior turns and the app must be stateless. If the app holds it, the tool has to read a session identifier out of the first response and send it back on each later request. Pick one; mixing them grades unrelated single turns.

open as a page

A promptfoo red-team run over an internal support assistant comes back with zero failures, and the configuration enabled three plugins. What can and cannot you conclude from that report?

level: middleimportance: must knowfreq 61%

basics

~20 s

You can conclude that on this run, for those three harm classes, at that suite size, nothing failed. You cannot conclude the assistant is safe. Harm classes with no plugin enabled generated no cases, so the report is silent about them. Silence is a missing denominator, not a pass.

open as a page

In a promptfoo red-team configuration, attack plugins and attack strategies are two separate lists. If you enable one more strategy, how does the generated test-case count change, and why does that matter for a suite that reruns on every release?

level: middleimportance: must knowfreq 72%

basics

~20 s

Strategies do not add cases, they re-express them. Each strategy takes the cases the plugins generated and emits a transformed variant of every one, so the suite scales with plugins times strategies rather than growing by a fixed amount. Every scheduled rerun re-pays that multiplied count in target calls, grader calls and wall-clock time.

open as a page

Where does the application description you write for a promptfoo red-team run go by default, and what does that mean for what you are allowed to put in it?

level: middleimportance: must knowfreq 60%

basics

~20 s

By default promptfoo generates red-team cases through its hosted generation service, so the description of your application leaves your machine. Treat it as text you would hand a vendor: no secrets, no customer records, no internal hostnames. There is a documented switch that turns remote generation off and generates locally instead.

open as a page

A promptfoo red-team run against a live app wired through the HTTP provider finishes with nearly every case passing. Before you report the app as clean, how do you rule out a mis-wired target?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Treat a spotless report as a wiring hypothesis first. Read the graded output of a few cases as text, confirm it is the reply and not the envelope or an error body, and run two controls: a case that must produce a substantive answer and one you know the app refuses. If neither behaves, the target is mis-wired.

open as a page

Your promptfoo red-team suite costs too much to run on every merge, so provider-by-prompt cells have to go. How do you decide which pairings survive, and what can the trimmed run no longer tell you?

level: seniorimportance: must knowfreq 40%

basics

~20 s

Keep the pairing you actually ship on every run and rotate the rest, so each provider and each prompt appears somewhere without crossing them all. Then state the loss out loud: a rotated matrix says a failure happened, not which provider-and-prompt pairing owns it. Re-cross the full matrix before a release.

open as a page

A promptfoo red-team run against your internal HR assistant reports almost no failures. The application description in the config reads, in full: 'an HR chatbot'. Is that result evidence the assistant is safe, and what do you do next?

level: seniorimportance: must knowfreq 55%

basics

~20 s

No. promptfoo's graders judge each reply against the description you supplied. If it never says what the assistant must refuse, who may ask, or what records it reaches, an off-policy answer breaks no declared rule and is scored a pass. Rewrite the description with roles, data and refusals, then re-run.

open as a page

A promptfoo eval reads its system prompt from a file the product team edits in place, and each run records only the resulting pass rate. Why does the week-over-week trend become uninterpretable, and what do you change?

level: middleimportance: should knowfreq 30%

basics

~20 s

The pass rate stops being a trend and becomes a coincidence: a drop could be the edited prompt, a change in the model behind the provider, or a different generated case set. Pin each prompt as a versioned artefact, record which version every run used, and move one axis at a time.

open as a page

You have budget to widen a promptfoo eval by exactly one axis: a second provider or a second prompt variant. What does each one buy you, and how do you choose?

level: middleimportance: should knowfreq 33%

basics

~20 s

A second provider tells you whether a weakness lives in the model or in your prompt. A second prompt variant tells you whether your wording is what is holding. If you ship on exactly one provider, add the prompt variant; if you are still choosing a model, add the provider. Either one doubles the run.

open as a page

How do you decide which promptfoo red-team plugins apply to a given application, and what goes wrong if you simply enable the whole catalogue?

level: middleimportance: should knowfreq 50%

basics

~20 s

Work from what the application can actually do: its tools, the data it can reach, who talks to it, and what it must never say. Enable the harm classes it could genuinely commit. Enabling everything costs generation and inference on irrelevant classes and buries real failures in noise nobody triages.

open as a page

In a promptfoo red-team report, the same underlying plugin case passes when it is asked plainly but fails under one wrapped attack strategy. How do you interpret that, and why is a single overall pass rate a poor way to report it?

level: middleimportance: should knowfreq 46%

basics

~20 s

It means the refusal is keyed to the surface form of the request rather than to its intent: change the framing and the same harm goes through. Pool it into one overall pass rate and that signal disappears, because the plain copies dilute the wrapped failures. Report per plugin and per strategy instead.

open as a page

A promptfoo red-team plugin generates 10 cases against your assistant, and the behaviour it probes appears in roughly 1 relevant response in 20. Roughly what chance does that run have of reporting zero hits, and how many cases would you need before a clean result means something?

level: middleimportance: should knowfreq 45%

basics

~20 s

Treating each case as an independent try, the chance of missing every time is 0.95 to the tenth, about 60 percent. So a clean run of ten mostly reflects the sample size. To be roughly 95 percent sure of seeing a one-in-twenty behaviour you need about sixty cases.

open as a page

Two promptfoo red-team runs over the same saved adversarial suite report 12% and 10% failures for your app. Why is "two points better" a weak summary of what the change did, and what view do you look at instead?

level: seniorimportance: should knowfreq 38%

basics

~20 s

A net delta hides composition. Fixing nine cases while breaking seven also reads as two points better, and the seven new failures may be worse than the nine fixed ones. Join the runs case by case, split the moves into newly-failing and newly-passing, and group them by plugin category and severity.

open as a page

You changed only the target model in a promptfoo red-team run — same saved adversarial cases, same plugins and strategies, same grader — and the failure rate moved several points. What could have changed besides the model, and how would you rule each out before reporting the model as the cause?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Multi-turn and adaptive strategies build follow-ups from the target's own replies, so the prompts actually sent differ per model even from identical seeds. The grader is a model and may score a new refusal style differently. Caching, truncation and errored cases also differ. Rule them out by diffing sent transcripts and re-running the baseline twice.

open as a page

A team quotes a monthly trend of failures from a scheduled promptfoo red-team job, and this month it dropped by half. The report does not say whether the suite was regenerated or whether responses were replayed from cache. What do you check to explain the drop, and what should each run record so the question is answerable next time?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Check whether the graded population changed: case count and case identity, then whether responses were served from cache, then whether the grader or its judge model changed. Only after those hold constant does a target change explain the drop. Each run should record suite identity, case count, target identity and cache mode alongside the number.

open as a page

Every case in your promptfoo suite gained an llm-rubric assertion, and the suite's grading spend and wall-clock time have roughly tripled. How do you bring the grading cost down without losing safety signal?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Stop grading what a rule can decide. Put deterministic checks on everything with a signature and let a case that already fails one skip the grader. Keep model grading for the semantic residue only, cache verdicts for unchanged outputs, and run the graded tier on a schedule rather than on every commit.

open as a page

A promptfoo red-team run reports 96% of cases passing, and almost every safety check in the config is a model-graded llm-rubric assertion. What has that 96% actually measured, and what do you check before quoting it to anyone?

level: seniorimportance: should knowfreq 50%

basics

~20 s

It measured how often a grading model, reading your rubric, declined to object. That is not the same as safe. Before quoting it, read a sample of passing transcripts, check what the rubric text literally asks, confirm errored grader calls are not counted as passes, and state the denominator: cases this suite generated.

open as a page

You run promptfoo test cases concurrently against an app that keeps conversation state server-side. What goes wrong if every case ends up on the same session identifier, and how would you prevent and detect it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The cases share one conversation, so each is graded on history other cases wrote. That produces both false hits, where an earlier case primed the refusal or the compliance, and false passes, and it makes the run unreproducible. Prevent it by scoping a fresh session per case; detect it by rerunning serially and diffing outcomes.

open as a page

showing 1–30 of 42