skip to content

promptfoo

You will learn how promptfoo blends prompt evaluation with a red-team plugin catalog to generate adversarial tests and gate them in CI. Interviewers like it because it bridges dev-time eval and security testing, showing you can make red-teaming continuous.

on this pageshow

explore

questions

page 2 of 2

In a promptfoo red-team run, one enabled harm-class plugin produces a large batch of failures that triage decides are not real problems for this application. Do you disable that plugin, and what does disabling it cost you?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Not before separating two causes: the application genuinely cannot commit that harm, or it can and the failures are being judged against the wrong expectation. Only the first justifies disabling. Disabling deletes the class from every future run too, so the next report is silently narrower while the headline number improves.

open as a page

You enable a multi-turn attack strategy in a promptfoo red team against an HTTP chat endpoint you wired up yourself. It reports almost no failures, while a single-shot strategy against the same endpoint finds several. What target-side configuration do you check first, and why?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Check that the target definition carries conversation state between turns. If every request hits the endpoint as a fresh conversation, a multi-turn strategy's gradual escalation is thrown away and each turn lands as an isolated plain ask, so it under-reports by construction rather than because the target is strong.

open as a page

A promptfoo red-team suite regenerates and re-runs nightly against a metered chat endpoint, and its bill has become the argument for deleting it. Someone proposes cutting the per-plugin case count to two across the board. What do you do instead?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Do not thin everything evenly - that leaves every harm class too small to conclude anything while still costing money. Split the schedule: a small nightly tier on the surfaces that change, and a deep sweep with high case counts before releases. Stop regenerating cases you could reuse, and label the thin tier as a smoke test.

open as a page

Your product's promptfoo red-team results are unacceptable and you can move three things: the backing model, the application's system prompt, or add a guardrail in front. How do you sequence the experiments over one frozen generated suite, and what do you refuse to change while the campaign runs?

level: principalimportance: should knowfreq 28%

basics

~20 s

Sequence by cost to test and cost to reverse: prompt first, guardrail second, model last, one variable per comparison over the same saved suite. Freeze the suite, the grader and the counting convention for the whole campaign, and hold back cases the prompt is never tuned against so improvement is not just fitting the suite.

open as a page

You own a pinned, committed promptfoo red-team suite that several teams run against different applications. How do you decide how often it is regenerated, and what do you do so numbers reported before and after a regeneration remain meaningful?

level: principalimportance: should knowfreq 24%

basics

~20 s

Refresh on events, not on a calendar alone: a new application capability, a catalogue update, or evidence that scores are being tuned against the frozen cases. Treat each refresh as a reviewed commit, run old and new suites together once to bridge the series, and break the trend line at the boundary rather than interpolating across it.

open as a page

How do you decide which safety properties in a promptfoo config may be model-graded at all, and which must be checked deterministically before anything is allowed to gate a release?

level: principalimportance: should knowfreq 34%

basics

~20 s

Anything with a machine-checkable signature is deterministic and may gate: planted canaries, forbidden artefacts, schema violations, unexpected tool calls. Model grading is for open-ended semantic harm, and it is advisory or trend evidence, not a hard gate, unless it comes with fixtures that prove it still discriminates.

open as a page

Your team wants to point a promptfoo HTTP provider at the running production deployment instead of a staging copy. As the lead, what do you require before agreeing, and what do you accept losing if you insist on staging?

level: principalimportance: should knowfreq 34%

basics

~20 s

Before production I require side-effect containment, a scoped credential, a rate and spend ceiling, a stop switch, and agreement with the app's own responders so the traffic is expected. Choosing staging instead costs fidelity: a different system prompt, guard config, model build or retrieval corpus means the number describes a system users never touch.

open as a page

A long-running promptfoo red-team matrix has one provider replaced by a different one. What happens to the pass-rate history, and how do you stop a team that watches that trend from drawing the wrong conclusion?

level: principalimportance: should knowfreq 22%

basics

~20 s

Treat it as a new series, not a continuation. The old points describe a provider you no longer call, so mark the break, keep the retired provider's last full run as the baseline, and re-run a frozen case set against the new provider before anyone reads a trend across the swap.

open as a page

Your team runs a promptfoo red-team suite on every release, and every attack strategy you enable is re-paid in full on each run. How would you decide which strategies stay in the per-release run and which move to a periodic deeper run?

level: principalimportance: should knowfreq 34%

basics

~20 s

Split by what each strategy is for. The per-release run should be a cheap, stable regression arm: deterministic framings, fixed cases, fast enough to block a release. Expensive iterative and conversational framings go to a scheduled deeper run, where a slow, stochastic result is investigated by a person rather than gating a deploy.

open as a page

Several teams run promptfoo red-team suites on a schedule, each with its own written description of the application under test. How do you keep those descriptions from quietly rotting, and who should own them?

level: principalimportance: should knowfreq 35%

basics

~20 s

Treat each description as a reviewed, version-controlled artifact owned with the application. When an app gains a tool, role or data source and the text does not, promptfoo's generator stops attacking the new surface and its graders stop failing it, so the pass rate rises while real risk grows. Diff it every release.

open as a page

You own promptfoo red-team suites for a dozen applications under one fixed monthly spend. How do you decide the per-plugin case counts across that portfolio, and why is a single uniform number the wrong answer?

level: principalimportance: should knowfreq 30%

basics

~20 s

Treat it as allocating a fixed recurring budget, not picking a setting. Give high counts to the harm classes where a miss is expensive and to apps with real exposure, and thin counts elsewhere. A uniform number overspends on harms an app cannot commit while under-sampling the ones it can, and hides both behind one green score.

open as a page

You own adversarial-testing coverage for a dozen LLM applications, all scanned with promptfoo's red-team mode. How do you govern which harm-class plugins each application runs, and what breaks when the available catalogue changes between quarters?

level: principalimportance: nice to knowfreq 29%

basics

~20 s

Set a baseline of harm classes every application must run, plus per-application additions derived from its capabilities, with every exclusion carrying a written reason and an owner. Record the enabled set with each report. When the catalogue grows, new classes change the denominator, so cross-quarter comparisons are invalid unless you re-baseline deliberately.

open as a page

showing 31–42 of 42