A promptfoo red-team run against a live app wired through the HTTP provider finishes with nearly every case passing. Before you report the app as clean, how do you rule out a mis-wired target?
answer
- everything fails open
- read the graded string
- force 401 and 500
- known-red control case
- identical latencies = constant response
basics
~20 sTreat a spotless report as a wiring hypothesis first. Read the graded output of a few cases as text, confirm it is the reply and not the envelope or an error body, and run two controls: a case that must produce a substantive answer and one you know the app refuses. If neither behaves, the target is mis-wired.
solid answer
~60 sA near-perfect pass rate has two explanations and only one of them is good news, so I try to falsify the wiring before believing the app. The checks, cheapest first: 1. **Read raw output.** Pull the stored response for several cases. If the graded string still has braces or is empty, the extraction is wrong and every number is void. 2. **Check the error path.** Force a 401 and a 500 against the target. If those cases score as passes, the run has been grading failed calls as safe behaviour — the single most common way a broken target produces a clean report. 3. **Two controls.** A benign case that must return substantive text proves replies reach the grader; a case the app is known to refuse proves the grader can produce a failure at all. 4. **Check session continuity** if the cases were multi-turn — a nonce in turn 1 recalled in turn 2. 5. **Look at the shape of the run:** near-identical latencies or response lengths across every case mean you are seeing one constant response, not an app answering. Only when a failure can be produced on demand does a pass mean anything.
go deeper
Knows to look at a couple of actual responses instead of trusting the summary table.
Checks extraction, auth and error handling, and understands that error bodies can score as passes.
Runs a falsification protocol with known-red controls, verifies session continuity, reads run-shape signals like uniform latency, and refuses to report a pass rate until a failure can be produced on demand.
Makes the controls a gate on the harness itself so no run is publishable without them, and separates 'the app is clean' from 'this suite exercised these behaviours' in how results are communicated.
## Everything in this chain fails open - A wrong extraction returns a string. - An expired credential returns a body. - A dead session returns a fresh first turn. - A throttled request returns a polite 429 page. - A WAF challenge returns HTML. None of those raise; every one of them produces something the grader can score, and what most assertions reward is the *absence* of disallowed content — which broken responses have in abundance. A near-perfect pass rate is therefore weak evidence about the application and strong evidence that you have not yet tested the harness. The correct first hypothesis for a spotless live-target report is not "strong guardrails", it is **"mis-wired target"**. ## Triage, cheapest first 1. *Is the graded string the reply?* Open several stored results — different cases, not just the first — and read them as text. Braces, an empty string, or the identical string on every row each point at the response transform rather than at safety. 2. *Is the target actually reached and authenticated?* Point the same config at a URL that 404s, then corrupt or revoke the credential on purpose. Both runs must go red. If a 401 body scores as a pass, you have learned that your assertions reward the absence of bad content and an error envelope contains none, so failure and safety are the same event to the grader. 3. *Can a failure be produced at all?* This is the **highest-value control**. Include a case whose outcome you already know from manual use — something the app is known to refuse, or a benign question it must answer substantively. A suite that has never once gone red against this target has not demonstrated that it can, and until it has, "no findings" is untested plumbing rather than a result. 4. *Did the conversation happen?* For multi-turn cases, verify recall across turns and confirm the target's own logs show one conversation per case rather than one for the run. 5. *Does the data look like an app?* Compare response lengths, latencies and token counts across cases. A live model varies per prompt; a stuck endpoint, a cached answer, a challenge page or an empty extraction is suspiciously uniform. ## What the falsification costs, against what it saves Three or four extra cases, a couple of minutes, a few cents of API spend — and once they live in the suite, that is the recurring cost forever. The alternative is a full rerun of a suite that may be hundreds of target calls plus a judge call per graded case, the engineer-days spent triaging a report that described an envelope, and a retracted claim to whoever consumed it. This is the **cheapest insurance** in the whole pipeline, and it is the reason the controls belong in the suite rather than in someone's memory. ## Where the number misleads The pass rate's real **denominator** is "cases that produced a gradeable string", not "cases the application actually answered". Errors, empty extractions, throttles and challenge pages all land in the numerator as passes, so the metric improves monotonically as the harness degrades — hold that inversion in mind, because it is the opposite of how every other test suite you have used behaves. Second, even with sound wiring, the pass rate is about the cases you sent. A clean result across the plugins and strategies you ran is not coverage of behaviours you never generated, and reporting it as "the app is safe" quietly swaps the denominator from "cases run" to "things that could go wrong". ## How I report it - If the wiring **cannot be shown sound**, the output is not "the app passed" but "this run is uninterpretable — here is which control failed and what a redo costs". - If it **can**, the sentence has two clauses: the harness was demonstrated capable of producing a failure, and on the N cases exercised it produced M. The durable fix is to **promote the controls into the suite** so that no future run can come back clean without simultaneously proving it could have come back dirty. A harness that gates itself is worth more than a human remembering to be suspicious.
- Why is 'an expired credential' one of the first things you check?Because the error body contains no disallowed content, so checks that reward its absence pass. The whole run then grades an auth failure as safe behaviour and looks better the more broken it is.
- How do you keep this from recurring on the next run?Bake the controls into the suite: a benign case that must answer, a case known to be refused, and a session-recall case. A run where those three do not behave is failed by the harness, not judged by a human.
- The controls pass and the report is still nearly clean. What do you say?That the wiring is sound and this suite found little — bounded by what it exercised, since a clean result on the cases run is not coverage of behaviours you never sent.
A clean report from an untested harness is a smoke detector that has never been tested: its silence only becomes good news after someone has held a lit match under it and heard it scream.
saying these in an interview costs you the question
- Reporting 'no findings' from a live-target run with no control case that ever fails.
- Treating HTTP 200 counts as proof the eval was valid.
- Never checking how error responses are scored.
- Explaining a spotless report as strong guardrails without inspecting a single raw output.