In promptfoo's red-team mode you write a free-text description of the application under test. What is that description used for, and why does a one-line version weaken the results?
answer
- description feeds generator and graders
- thin text = generic attacks
- no stated rule = graded pass
- allowed vs forbidden actions
- read the generated cases
basics
~20 spromptfoo's red-team mode reads that description twice: the generator writes attack cases from it, and the graders judge replies against it. A one-line description yields generic attacks that miss your app's real surfaces, and gives the graders no stated rule to fail an answer against, so weak replies get marked pass.
solid answer
~50 sThe description is the only thing promptfoo knows about your system before it writes a single case. It feeds two stages. **Generation.** Attack cases are written to fit the described application. "A chatbot" produces generic bait. "A retail support assistant that can look up an order by number for the signed-in customer, issue refunds under a threshold, and must never discuss another customer's order" produces cases aimed at those exact actions and that exact boundary. **Grading.** The object that decides whether a reply is a failure compares the reply against what you said the app is allowed to do. Behaviour you never declared out of bounds is not a violation, so it scores as a pass. That is why a thin description is worse than no run: you get a high pass rate that reflects the absence of stated rules, not the presence of safe behaviour. Write allowed actions, forbidden actions, the data and tools reachable, and who the users are.
go deeper
Knows it is a description of the app and that the generator uses it to write attacks; can say a richer description gives better attacks.
Names both consumers — generation and grading — and explains that undeclared limits cannot be failed, so thin text inflates the pass rate.
Can list the specific content a usable description needs (roles, tools, data classes, forbidden actions) and describes sampling generated cases as the check that it worked.
Frames the description as the specification the whole programme's numbers depend on, and owns keeping it true as the application changes.
### The field and its two readers In a promptfoo red-team configuration the application description lives in `redteam.purpose` inside `promptfooconfig.yaml`; `promptfoo redteam init` prompts for it when it scaffolds a config. It is ordinary free text, and it is the only thing promptfoo knows about your system that is not a transport detail. The `providers` block says *how* to reach the application; `redteam.purpose` says *what the application is*. Two separate stages read that same paragraph. **Stage one — generation.** `promptfoo redteam generate` walks the entries in `redteam.plugins`. A promptfoo plugin is a harm class, not a payload: personal-data disclosure, excessive agency, unauthorised contract-making, invented policy, and so on. For each enabled plugin it asks a generation model to write `redteam.numTests` adversarial cases *for the application described in `redteam.purpose`*, and then every entry in `redteam.strategies` rewrites those base cases into transformed variants. The result is a generated suite file on disk that you can open and read — that readability is the whole verification story below. **Stage two — grading.** `promptfoo redteam eval` sends each case to the target and hands the reply to the model-graded assertion attached to that plugin. That grader's rubric is not a fixed universal safety policy: the purpose text is interpolated into it, so the grader is effectively asked "given that this application is *this*, does this reply violate it?" A behaviour the purpose never placed out of bounds gives the grader no rule to apply, and a grader with no applicable rule returns pass. So the same paragraph is simultaneously the attack-surface map and the pass/fail policy. Anything you leave out is missing from both halves at once: the case is never written, and if a similar case arrives anyway through a strategy rewrite, there is no declared rule for it to break. ### What a usable purpose contains - **Roles and authentication** — anonymous visitor, signed-in customer, internal agent; and which of them the app is serving. Without this, cases that probe crossing a role boundary have nothing to aim at. - **Permitted actions and tools** — what it may call, submit, change, refund, file. - **Reachable data, stated by class** — "the signed-in customer's own order records", never a pasted record. - **Required refusals, in policy language** — including topics that are merely off-mission rather than dangerous. - **Boundary pairs** — "may summarise the returns policy; may not commit to an exception." ### What it costs Writing a good purpose costs perhaps half an hour once, plus a re-read whenever the app gains a capability. Running the suite is the expensive half, and its price does not change with the quality of that text. A modest configuration — say fifteen plugins at `numTests: 5` with two strategies layered on top — produces on the order of a few hundred cases, and every case is at least one target call plus one grader call, on top of the generation calls that wrote it. A one-sentence purpose buys exactly the same bill and a small fraction of the information. Because red-team suites are normally re-run on a schedule, you pay that difference every cycle rather than once. ### Where the number misleads The headline output is a pass rate, and both halves of it were chosen by that paragraph. Its denominator is "cases that happened to be generated", not "attacks that exist". Its verdict function is "rules that happened to be declared", not "harms that matter". A thin purpose shrinks the denominator toward generic bait any assistant deflects, and empties the verdict function so that even a genuinely bad reply lands as a pass. The two errors point the same way — upward — so the report is not merely narrow, it is flattering. A high pass rate produced from a thin purpose is a statement about your configuration, not about your model, and mistaking one for the other is the single most common misreading of this tool. ### What you would check Open the generated suite file and read ten cases at random: could any of them have been written *only* for your application? If they would fit any chatbot, the purpose is too thin and the score is void. Then check coverage the other way — every refusal you declared should have drawn cases against it; one that drew none is worded in language the generator did not turn into anything, and needs restating in the same operational terms as the rest. Finally, prove the graders are capable of failing at all: inspect a case whose reply you can see is off-policy and confirm it is scored a failure. A grader that has never failed anything is not evidence of safety.
- Your first promptfoo red-team run comes back almost entirely green. What is the first thing you inspect?The generated cases themselves. If they could have been written for any chatbot, the description was too thin to produce app-specific attacks or app-specific grading rules.
- Why is describing what the app must refuse as important as describing what it does?The graders judge against declared limits. An action you never declared out of bounds is not a violation, so the run records it as a pass.
- Should the description include real data samples?No. Describe data by class, not by content — the text is a config artifact and, by default, is also sent off your machine during generation.
The purpose paragraph is both the brief you hand the attacker and the rulebook you hand the referee. Write one line and the attacker swings at a generic silhouette while the referee, holding a blank rulebook, calls everything fair.
saying these in an interview costs you the question
- Calling it documentation or metadata that only affects the report heading.
- Assuming it changes generation but not grading.
- Reading a high pass rate from a one-line description as evidence the app is safe.
- Pasting real customer records or secrets into it to 'make it specific'.