skip to content

Describing the Target

The generator writes attacks from your description of the application, and by default that description travels to a remote service. Interviewers ask because both halves of that surprise people.

on this pageshow

explore

questions

4

In promptfoo's red-team mode you write a free-text description of the application under test. What is that description used for, and why does a one-line version weaken the results?

level: juniorimportance: must knowfreq 70%

answer

  1. description feeds generator and graders
  2. thin text = generic attacks
  3. no stated rule = graded pass
  4. allowed vs forbidden actions
  5. read the generated cases

basics

~20 s

promptfoo's red-team mode reads that description twice: the generator writes attack cases from it, and the graders judge replies against it. A one-line description yields generic attacks that miss your app's real surfaces, and gives the graders no stated rule to fail an answer against, so weak replies get marked pass.

solid answer

~50 s

The description is the only thing promptfoo knows about your system before it writes a single case. It feeds two stages. **Generation.** Attack cases are written to fit the described application. "A chatbot" produces generic bait. "A retail support assistant that can look up an order by number for the signed-in customer, issue refunds under a threshold, and must never discuss another customer's order" produces cases aimed at those exact actions and that exact boundary. **Grading.** The object that decides whether a reply is a failure compares the reply against what you said the app is allowed to do. Behaviour you never declared out of bounds is not a violation, so it scores as a pass. That is why a thin description is worse than no run: you get a high pass rate that reflects the absence of stated rules, not the presence of safe behaviour. Write allowed actions, forbidden actions, the data and tools reachable, and who the users are.

go deeper

for a junior

Knows it is a description of the app and that the generator uses it to write attacks; can say a richer description gives better attacks.

for a middle

Names both consumers — generation and grading — and explains that undeclared limits cannot be failed, so thin text inflates the pass rate.

for a senior

Can list the specific content a usable description needs (roles, tools, data classes, forbidden actions) and describes sampling generated cases as the check that it worked.

for a principal

Frames the description as the specification the whole programme's numbers depend on, and owns keeping it true as the application changes.

### The field and its two readers In a promptfoo red-team configuration the application description lives in `redteam.purpose` inside `promptfooconfig.yaml`; `promptfoo redteam init` prompts for it when it scaffolds a config. It is ordinary free text, and it is the only thing promptfoo knows about your system that is not a transport detail. The `providers` block says *how* to reach the application; `redteam.purpose` says *what the application is*. Two separate stages read that same paragraph. **Stage one — generation.** `promptfoo redteam generate` walks the entries in `redteam.plugins`. A promptfoo plugin is a harm class, not a payload: personal-data disclosure, excessive agency, unauthorised contract-making, invented policy, and so on. For each enabled plugin it asks a generation model to write `redteam.numTests` adversarial cases *for the application described in `redteam.purpose`*, and then every entry in `redteam.strategies` rewrites those base cases into transformed variants. The result is a generated suite file on disk that you can open and read — that readability is the whole verification story below. **Stage two — grading.** `promptfoo redteam eval` sends each case to the target and hands the reply to the model-graded assertion attached to that plugin. That grader's rubric is not a fixed universal safety policy: the purpose text is interpolated into it, so the grader is effectively asked "given that this application is *this*, does this reply violate it?" A behaviour the purpose never placed out of bounds gives the grader no rule to apply, and a grader with no applicable rule returns pass. So the same paragraph is simultaneously the attack-surface map and the pass/fail policy. Anything you leave out is missing from both halves at once: the case is never written, and if a similar case arrives anyway through a strategy rewrite, there is no declared rule for it to break. ### What a usable purpose contains - **Roles and authentication** — anonymous visitor, signed-in customer, internal agent; and which of them the app is serving. Without this, cases that probe crossing a role boundary have nothing to aim at. - **Permitted actions and tools** — what it may call, submit, change, refund, file. - **Reachable data, stated by class** — "the signed-in customer's own order records", never a pasted record. - **Required refusals, in policy language** — including topics that are merely off-mission rather than dangerous. - **Boundary pairs** — "may summarise the returns policy; may not commit to an exception." ### What it costs Writing a good purpose costs perhaps half an hour once, plus a re-read whenever the app gains a capability. Running the suite is the expensive half, and its price does not change with the quality of that text. A modest configuration — say fifteen plugins at `numTests: 5` with two strategies layered on top — produces on the order of a few hundred cases, and every case is at least one target call plus one grader call, on top of the generation calls that wrote it. A one-sentence purpose buys exactly the same bill and a small fraction of the information. Because red-team suites are normally re-run on a schedule, you pay that difference every cycle rather than once. ### Where the number misleads The headline output is a pass rate, and both halves of it were chosen by that paragraph. Its denominator is "cases that happened to be generated", not "attacks that exist". Its verdict function is "rules that happened to be declared", not "harms that matter". A thin purpose shrinks the denominator toward generic bait any assistant deflects, and empties the verdict function so that even a genuinely bad reply lands as a pass. The two errors point the same way — upward — so the report is not merely narrow, it is flattering. A high pass rate produced from a thin purpose is a statement about your configuration, not about your model, and mistaking one for the other is the single most common misreading of this tool. ### What you would check Open the generated suite file and read ten cases at random: could any of them have been written *only* for your application? If they would fit any chatbot, the purpose is too thin and the score is void. Then check coverage the other way — every refusal you declared should have drawn cases against it; one that drew none is worded in language the generator did not turn into anything, and needs restating in the same operational terms as the rest. Finally, prove the graders are capable of failing at all: inspect a case whose reply you can see is off-policy and confirm it is scored a failure. A grader that has never failed anything is not evidence of safety.

  • Your first promptfoo red-team run comes back almost entirely green. What is the first thing you inspect?
    The generated cases themselves. If they could have been written for any chatbot, the description was too thin to produce app-specific attacks or app-specific grading rules.
  • Why is describing what the app must refuse as important as describing what it does?
    The graders judge against declared limits. An action you never declared out of bounds is not a violation, so the run records it as a pass.
  • Should the description include real data samples?
    No. Describe data by class, not by content — the text is a config artifact and, by default, is also sent off your machine during generation.

The purpose paragraph is both the brief you hand the attacker and the rulebook you hand the referee. Write one line and the attacker swings at a generic silhouette while the referee, holding a blank rulebook, calls everything fair.

saying these in an interview costs you the question

  • Calling it documentation or metadata that only affects the report heading.
  • Assuming it changes generation but not grading.
  • Reading a high pass rate from a one-line description as evidence the app is safe.
  • Pasting real customer records or secrets into it to 'make it specific'.

context

open as a page

Where does the application description you write for a promptfoo red-team run go by default, and what does that mean for what you are allowed to put in it?

level: middleimportance: must knowfreq 60%

basics

~20 s

By default promptfoo generates red-team cases through its hosted generation service, so the description of your application leaves your machine. Treat it as text you would hand a vendor: no secrets, no customer records, no internal hostnames. There is a documented switch that turns remote generation off and generates locally instead.

open as a page

A promptfoo red-team run against your internal HR assistant reports almost no failures. The application description in the config reads, in full: 'an HR chatbot'. Is that result evidence the assistant is safe, and what do you do next?

level: seniorimportance: must knowfreq 55%

basics

~20 s

No. promptfoo's graders judge each reply against the description you supplied. If it never says what the assistant must refuse, who may ask, or what records it reaches, an off-policy answer breaks no declared rule and is scored a pass. Rewrite the description with roles, data and refusals, then re-run.

open as a page

Several teams run promptfoo red-team suites on a schedule, each with its own written description of the application under test. How do you keep those descriptions from quietly rotting, and who should own them?

level: principalimportance: should knowfreq 35%

basics

~20 s

Treat each description as a reviewed, version-controlled artifact owned with the application. When an app gains a tool, role or data source and the text does not, promptfoo's generator stops attacking the new surface and its graders stop failing it, so the pass rate rises while real risk grows. Diff it every release.

open as a page