Where does the application description you write for a promptfoo red-team run go by default, and what does that mean for what you are allowed to put in it?
answer
- hosted generation by default
- description = the egress payload
- classes and roles, never records
- switch disables remote generation
- local generation, weaker cases
basics
~20 sBy default promptfoo generates red-team cases through its hosted generation service, so the description of your application leaves your machine. Treat it as text you would hand a vendor: no secrets, no customer records, no internal hostnames. There is a documented switch that turns remote generation off and generates locally instead.
solid answer
~50 sTwo facts surprise people, and interviewers ask about both. First, red-team case generation is not purely local by default: promptfoo calls a hosted service to write the adversarial cases, and the description of your application goes with the request. That is a data-egress event, and in a regulated shop it needs the same review any other third-party call gets. Second, you can turn it off. promptfoo documents a switch that disables remote generation, after which cases are produced locally using a provider you configure. That is not free: the quality and variety of generated cases depend on whatever model you point it at, so a local-only run and a hosted run are not interchangeable results and should not be compared as if they were. The practical stance: write the description at the level of roles, capabilities and data *classes* — never sample records, credentials, internal URLs or unreleased product names — and decide deliberately, once, per environment which generation mode you run.
go deeper
Knows the description is config text and that some part of the run talks to a remote service, so secrets do not belong in it.
States that generation is hosted by default, that the description is the payload, and that a documented switch makes generation local at a cost in case quality.
Reviews the config as a disclosure surface, verifies empirically what leaves the host, and treats a change of generation mode as breaking comparability of scores.
Sets one policy per environment for generation mode and description content, and makes that policy a precondition for trusting any number the tool reports.
### The default path `promptfoo redteam generate` does not, by default, write adversarial cases using a model you supply. It calls promptfoo's hosted generation service, and the payload of that call carries `redteam.purpose` — your description of the application — along with the plugin identifiers and strategies you enabled. The generated cases come back and are written into the local suite file. Everything after that (dispatching cases to the target, grading replies) runs from your machine, but the description has already left it. Nothing about this is a defect. It is how the tool produces strong, varied attacks without you standing up and maintaining a capable generator yourself. But it makes the answer to "can I paste our architecture notes in here?" *no, by default*, and it makes "has anyone reviewed this config for what it discloses?" a legitimate review question rather than a pedantic one. **Why the purpose is the sensitive string.** Every other field in the config is mechanical — an endpoint, a header name, a plugin id. The purpose is the one field whose value is a description of your business logic, your data, your roles and your internal policy, written by someone who was trying to be as specific as possible because specificity is what makes the attacks good. The incentive of the field pulls exactly against disclosure hygiene, which is why this trips people. ### Writing it so it can safely leave Describe classes and roles, not instances: - "reads the signed-in customer's own order records" rather than a pasted order. - "calls an internal refund service" rather than that service's hostname. - "three user tiers: anonymous, customer, agent" rather than a tenant or customer list. This is not only privacy hygiene. A pasted instance makes the generator overfit to that one example and produce near-duplicates of it; a described class produces cases spread across the class. The safe version is usually the higher-yield version. ### Turning generation local promptfoo documents an environment switch that disables remote red-team generation (`PROMPTFOO_DISABLE_REDTEAM_REMOTE_GENERATION=true`); cases are then produced through the generation provider you configure. State the consequences plainly: - Case quality and variety now track your local generator, which is usually the weaker model. - Some plugins and strategies are implemented on the hosted side. With remote generation off, those can yield fewer cases or none at all — read the per-plugin counts in the generate output instead of assuming you got the same suite in a different place. - Grading is a separate stage with its own provider, and the target itself is often a hosted API. Flipping one generation switch does not make the run offline, and claiming an air gap on the strength of it is a claim you have not checked. ### What it costs The money cost of the run barely moves; the costs are elsewhere. Locally generating a few hundred cases with a self-hosted model spends its tokens and its wall-clock, and on modest hardware that stage can dominate the run. The larger standing cost is that you now maintain two incomparable score series — anything generated hosted cannot be trended against anything generated locally. And the review cost is real and one-off: getting the default egress approved for one environment, once, is usually cheaper than permanently running a degraded generator. ### Where the number misleads Two readings go wrong. First, comparing a local-generation pass rate with an earlier hosted one as though the app improved: the generator changed, so the suite changed, so the denominator changed — the delta is an artefact. Second, reading a *smaller* generated case count as "fewer attacks were needed" rather than "coverage silently dropped for the plugins that only generate remotely". A missing plugin does not show up as a failure; it shows up as a harm class with no denominator at all, which reads on the report as nothing. ### What you would check Run one full `promptfoo redteam generate` followed by `promptfoo redteam eval` on a host whose outbound traffic you are logging — an egress proxy or firewall log — and read what actually left, from which stage. Diff per-plugin generated case counts between the two generation modes so a silently empty plugin is visible. Confirm the environment you ran in (laptop, shared CI runner, isolated evaluation host) is one where that egress is permitted. And put the purpose text under code review like any other config, because it is the artefact in the repository that discloses the most.
- Your organisation forbids sending application details to third parties. Can you still run promptfoo's red-team generation?Yes — disable remote generation and generate through a provider you host. Expect weaker, less varied cases, and do not compare that run's score with earlier hosted runs.
- Why is 'describe data by class, not by example' good for attack quality as well as for privacy?A pasted example makes the generator overfit to that one instance; a described class produces cases spread across the whole class.
saying these in an interview costs you the question
- Assuming everything runs locally because the CLI runs locally.
- Pasting sample records, credentials or internal hostnames into the description to make it specific.
- Claiming an air gap after flipping one setting, without checking what each stage calls.
- Comparing a hosted-generation pass rate with a local-generation one as if they were the same suite.