skip to content

Your team wants to point a promptfoo HTTP provider at the running production deployment instead of a staging copy. As the lead, what do you require before agreeing, and what do you accept losing if you insist on staging?

level: principalimportance: should knowfreq 34%

answer

  1. blast radius vs fidelity
  2. scoped, revocable test principal
  3. downstream actions and tool calls
  4. documented environment diff
  5. label the environment on every result

basics

~20 s

Before production I require side-effect containment, a scoped credential, a rate and spend ceiling, a stop switch, and agreement with the app's own responders so the traffic is expected. Choosing staging instead costs fidelity: a different system prompt, guard config, model build or retrieval corpus means the number describes a system users never touch.

solid answer

~60 s

Pointing an eval at a running app is not just a URL change; the tool sends adversarial input to something that acts on it. **Preconditions for production.** Know what the app does downstream — if it files tickets, emails, or drives tools, the run writes real artefacts. Require a scoped, revocable credential and a dedicated test principal so traffic is attributable and killable. Set a request-rate and spend ceiling, since suites multiply cases by turns by retries. Tell the people who watch abuse signals, or you will consume their incident budget testing your own app. Agree where transcripts land: adversarial prompts stored in production logs are now content someone must handle. **What staging costs.** Environments drift in exactly the places that decide outcomes — the system prompt, the guard configuration and thresholds, the model build behind the endpoint, the retrieval corpus, rate limits. A pass in staging is evidence about staging. **The judgment I actually make:** run in staging by default, require a documented, reviewed diff between the two environments, and reserve production for a narrow, low-side-effect subset where that diff says the answer could differ.

go deeper

for a junior

Knows staging is the safer default and that a live app can be affected by the traffic you send it.

for a middle

Lists concrete preconditions — scoped credentials, rate and spend limits, side effects — and can say why staging results may not transfer.

for a senior

Weighs blast radius against fidelity per suite, coordinates with the app's own responders, and handles transcript retention and a stop switch.

for a principal

Turns it into a decidable policy: a reviewed environment diff, staging by default, a narrow scoped production subset with a named owner, and an environment label attached to every reported number.

## Frame it as blast radius versus fidelity, not as a preference **Blast radius.** A live target is not a text generator, it is an application that may act. Before I agree to production I want written answers to five questions: 1. what the app writes or triggers downstream (tickets, emails, webhooks, tool calls, orders); 2. whether tool execution can be disabled or sandboxed for the test principal; 3. what the worst single request can do; 4. how we stop the run mid-flight; 5. and who is paged if it goes wrong. Alongside those, a **dedicated principal** holding a scoped, revocable credential — so the traffic is attributable in the app's own logs and killable in one action, rather than by asking somebody to go and find it. **The ceiling, with the arithmetic.** Suites are multipliers, and an estimate that drops a factor is wrong by an order of magnitude. The real cost is cases x turns per case x (one target call plus any judge calls per graded turn) x retries, and an adversarial suite generated from a plugin catalogue routinely produces far more cases than anyone eyeballed. Two hundred cases at five turns with a graded judge per turn is on the order of two thousand calls, arriving at a request rate the production rate limiter has an opinion about. A **hard ceiling** on requests per minute and on total spend, enforced by the harness rather than by intention, is the difference between an eval and an unintentional load test on a system with real users on it. **Contamination of the app's own data.** Adversarial transcripts land wherever conversations land: analytics, quality-review queues, human labelling pipelines, anything feeding future tuning. Decide **retention and deletion** before the run. It is far cheaper than afterwards, and skipping it is how a red-team exercise turns into a data-handling incident. **The defenders.** If the app has abuse detection, an adversarial suite trips it. - **Coordinate**, and the alerts are expected and attributable — at the price of not learning whether detection would have worked unprompted. - **Do not coordinate**, and you have run a detection exercise, which is a legitimate thing to run but needs its own scope and approval rather than happening as a side effect of an eval. **Fidelity, the real argument for production.** Staging drifts from production in exactly the variables that decide outcomes: - the system prompt; - the guardrail configuration and its thresholds; - the model build behind the endpoint; - the retrieval corpus; - rate limiting; - and often whether a hosted moderation service is enabled at all. A pass in staging is evidence about staging. ## Where the number misleads, in both directions A **staging pass rate** presented without its environment label reads to every recipient as a statement about the deployed system; it is a statement about a system no user touches, and nothing in the report shows the gap. A **production pass rate** can overstate too: obtained with tools sandboxed and detection suppressed, it describes a path a real request never takes, and a run whose rate ceiling silently throttled a third of its cases has scored those errors somewhere — usually as passes. Any pass rate that does not name the environment, the privilege level of the credential used, and whether tool execution was live is a number nobody can act on. ## What I require, and what I accept losing **Staging by default**, plus a reviewed **environment diff** produced by the owning team, listing every variable above. If that diff is empty on the axes this suite exercises, staging is sufficient and production adds risk for nothing. If it is not empty, the diff names precisely which findings need production confirmation, and I reserve production for that narrow, low-side-effect, preferably read-only subset: low concurrency, a stop switch, a named owner, a scoped principal, an agreed window. What I accept losing by insisting on staging is fidelity on exactly the axes the diff shows as open — and the discipline is to write that sentence into the report rather than let the reader assume otherwise.

  • What single artefact makes the staging-versus-production argument decidable?
    A reviewed environment diff listing system prompt, guard configuration and thresholds, model build, retrieval corpus and rate limits. If it is empty on the axes the suite exercises, staging is enough; if not, it names exactly which findings need production confirmation.
  • The app takes real actions on tool calls. Does that rule production out?
    Not automatically, but it moves the requirement: the test principal must have tools sandboxed or disabled, or the subset must be limited to read-only paths. Otherwise the eval is not a test, it is the tool doing the thing.
  • Should you tell the app's abuse-detection owners before the run?
    Yes by default, so alert noise is expected and attributable. If you deliberately do not, that is a detection exercise with its own scope and approval, not a side effect of an eval.

saying these in an interview costs you the question

  • Treating production versus staging as only a URL change.
  • Running adversarial suites against production with no rate ceiling, stop switch or named owner.
  • Reporting a pass rate with no environment label and no diff between environments.
  • Ignoring where adversarial transcripts are stored and who has to handle them afterwards.

context