skip to content

Before a garak engagement on a customer-facing assistant, you must choose where the generator points: the vendor model endpoint behind the app, a staging deployment of the app, or live production. How do you argue the choice, and what does each option's number actually mean?

level: principalimportance: should knowfreq 35%

answer

  1. three endpoints, three denominators
  2. model baseline = attribution, not risk
  3. staging validity = drift from production
  4. production = logs, alerts, bans, spend
  5. credential defines authorised scope

basics

~20 s

Name what each choice measures. The bare model measures the model; a staging deployment measures the app's prompts, retrieval and filters; production measures the live system plus its abuse controls, at the cost of polluted logs, alerts and real spend. Default to a staging build that mirrors production, and state the choice in the report.

solid answer

~60 s

Treat it as a measurement question, not a convenience question. - **Vendor model endpoint.** Removes the application. Useful as a baseline for what the model does unprotected, and for attributing a hit to the model rather than to a retrieved document. It says nothing about shipped risk. - **Staging deployment.** Measures the product as built — system prompt, retrieval corpus, tools, output filters — with no customer blast radius. This is the default, and its validity rests entirely on how faithfully staging mirrors production: same prompt version, same filter configuration, comparable retrieval content. - **Production.** The only place where the real gateway, rate limits, abuse detection and incident tooling are in the loop. It is also where thousands of attack prompts land in real logs, page a real on-call, risk an account ban mid-run, and bill real money. The defensible plan is usually a small production confirmation of findings already established in staging, agreed in writing with the operators, with time windows, a kill switch and a named contact.

go deeper

for a junior

Knows scanning production has consequences and that a non-production target is usually preferred.

for a middle

Distinguishes what a model-level scan measures from what an application-level scan measures.

for a senior

Checks staging-versus-production drift, plans a narrow authorised production confirmation, and records endpoint and version in the report.

for a principal

Owns the tradeoff end to end: authorisation tied to the credential, blast radius on logs and quota, cost, and what each resulting number is permitted to claim.

### One config key, three different measurements The `uri` in garak's REST generator config is a scoping decision disguised as a URL. Each candidate answers a different question, with a different denominator, and a report that does not name which was scanned is not interpretable. | Target of `uri` | Question it answers | What it cannot say | |---|---|---| | Vendor model endpoint behind the app | How does the underlying model behave with no application scaffolding? | Nothing about shipped risk | | Staging deployment of the app | Does the product, as engineered, resist these probes? | Nothing about live gateway, quota or abuse controls | | Live production | Does the deployed system, with its real defences and operations, resist them? | Nothing cheaply, safely or repeatably | **Model baseline.** Removing the application removes its system prompt, retrieval corpus, tool surface and output filter from the measurement. That is precisely its value: it is an *attribution* instrument. A hit that reproduces at the bare model is the model's; a hit that appears only through the app implicates the app's own context, which usually means the fix is app-side. It is never quotable as product risk. **Staging.** This is where most of an engagement belongs, because it is repeatable and has no customer blast radius. Its entire validity is drift. A staging system prompt one revision behind production, a smaller or synthetic retrieval corpus, a guardrail disabled to make development bearable, a different model version pinned in config — each silently invalidates the number without changing anything you can see in the report. Diff those four explicitly against production and record the deltas alongside the result. **Production.** Uniquely includes the gateway, the real rate limit, abuse heuristics, alerting and whether anyone notices — none of which staging usually has. ### What production actually costs Thousands of attack prompts land in real conversation logs that humans and downstream analytics will read, and that may be retained under a policy nobody consulted. The sweep consumes production quota that customers are also using. The completions are billed to a live cost centre. On-call may be paged by your own traffic. And the account can be suspended partway through, which does not merely stop the run — it destroys its denominator. ### Where the number misleads The failure specific to production is the flattering truncation. When abuse controls start returning 403s or a WAF interposes a block page, those bodies are stored as replies unless you have handled them. They contain no policy violation, so detectors find nothing, and the failure rate *falls* over the second half of the run. The report reads as a system that got safer under sustained attack; what actually happened is that the tail of the sweep never reached the model. Distinguishing the two requires inspecting stored replies and HTTP statuses, not reading the rate. The staging failure is the opposite and quieter: a clean staging result quoted as the product's posture when the production prompt revision, filter configuration or retrieval content differs. And the model-baseline failure is the oldest one — a vendor-endpoint number presented as the shipped product's risk, which is a category error, not a rounding error. ### The defensible plan Run the bulk of the sweep against a staging build that mirrors production, with the mirroring verified rather than assumed. Then run a small, narrowly scoped production pass to confirm findings that matter, agreed in writing with the operators: a time window, a request ceiling, a kill switch, and a named contact who can abort. Authorisation attaches to the credential in `headers` and the `uri` it is used against — that pair is what defines the environment and tenant you touched, so scope the written permission to exactly that pair and keep the run log as evidence. ### What the report must state The endpoint class scanned; the app or prompt revision it was running; the credential's tenant; the probe families included and excluded; the total request count; and the wall-clock. Without those six, two scans of "the same assistant" are not comparable and neither can be re-run. If the run was cut short — by suspension, by a deploy, by quota — say so, and treat the suspension itself as a genuine finding about live abuse controls that staging could never have produced.

  • A hit reproduces on the staging app but not on the bare model endpoint. What does that suggest?
    The application's own context — its system prompt, tools or retrieved documents — is contributing to the behaviour, so the fix is likely app-side rather than model-side.
  • What single check most improves the credibility of a staging-only result?
    An explicit, recorded diff of staging against production for prompt revision, guardrail configuration and retrieval corpus, so readers know which differences the number is exposed to.
  • Production scanning gets your account suspended mid-run. What do you report?
    Two things: the partial coverage with its request count and excluded families, and the suspension itself as a positive finding about the live abuse controls that staging could not have shown.

saying these in an interview costs you the question

  • Scans production because it was the easiest URL to obtain.
  • Presents a bare-model result as the shipped product's risk.
  • Assumes staging mirrors production without checking prompt revision, filters and retrieval content.
  • Has no abort contact, time window or written authorisation tied to the credential used.
  • Omits the scanned endpoint and app version from the report.

context