A deployed assistant's HTTP endpoint answers with a JSON envelope rather than a bare string. What must garak's HTTP generator be told about that envelope, and what does a scan report look like when that setting is wrong?
answer
- generator returns strings, endpoint returns envelope
- wrong selector, clean-looking report
- empty replies look safe
- streaming truncation at first chunk
- print one benign reply verbatim
basics
~20 sIt must be told where in the response body the assistant's text lives, because the endpoint returns an envelope, not a bare string. If that selector is wrong, garak stores the wrapper, an error blob or an empty string, and every attempt is judged against text the model never produced.
solid answer
~50 sgarak's contract is simple: a generator returns reply **strings**. The HTTP generator therefore needs a selector saying which part of the response body is the assistant's message — a path into the JSON, or an equivalent extraction rule. When that selector is wrong, nothing fails loudly. The run completes, the report renders, and the numbers are garbage in a specific direction: - Extract the **whole envelope** and every stored reply carries metadata, ids and role labels that the model never wrote. - Extract an **empty or null field** and every attempt is an empty string. - Extract the **wrong turn** — an echo of your own prompt, or a system field — and you are scoring your own attack text. Detectors judge whatever string was stored, so both hits and misses become meaningless. The fix is procedural, not clever: send one benign prompt, print the stored reply verbatim, and compare it to the reply the app shows a human user.
code
json · 9 lines{
"conversation_id": "c-8831",
"status": "ok",
"output": [
{ "role": "user", "content": "<the probe prompt, echoed back>" },
{ "role": "assistant", "content": "<the reply garak must score>" }
],
"usage": { "input_tokens": 118, "output_tokens": 240 }
}go deeper
Knows the reply text has to be pulled out of a JSON envelope and that the path is configured on the generator.
Enumerates the wrong-selector failure modes and explains why the run still produces a clean-looking report.
Makes a verbatim benign round-trip a gate before any sweep, and covers streaming accumulation and non-2xx bodies.
Treats reply-extraction verification as a required, recorded step of the methodology so a report's numbers are defensible to an auditor.
### The contract the extraction setting satisfies Downstream of the generator, garak deals only in strings. Everything that makes an HTTP reply structured — status, conversation and message ids, role labels, tool calls, citations, safety metadata, token usage, streamed chunks — must be collapsed by the adapter into one message string per generation. That collapse is lossy, and it is yours to specify. In garak's REST generator config it is two keys: `response_json` says the body is JSON rather than raw text, and `response_json_field` says which part of it is the assistant's message — a top-level key name, or a JSONPath expression such as one selecting the last element of an `output` array and its `content`. If the endpoint streams, the generator must also accumulate chunks into a whole message before returning. ### What it costs to get wrong Nothing, at run time — and that is the problem. A wrong `response_json_field` costs you the entire sweep: the full request count, the full completion bill, and the full wall-clock, all spent producing a corpus that cannot be triaged into correctness afterwards. On a 15,000-request run against a metered endpoint that is hours and a real line item, discovered only when someone finally reads a stored reply. There is no partial recovery: detectors cannot be retuned to fix text the model never produced. The run is redone. ### How the number misleads, failure mode by failure mode - **Whole envelope stored.** Substring and pattern detectors now match against JSON key names and metadata as well as content. You get false hits where a key or a role label coincides with a matched term, and false misses where the real text is present but re-encoded or escaped. - **Empty or null field.** Every attempt is an empty string, every detector abstains, and the report is beautifully clean. A sweep of a chatty production assistant that returns near-zero hits on *every* probe family is far more likely to be a broken selector than a uniquely robust deployment — uniform cleanliness across unrelated probes is the tell. - **Wrong turn selected.** If the path lands on the echoed user turn rather than the assistant turn, garak is scoring your own probe text. Detectors that look for markers of the attack itself then fire on nearly every attempt, and the deployment looks catastrophically broken. - **First streamed chunk only.** Replies are truncated at a fixed prefix. Long-form violations that appear after a polite preamble are never seen, so the run under-reports exactly the failures that matter most. - **Non-2xx bodies stored as replies.** Throttle and gateway-error payloads enter the corpus as though the model had said them. They contain no violation, so they read as passes and dilute the rate. Notice the direction: three of the five bias the number *downward*, toward "safe". A generator misconfiguration is a systematically optimistic error, which is why it survives review — nobody interrogates a good result. ### Why this is a generator problem, not a detector problem A detector's false-positive rate is a property of the detector, argued on its own terms and triaged case by case. This failure is upstream of all of that: the detector is behaving correctly on a string that never came from the model. No threshold, no detector swap and no post-hoc triage repairs it, because the evidence itself is wrong. ### What you check, in minutes Run one benign prompt through the configured generator, then open the run's JSONL log and read the stored output verbatim. It must be byte-identical to what a human user sees in the product. Repeat with a prompt whose answer you can predict, so you can tell a real reply from a plausible-looking wrapper. Repeat once more with a prompt that provokes a long answer, to prove the stream is accumulated rather than truncated. Finally, force one throttle or error and confirm you know where that body ends up. Only then is a multi-hour sweep worth launching.
- A scan of a chatty production assistant returns almost no hits on any probe. What do you check first?The stored reply strings. Uniformly empty, truncated or envelope-shaped replies point at generator extraction, not at an unusually robust deployment.
- The endpoint streams tokens. What extra care does the generator need?It must accumulate the stream into one complete message before returning, otherwise every stored reply is a truncated prefix and later content is never scored.
It is like grading an exam where the scanner captured the cover sheet instead of the answer pages: every script is marked carefully and consistently, and every mark is meaningless. The marker is not at fault, and no amount of re-marking fixes it.
saying these in an interview costs you the question
- Says a wrong reply field would obviously crash the run.
- Proposes tuning detectors to compensate for text the model never produced.
- Treats a near-zero hit count as good news without checking stored replies.
- Ignores streaming, so only a truncated prefix is ever scored.