In promptfoo, you point the HTTP provider at your own chat service, which answers with a JSON envelope. What is the job of the response transform, and what do the assertions grade if you never declare one?
answer
- envelope vs reply field
- no transform = grade the raw body
- clean report, wrong string
- error body grades as safe
- read one stored output by eye
basics
~20 sThe transform picks the assistant's text out of the HTTP response body and hands that string to the graders. Without one, promptfoo grades whatever the raw body serialises to: the whole envelope, metadata included. Assertions then match against wrapper fields, so the report looks plausible but describes the envelope, not the reply.
solid answer
~50 sThe HTTP provider knows how to send a request and read a body; it does not know which part of your body is the model's answer. The response transform is the declaration that closes that gap, mapping the parsed body to the single string every assertion, grader and multi-turn strategy will then treat as the output. Leave it out and nothing errors. The provider falls back to the raw body, so a substring assertion can match a field name, a JSON key or a trace id that happens to contain the token you looked for, and a model-graded check is asked to judge a blob of metadata. Both produce clean-looking rows. The practical rule: after wiring a new target, open one stored result and read the graded output as text. If it still has braces in it, the transform is wrong, and every number in that run is about your envelope rather than your app.
go deeper
Says the transform tells promptfoo which field of the HTTP response is the assistant's reply, and that without it the whole body gets graded.
Adds that nothing errors when it is missing — assertions match against wrapper fields and produce clean-looking but meaningless rows — and knows to eyeball one stored output.
Covers error bodies and streaming, explains why an empty-string fallback turns failures into passes, and separates the app's own guard verdict from the model text.
Treats the transform as owned code with its own regression risk: a control case per target, a shape change in the app as a breaking change to the eval, and a rule that no run is reported before one raw response is read.
**What the provider knows, and what it cannot know.** promptfoo's HTTP provider is declared in the eval config under `providers` and carries a `url`, a `method`, `headers` and a `body` template into which each test case's variables are rendered. Once the response comes back, the provider hands it to the eval loop, and there its knowledge stops. For a first-party model provider such as `openai:chat:...` the response schema is fixed, so promptfoo pulls the completion out itself. Your own service has no such contract: the reply might sit at `message`, at `answer`, at `data.text`, or arrive as chunks you have to join, alongside citations, a moderation verdict, latency, token usage and a conversation id. The response transform — `transformResponse` in current promptfoo releases, `responseParser` in older ones — is the single place you declare which of those is the thing under test. It is a small expression or function that you own, evaluated once per HTTP response, whose return value becomes `output`: the one string that every assertion, every `llm-rubric` judge, every multi-turn strategy and every stored result row will treat as what the application said. ```json {"conversation_id":"c-8812","blocked":false,"data":{"text":"I can't help with that."},"usage":{"tokens":142}} ``` **What happens when you declare nothing.** Nothing errors. The provider falls back to the serialised raw body, so the graded string is that whole object. A `contains` assertion hunting a refusal phrase still matches — and it would match just as happily if `data.text` were empty and the phrase appeared in a citation or an error label. A model-graded rubric is worse: it is handed a blob of metadata and asked whether the output is harmful, and a judge shown JSON keys will nearly always say it is not. That "not harmful" is recorded as a pass. **What it costs.** At runtime the transform is free — local code, no extra call, negligible time. The cost is ownership and blast radius. Do the arithmetic on a modest suite: 300 cases, a two-turn strategy, one judge call per graded turn is roughly 600 target calls and 600 judge calls per run. A wrong transform does not spoil one row, it spoils all 1,200, and you pay for them again on the rerun on top of the hours spent triaging a report that was never about the app. It is also code that drifts: the day someone adds an error envelope or turns on streaming, the transform keeps returning *something*, and no test tells you the meaning changed. **Where the number misleads.** Four specific readings go wrong, and they all move the pass rate in the flattering direction. - *Error bodies.* A 401 or 500 envelope has no reply field. A transform that returns an empty string on it turns every failed call into a case that trivially satisfies "did not produce disallowed content". The pass rate then rises as the target gets more broken, which is the single most common way a live-target run reports clean. - *Wrapper safety flags.* If `blocked` sits inside the graded string, an input guard firing and the model declining look identical to the grader. That is exactly the distinction remediation depends on: one says the filter works, the other says the model happened to behave. - *Substring luck.* Field names, trace ids and enum values are text. Assertions keyed on tokens match them without the reply being read at all. - *Tool-call and streaming turns.* If the first response is a tool invocation with empty text, or a stream that ends early, the graded string is empty or truncated and the assertion passes on nothing. **What I check before quoting a number.** Open several stored raw responses — different cases, not just the first — and read the graded output as plain text; braces, an empty string, or the identical string on every row all point at extraction rather than safety. Add a benign control case whose correct answer you already know, so a substantive non-empty reply must reach the grader for the suite to go green. Force a 401 and a 500 on purpose and confirm those rows go red rather than green. Decide explicitly what the transform does on an error body — raising or returning a sentinel that fails an assertion beats returning an empty string. And confirm the config key against the promptfoo version you are running rather than copying a config from a blog post, because these names have moved between releases.
- Your app streams the reply as chunks. What changes about the transform?It has to join the chunks into one final string before the graders see it, and decide what to do with a stream that ends early — an incomplete reply graded as the full answer is a false pass.
- How would you make a wrong transform loud instead of silent?Add a control case whose correct output you know — a benign question that must return substantive text — and assert on it. If the graded string is the envelope or empty, that case fails immediately instead of everything passing.
- Should the transform strip the app's own 'blocked' flag out of the graded output?Yes for the text under test, but keep the flag as separate metadata. Mixing it into the graded string makes a guard block and a model refusal look identical in the report.
It is the difference between grading the letter and grading the envelope it arrived in. The postmark, the return address and the stamp are all text, so a checker looking for a word can find one without the letter ever being opened.
saying these in an interview costs you the question
- Assuming the tool 'just knows' which field holds the reply because the body is JSON.
- Reporting a run as clean without ever looking at one graded output string.
- Writing a transform that returns an empty string on an error body, so failed calls score as safe.
- Grading the whole envelope and then treating a wrapper flag such as 'blocked' as if it were model text.