Your only access to a support assistant is an authenticated HTTP API with a bespoke JSON body, a rotating session token and a streamed reply. How do you get it under PyRIT as a prompt target, and what does that wrapper silently bound about the run?
answer
- generic HTTP target: request template plus placeholder
- parse the reply, aggregate the stream
- token rotation and pacing are yours now
- 200 with an error body scores as a miss
- smoke-test one refusal, one normal answer
basics
~20 sUse PyRIT's generic HTTP target: give it the request template with a placeholder where the prompt goes, plus logic that pulls the assistant text out of the reply. You then own auth refresh, retries and rate limiting yourself. Anything the API never returns - system prompt, tool calls, retrieved context - stays untested.
solid answer
~50 sThe practical path is the generic HTTP-request target: you supply the request as a template with a placeholder for the prompt, plus parsing that turns the response body into the assistant's text. For a streamed reply you aggregate the chunks before handing it back, otherwise the scorer sees a fragment. Three things become your job rather than the framework's. Auth: a rotating token needs refreshing inside the target, or a long run dies partway and everything after the expiry scores as a miss. Throughput: the app's limits are tighter than a model API's, so pace requests and treat throttling as a retry, not a response. Parsing: an error or refusal returned in a normal-looking body is the dangerous case - the extractor yields an empty string and the attempt is silently recorded as a clean pass. The bound: you are testing the product surface, so anything it does not surface stays outside what any number from this run covers.
go deeper
Knows PyRIT can talk to an HTTP endpoint that is not a standard model API, and that someone has to describe the request and read the response.
Can wire the request template and response extraction, aggregate a streamed reply, and explain why the scorer needs the assembled text.
Handles token rotation and pacing inside the target, smoke-tests the parser against a known refusal and a known answer, and states what the product surface leaves untested.
Decides whether wrapping the product API is even the right instrument for the engagement, and what the report may and may not claim about the application from it.
## Getting it wired A bespoke API does not need a bespoke framework integration. PyRIT's `HTTPTarget` takes a full raw HTTP request - method, URL, headers, body - as a template, with a **placeholder token** marking where the prompt is substituted, plus a **callback** that receives the response and returns the assistant's text. When the shape is too odd for that, the fallback is to subclass the prompt-target base and implement the send method yourself; either way the rest of the run is unchanged, because everything upstream only knows the send-and-return contract. The order of work matters. 1. **Capture a real request** from the app first - browser devtools or an intercepting proxy - and replay it verbatim by hand before you template it. Guessing the body shape is the single most common way to lose an afternoon to a target that returns nothing. 2. **Then substitute the prompt** *safely*: a prompt containing a quote, a newline or a brace will break a naive string template inside a JSON body, and the resulting malformed request usually comes back as a clean-looking error rather than a crash. ## Streaming A chunked or server-sent-event reply must be consumed and joined inside the target before it is returned. Handing half-assembled text to a scorer produces errors in both directions: - **false misses**, because the violating part never arrived; - and **false hits**, because a truncated refusal ("I can't help with that, but here is") reads as the beginning of compliance. ## Auth and pacing are now yours A rotating session token needs a **refresh path** inside the target and a run that survives a mid-run rotation; without it, everything after the expiry returns an auth error body and scores as a miss. Product endpoints also throttle far more tightly than model APIs - single-digit requests per minute per user is common - so pace the target and back off on throttling instead of retrying immediately. ## What it costs The cost line here is **engineer time, not tokens**. Capture, template, escape, parse, aggregate the stream, handle token refresh, then smoke-test: realistically half a day to a day for a non-trivial API, and it is work you redo when the product changes its response shape. The runtime cost is dominated by the product's **rate limit**: a 200-attempt run at three calls per turn against an endpoint allowing a few requests a minute is hours of wall-clock, not minutes - and the attacker and scorer calls burn *your* model quota in parallel while the target calls burn the product's. Budget the run in wall-clock against the tightest limit, and pilot with a handful of objectives before committing to the full suite. ## Where the number misleads - **The dangerous case is not a crash, it is a plausible clean sheet.** An app that returns moderation refusals or internal errors as HTTP 200 with the text in a field your extractor does not read hands the scorer an empty string on every attempt. Empty strings contain no violation, so they score as misses, the run completes normally, and the report says the endpoint held. Nothing in the pipeline flags it, because from the target's point of view the request succeeded. - **The rate-limit variant distorts the denominator instead.** Attack-success rate is hits divided by attempts, and *attempts* is counted by the harness, not by the endpoint. Retry a throttled call immediately and one rate-limited request becomes a burst of recorded attempts the application never processed, pushing the reported rate down against a target that mostly never saw the prompts. The same applies to attempts sent after a token expiry. ## What I check - **Before any long run:** send one prompt that should get an ordinary answer and one that should be refused, then read both stored responses out of memory and confirm each is readable, non-empty and complete. - **During and after:** check the distribution of stored response lengths - all-empty or all-identical means the parser is wrong, not that the model is consistent - and reconcile the harness's attempt count against the application's own request log or your proxy log, so throttled and errored calls are not silently counted as tested prompts. - Deliberately run past the token's lifetime in a short pilot to prove the refresh path works. ## What the wrapper bounds This target sees the product's chat surface and nothing else. The system prompt, tool calls, retrieved documents and any other entry point into the same model are invisible to it. A result from this run is a statement about that chat surface on that date - not about the model, and not about the application as a whole - and the report should say so explicitly rather than let a reader generalise it.
- What is the cheapest check that a hand-wired HTTP target is actually working before you start a long run?Send one prompt that should get a normal answer and one that should get a refusal, then read both stored responses. If either comes back empty or truncated, the parser is wrong, not the model.
- Why is an immediate retry on a throttling response worse than pacing the target?It converts one rate-limited call into a burst of failures that get recorded as attempts, so the run reports low success against an endpoint that mostly never received the prompts.
An error field your parser never reads is a microphone that was never plugged in: the recording is silent, and you write down that nobody in the room said anything objectionable.
saying these in an interview costs you the question
- Trusting a run whose target was never smoke-tested against a known refusal and a known normal answer
- Letting an unparsed error body be scored as a non-violation
- Passing streamed chunks to the scorer without assembling them
- Retrying throttled requests immediately and reporting the failures as attempts
- Reporting the result as coverage of the application rather than of its chat surface