PyRIT
You will learn PyRIT's orchestrator/target/converter/scorer model and how it automates multi-turn probing of an LLM target. Interviewers probe it because it is the reference open framework for structured, repeatable AI red-team runs.
on this pageshowhide
explore
- Execution Objects23 questions
- Attack Strategies4 questions
- Prompt Targets4 questions
- Prompt Converters5 questions
- Scorers5 questions
- Conversation Memory5 questions
- Driving a Run15 questions
- Multi-Turn Loops5 questions
- Throughput and Cost5 questions
- Untargeted Surfaces5 questions
- Writing Your Own Pieces10 questions
- Custom Scorers5 questions
- Custom Targets5 questions
questions
page 2 of 2An application team disputes one item in your report from a PyRIT engagement. How do you get from that report item back to the exact stored exchange that produced it, and what does the transcript prove and not prove?
basics
~20 sEach stored turn carries identifiers tying it to its conversation and run, so a report item can point at the exact exchange that produced it. Keep those identifiers in the finding. The transcript proves what happened once against that configuration - not that it reproduces, since the target samples and may have changed.
A finished PyRIT engagement is stored, and you now want to relabel it with a stricter judgment of what counts as a hit without spending another call against the target. What makes that possible, and what can relabelling stored transcripts not tell you?
basics
~20 sBecause the run stored every prompt and response, you can point a different scorer at the saved transcripts and relabel them without spending another target call. What that cannot tell you is what the attack would have done under the new judgment: the strategy branched on the old verdicts, so the turns it never explored simply do not exist.
Your only access to a support assistant is an authenticated HTTP API with a bespoke JSON body, a rotating session token and a streamed reply. How do you get it under PyRIT as a prompt target, and what does that wrapper silently bound about the run?
basics
~20 sUse PyRIT's generic HTTP target: give it the request template with a placeholder where the prompt goes, plus logic that pulls the assistant text out of the reply. You then own auth refresh, retries and rate limiting yourself. Anything the API never returns - system prompt, tool calls, retrieved context - stays untested.
Your PyRIT run is too slow, so you raise its concurrency. What does that change about the run's wall-clock time and about its cost?
basics
~20 sConcurrency changes arrival rate, not work. The same calls are made, so spend is unchanged; you only reach the quota sooner. Wall clock improves until the tightest-quota of the three endpoints saturates, then stops improving, and past that point the other two legs idle behind the bottleneck.
A PyRIT run bills an adversarial model, the prompt target and a scorer. If you must cut spend, which of those three can you move to a cheaper or locally-hosted model, and what does each substitution cost you?
basics
~20 sThe target cannot change — it is the system under test, and swapping it invalidates the run. The adversarial model and the scorer can. A weaker adversarial model finds fewer breaches, so results understate risk. A weaker scorer misjudges responses, which corrupts the number the whole run reports.
A PyRIT multi-turn run halts the moment its objective scorer marks a turn as success. What does that early stop leave out of your engagement report, and how would you run it differently?
basics
~20 sStopping at first success gives an existence proof and nothing else: no severity, no repeatability, no sense of whether it took two turns or nine. One run is a single sample of a stochastic search. Repeat each objective several times, record turns-to-success across runs, and hold the budget fixed across targets you compare.
A PyRIT run against a support assistant returned no hits, but an incident later shows the same behaviour was reachable in production. How do you work out whether the target wiring, rather than the attack content, produced the clean result?
basics
~20 sCheck reachability before blaming the prompts. Confirm the incident path was in the target list at all, replay a positive control through each target, read the stored attempts for empty bodies, errors, auth rejections and timeouts, and compare the target's endpoint, tenant and deployment identifier against what production served.
PyRIT is wired to a support assistant's chat endpoint and the run is clean. Which other intake paths of that application would you wire as separate prompt targets before calling the surface covered, and what does wiring each one require?
basics
~20 sAny path whose content reaches the same model: document or ticket ingestion, file and attachment parsing, transcribed voice or extracted image text, scheduled and batch jobs, webhook consumers, admin or internal endpoints, other tenants and older routed API versions. Each needs its own send path, matching credentials, and somewhere the effect can be observed.
Mid-engagement you fix a bug in your custom PyRIT scorer. You re-score last week's stored transcripts instead of re-running the attacks. What does that recover, and what does it not?
basics
~20 sRe-scoring recovers verdicts on responses you already collected: single-turn runs are fully repaired and missed hits reappear. It cannot recover the run itself. In multi-turn attacks the old scorer chose when to escalate and when to stop, so the transcripts hold only the paths it steered into; turns never taken stay unexplored.
You have a fixed call allowance against a metered chat endpoint for one PyRIT engagement. How do you split it between many single-send runs across a broad objective list and a few deep multi-turn runs, and how do you choose the per-run turn cap?
basics
~20 sSpend the first slice on broad single-send runs to find where the target is soft, then concentrate deep multi-turn runs on those objectives. Set the turn cap from observed successes, not from a guess, and reserve a slice for repeats, because one run against a non-deterministic endpoint is a sample rather than a result.
Your team's standing PyRIT suite runs every seed prompt through a matrix of prompt-converter variants. How do you decide how many variants the suite carries, and what does a per-converter success number actually measure when you publish it?
basics
~20 sSize the matrix by distinct transform classes, not by count: variants multiply target and scorer calls linearly and mostly re-test the same input path. A per-converter number measures which surface form got past the input handling and the judge — not distinct model weaknesses, and not attack-surface coverage.
Your team runs PyRIT campaigns against several internal endpoints on a recurring schedule. What should the scorer be allowed to decide on its own, and what has to reach a human before it counts?
basics
~20 sLet the scorer decide only inside the run: when to stop a conversation and how to rank what gets read first. Anything leaving the team as a finding needs a human who read the transcript. Keep a fixed labelled transcript set as the scorer's own regression test, re-run whenever the scorer changes.
For a PyRIT engagement you can point the prompt target at the raw model endpoint behind a product or at the product's own chat API. How do you decide which surface to wrap, and how does that choice change what your report may claim?
basics
~20 sDecide from the question being asked. The raw endpoint measures the model's own behaviour with no product defences; the product API measures what a user can actually reach, defences included. Each report claims only its own surface. Where budget allows, wrap both and read the gap between them as the value of the defences.
You have a fixed spend for a PyRIT engagement. How do you divide it between many objectives at a short turn budget and few objectives explored deeply, and how do you report what the money bought?
basics
~20 sSplit the budget: a broad shallow pass to find which objectives show movement, then deep turn budgets spent only on those. Cap spend per objective so one runaway conversation cannot eat the pass. Report cost per objective and per confirmed finding, and state which objectives were cut short rather than cleared.
You have a fixed pool of adversarial turns to spend across roughly 40 objectives in a PyRIT engagement. How do you allocate the turn budget between deep runs on a few objectives and shallow runs on many, and how do you report the choice?
basics
~20 sDo not spread it evenly. Run a short screening pass over all objectives, then spend deep budgets only on the ones whose transcripts showed the target trending rather than refusing flat. Report every negative with the budget that produced it, treat shallow negatives as untested rather than safe, and keep budgets fixed between retests.
Your team re-runs the same PyRIT target set against a fast-moving product every release. How do you keep the wired target set from silently drifting behind the surfaces the product actually ships?
basics
~20 sDerive the target set instead of remembering it. Reconcile wired targets against the artefacts the platform deploys from — routes, service registry, data-flow records — and flag any model-reaching path with no target and no expiring waiver. Make registering a target part of shipping a new intake path, and track wired-over-deployed as a first-class number.
Four teams each wrote their own PyRIT prompt target for their own service, and each now reports an attack-success rate to you. What do you require before you are willing to compare those numbers across teams?
basics
~20 sRequire every adapter to pass one shared conformance check: a known refusal, a known compliance, a forced error and a forced timeout, each landing in PyRIT's memory as a distinct expected outcome. Also require the same prompt set, scorer and turn budget. Otherwise the differences measure adapters, not services.
showing 31–48 of 48