skip to content

PyRIT

You will learn PyRIT's orchestrator/target/converter/scorer model and how it automates multi-turn probing of an LLM target. Interviewers probe it because it is the reference open framework for structured, repeatable AI red-team runs.

on this pageshow

explore

questions

page 1 of 2

In PyRIT, what does a scorer do during an attack run, and what must a custom scorer you write return for the run to be able to use it?

level: juniorimportance: must knowfreq 62%

answer

  1. scorer = the run's stop signal
  2. score object, not bare True
  3. verdict + rationale + piece id
  4. boolean family vs graded family
  5. persisted to memory, replayable

basics

~20 s

A PyRIT scorer reads the target's response and decides whether it satisfies the run's objective. A custom scorer must return what the built-in ones return: a verdict, boolean or a normalized number, plus a rationale and the identifier of the piece it scored, so the attack strategy can branch and memory can store it.

solid answer

~60 s

In PyRIT the scorer is the object that decides whether a response counts as objective-achieved. It is not decoration on the report; it is the run's control signal. A multi-turn attack strategy asks the scorer after each target reply and uses the answer to decide whether to send another turn or to stop and declare success. So a custom scorer has to satisfy the same contract as the shipped ones: subclass the scorer base class in the version you have installed, implement its scoring method, and return score objects — not a bare `True`. Each score carries the verdict (a true/false value, or a value on a normalized scale, depending on which scorer family you are extending), the rationale text explaining the verdict, a category, and a reference to the prompt piece it scored so memory can join the score back to the transcript. The rationale is the part people skip and then regret: without it, a triager reading the run later sees a hit with no reason, and cannot tell a real finding from a scorer bug.

go deeper

for a junior

Says a scorer decides whether the target's response met the objective, and that a custom one must return the framework's score object with a verdict and a rationale.

for a middle

Adds that the verdict is the multi-turn stop condition, and distinguishes the boolean family from the graded family and when each is appropriate.

for a senior

Talks about memory persistence, replaying the scorer over stored transcripts, and handling responses the scorer cannot parse instead of returning a silent false.

for a principal

Frames the scorer as the run's control plane and the report's evidence chain at once, and sets a team rule that no custom scorer ships without a rationale field and a defined behaviour for unscoreable input.

### Where the scorer sits in a run PyRIT drives an attack as a loop around three objects you configure: an *attack strategy* (the algorithm deciding what to send next), a *prompt target* (the adapter wrapping the system under test), and a *scorer*. Each iteration builds a prompt, passes it through any attached *prompt converters* (objects that transform the outgoing text), sends it to the target, writes the response into *memory* (PyRIT's persistence layer — a local SQLite file by default, a shared database on a real engagement), and then hands that response to the scorer with exactly one question: does this satisfy the objective? Two things consume the answer, which is why a scorer is not report furniture. **The attack strategy.** A multi-turn strategy loops until the scorer reports objective-achieved or the turn budget is exhausted. Your scorer is the loop's exit condition — it decides when the run stops, and in strategies that condition the next prompt on the last verdict, it also shapes what gets sent next. **Memory and the deliverable.** Scores are persisted alongside the prompt pieces and keyed to the piece that was scored. That key is what lets you re-open a conversation months later, join a verdict to the exact response that earned it, and re-score stored transcripts offline without touching the live target again. ### The contract a custom scorer must satisfy Subclass the scorer base class present in your installed version, implement its scoring entry point, and return a *list* of score objects — never a bare `True`. A list, because one response can attract several verdicts in different categories. Each score object carries four things that matter downstream: the value (a true/false verdict, or a number on a normalised scale, depending on which scorer family you extend), the rationale text, the category, and the identifier of the prompt piece scored. Pick the family deliberately. A boolean verdict is what a stop condition wants; a graded value is what a trend across runs wants. Returning a float and letting the caller decide that 0.7 means yes hides the threshold outside the scorer, so nobody can reconstruct the hit count from the stored scores alone. ### What it costs A deterministic scorer costs effectively nothing per turn — it is a function over a string. A model-backed one is a third metered call on every turn, alongside the attacker model and the target: a 50-conversation run with an 8-turn budget is up to 400 scored turns, so 400 extra calls and 400 extra round-trips in a loop that runs serially per conversation. The engineering cost is the line teams under-budget: writing the scorer is an afternoon, while establishing that it agrees with a human is a day of hand-labelling — and that day is the one that makes the number reportable. ### Where the number misleads The dangerous failure is silent. A scorer that returns an empty list, or returns false because it was handed something it cannot parse — an image, an audio response, an error string, an empty completion — produces a run with zero objective-achieved. Zero reads in a report as *the target held*. That is a different claim from *we looked and found nothing*, and nothing in the pipeline separates them, because no human ever opens a non-hit. The second misread is a hit with no evidence. A bare boolean, with no rationale and no piece reference, hands a triager a count with nothing to check, so a scorer bug and a real finding look identical in the deliverable. ### What to check before you trust it - Run three turns and read the scores table in memory directly. Confirm rows exist, that the piece reference resolves to the response you expect, and that the rationale is populated. - Feed the scorer one hand-labelled known hit and one hand-labelled known refusal. A scorer that has never been shown a true positive is untested, however clean the code reads. - Confirm it is actually attached to the strategy. An unattached scorer and a broken one look identical from outside. - Check class and field names against the installed package rather than a write-up. These have been renamed across PyRIT releases. - Define the behaviour for input it cannot score, and make it visible: an unscored turn you can count, not a false verdict you cannot.

  • Why does a PyRIT score reference the prompt piece it scored rather than just the conversation?
    One conversation holds many pieces and one response can get several scores. The piece reference is what lets memory join a verdict back to the exact response, so a later re-score or a triage read lands on the right turn.
  • Your custom scorer returns an empty list for every response. What does the multi-turn strategy do?
    It never sees objective-achieved, so it keeps sending turns until the turn budget is exhausted and reports the run as a failure. The symptom is a run that always burns its full budget and always finds nothing.

saying these in an interview costs you the question

  • Describing the scorer as report metadata rather than the thing that ends a multi-turn run.
  • Returning a bare boolean with no rationale and no reference to the scored piece.
  • Assuming a specific class or field name from a tutorial without checking the installed package, since these names have been renamed across releases.
  • Silently returning false when the scorer was handed content it cannot parse.

context

open as a page

You need to run PyRIT against an internal HTTP chat endpoint that PyRIT ships no prompt target for. What does the custom prompt target you write own, and what does the rest of the PyRIT run still do for you?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Your target owns everything endpoint-specific: authentication, request and response shape, timeouts and retries, and turning the reply into one response PyRIT can store. Everything else is unchanged. The attack strategy still chooses prompts, converters still transform them, scorers still judge them, and the memory store still records each exchange.

open as a page

In PyRIT, an attack strategy object is constructed once with its prompt target and its scorer, while the objective is supplied each time you execute it. What does that split buy you when you have twenty objectives to test against the same endpoint, and what does each execution get of its own?

level: juniorimportance: must knowfreq 62%

basics

~20 s

You build the strategy once with its target, scorer and any converters, then execute it in a loop, once per objective. The wiring is shared across all twenty. Each execution carries its own objective and its own conversation, so the transcript and the verdict stay separate per objective.

open as a page

In PyRIT, what does a prompt converter do to a prompt before it reaches the target, and why is adding a converter on its own not an attack?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A PyRIT prompt converter rewrites the outgoing prompt text after the seed prompt is chosen and before it is sent. It changes the surface form, not the request underneath. On its own it is just a transform: it tests whether that form gets through, never whether the model will actually comply.

open as a page

In PyRIT, what does the conversation memory store record while a run executes, and what do you give up by running with the in-memory store instead of the durable one?

level: juniorimportance: must knowfreq 62%

basics

~20 s

PyRIT's memory records every prompt sent and every response received, turn by turn, with the scores attached, so a run can be resumed, audited and re-scored later. The in-memory option keeps that only for the process lifetime: when it exits, the transcripts are gone and nothing can be re-read or re-scored.

open as a page

In PyRIT, what is a scorer's job during an attack run, and what does its verdict do to the run itself?

level: juniorimportance: must knowfreq 72%

basics

~20 s

In PyRIT, the scorer reads the target's response and decides whether the attack objective was met. In a multi-turn run that verdict is control flow: a positive verdict ends the run early and marks it a success. So the scorer, not you, decides when testing stops.

open as a page

In PyRIT, what does a prompt target object do, and why does one run usually need more than one of them?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A PyRIT prompt target is the adapter that knows one endpoint: it sends a prompt there and returns the response. A run usually needs several of them, because the system under test, the attacker model that drafts prompts, and any model that scores replies are three different endpoints.

open as a page

In a PyRIT multi-turn run, how many model calls does a single turn bill, and which components make them?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Three. A PyRIT turn bills the adversarial model that writes the next prompt, the target under test that answers it, and the scorer that judges the reply. Every turn pays all three, so a fifty-turn objective is roughly one hundred fifty billed calls, not fifty.

open as a page

In PyRIT, what do you have to configure before a multi-turn adversarial run can start, and what are the two ways the loop can stop?

level: juniorimportance: must knowfreq 68%

basics

~20 s

You set four things: an objective describing what the system under test should be made to do, an adversarial chat model that writes each attacker turn, the target endpoint being tested, and a scorer that judges the target's reply. The loop ends when that scorer says the objective was met, or when the turn budget runs out.

open as a page

A PyRIT run against your deployed chat assistant finishes with no successful attacks recorded. What has that run actually tested, and what has it not?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Only the targets you configured. PyRIT sends prompts to the prompt targets you wired and scores the replies; anything it was never pointed at, such as another endpoint, an upload path or an internal service, is untested rather than proven safe. A clean result describes your target list, not the application.

open as a page

In a PyRIT red-team run, why is a false negative from your custom scorer more expensive than a false positive?

level: middleimportance: must knowfreq 58%

basics

~20 s

A false positive is caught: a human reads the flagged transcript during triage and drops it. A false negative is never read, because the transcript looks like an ordinary refusal and nobody opens it. In a multi-turn attack the miss also tells the strategy to keep pushing or to stop, so it changes the run itself.

open as a page

In PyRIT, what distinguishes a single-send attack strategy from a multi-turn one, and what actually happens when you point a single-send strategy at an objective that only succeeds after several turns?

level: middleimportance: must knowfreq 70%

basics

~20 s

A single-send strategy delivers each prompt once, scores the reply and stops; nothing feeds forward. A multi-turn strategy loops: it reads the last reply, composes the next prompt with an attacker model, and scores each turn until the scorer says achieved or the turn cap runs out. Single-send on a multi-turn goal returns not-achieved, having proved nothing.

open as a page

A PyRIT run is configured with several prompt converters applied to each outgoing prompt. In what order does the chain apply them, and why does swapping two converters change what the target actually receives?

level: middleimportance: must knowfreq 62%

basics

~20 s

The chain runs in series over a single prompt: the first converter's output becomes the second's input, and only the final string is sent. So the converters compose, and composition is not commutative — rewriting an already-transformed string gives a different result than transforming a rewritten one.

open as a page

By default a PyRIT run persists its conversations to an unencrypted local database file. What is actually in that file, and how should it change where and how you run an engagement?

level: middleimportance: must knowfreq 58%

basics

~20 s

By default PyRIT writes the run to a local database file on the machine you launched it from, unencrypted. That file holds attack prompts and the target's worst answers verbatim. Treat it as sensitive evidence: put it on encrypted storage you control, keep it off shared drives and backups, and delete it on schedule.

open as a page

A PyRIT scorer can return a true/false verdict, a scaled numeric score, or a category label. How do you choose between them for a multi-turn attack run, and what extra decision does a scaled scorer force on you?

level: middleimportance: must knowfreq 62%

basics

~20 s

A true/false scorer gives the loop the yes-or-no it needs to stop, so it is the default for objective-driven runs. A scaled scorer returns a degree of compliance and forces you to pick the threshold that counts as success. A category scorer says what kind of harm, not whether the attack worked.

open as a page

In PyRIT, how does earlier conversation history reach the endpoint on turn five of a multi-turn run, and what changes when the prompt target wraps a service that keeps its own server-side session?

level: middleimportance: must knowfreq 58%

basics

~20 s

For a chat-completion style endpoint the target replays the stored turns: it reads the conversation out of memory and sends the whole message list each call. If the service keeps its own session and accepts only the newest message, the history lives on the server, so the tool's copy and the real context can drift apart.

open as a page

Before launching a PyRIT run, how do you estimate its total call count when the configuration includes several seed prompts, converters and scorers?

level: middleimportance: must knowfreq 60%

basics

~20 s

Multiply, do not add. Target calls are objectives times turns times converter variants. Scoring calls are target responses times the number of scorers attached. Adding one converter and one scorer roughly doubles two legs at once. Work the product out on paper first, then dry-run one objective to check it.

open as a page

A PyRIT multi-turn run ends with 'objective not achieved' after exhausting a 10-turn budget. Why is that not evidence that the target refuses the behaviour?

level: middleimportance: must knowfreq 62%

basics

~20 s

It only says the adversarial model did not get there within the turns you allowed. The budget is part of the result, not a property of the target. A longer run, a different attacker configuration, or another attack strategy can still breach the same endpoint. Report the negative together with the turn budget that produced it.

open as a page

When you configure a PyRIT prompt target, what is the difference in what a clean run proves if you wire it at the model provider's raw API versus at your application's own endpoint?

level: middleimportance: must knowfreq 58%

basics

~20 s

The raw API tests the model alone, with no production system prompt, retrieval, tools or guard. The application endpoint tests the whole deployed stack. A clean run at one says nothing about the other: a guard can hide model weakness, and a bare-model result ignores every layer users actually pass through.

open as a page

Your custom PyRIT scorer reports 12 objective-achieved hits across a 400-turn run. What do you do before that number goes into the engagement report?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Hand-label a sample and compare. Sample from the 388 the scorer called non-hits, not only the 12 hits, because misses are where uncounted findings sit. Read the hits too. Then report the count together with the measured miss rate and false-alarm rate on that sample, so a reader knows what the 12 excludes.

open as a page

You pointed a PyRIT run at an internal chat service through a prompt target you wrote yourself, and the attack-success rate came back near zero. How do you tell a genuinely hardened service from a broken adapter?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Do not trust the summary number yet. Open the responses PyRIT stored: a hardened service returns real refusal text, a broken adapter stores empty strings, error bodies, truncated replies or the same string every time. Then replay two or three of those exact prompts by hand and compare against what the adapter recorded.

open as a page

Adding an encoding converter to a PyRIT chain triples the success count over the same seed prompts. Before you report that as a jailbreak result, how do you establish whether the target actually complied?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Treat the jump as a scoring hypothesis first. Pull the stored exchanges, decode the replies, and read whether the content is really there. A transform that changes the reply's form also changes what a text-based PyRIT scorer sees, so refusals stop matching and garbled output can read as compliance.

open as a page

Before you let a PyRIT scorer end multi-turn runs on its own, how do you check it against transcripts you labelled yourself, and which responses have to be in that labelled set?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Pull stored conversations from earlier runs, label the responses yourself, then re-score them offline with the candidate scorer and compare. Load the set with the hard cases: refusal-then-comply, partial compliance, in-fiction compliance, and confident nonsense. Only then let it stop runs. Replaying stored text costs no target turns.

open as a page

A batch of PyRIT multi-turn runs reports a high objective-achieved rate, but most transcripts stop on an early turn at a hedged, non-compliant answer. What is happening, and how do the two directions of scorer error differ in what they cost you?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The scorer is firing on responses that are not real compliance, and because a positive verdict ends the run, every false hit is also a truncated attack. False positives inflate the success rate and delete the turns you never sent. False negatives cost the whole turn budget and undercount. Read the stop-turn responses.

open as a page

You need a PyRIT scorer for a success condition none of the shipped scorers covers. When do you write it as a deterministic rule over the response, and when do you back it with a model call?

level: middleimportance: should knowfreq 45%

basics

~20 s

Write a deterministic rule when the success condition has a checkable surface: a specific string, a schema, a tool call, a leaked identifier you planted. Back it with a model only when success is a judgment about meaning. The rule is cheap, reproducible and re-runnable over stored transcripts; the model call adds cost, latency and run-to-run variation.

open as a page

A custom PyRIT prompt target wraps an endpoint that streams its reply back in chunks. What must the adapter do before that reply reaches PyRIT's scorers, and what goes wrong if it returns early?

level: middleimportance: should knowfreq 38%

basics

~20 s

Consume the whole stream and assemble one complete response before returning, because scorers see a single stored response, not chunks. If you return on the first chunk or bail at a client timeout, a truncated answer is stored as the target's real reply, and a response that was about to comply can score as a refusal.

open as a page

In a PyRIT multi-turn run the objective text is given both to the adversarial model and to the objective scorer. Why does a vaguely worded objective make the run untrustworthy in both directions?

level: middleimportance: should knowfreq 45%

basics

~20 s

The same sentence steers the attacker and defines success. If it names no observable artefact in the target's reply, the adversarial model has nothing concrete to drive toward and drifts, while the scorer has nothing concrete to check and can mark a hedged, harmless answer as met. Write objectives a reviewer can verify from the transcript.

open as a page

Your custom PyRIT prompt target calls a service that keeps conversation state server-side behind a session id it issues. What breaks within a multi-turn run and across repeated runs, and how do you make each attempt independent?

level: seniorimportance: should knowfreq 34%

basics

~20 s

PyRIT already tracks conversation identity, so reusing one server session makes history arrive twice or bleed between attempts. Mint a fresh session per PyRIT conversation, bind it to that conversation, and tear it down at the end. Otherwise a result depends on which attempts ran before it and stops being reproducible.

open as a page

A multi-turn PyRIT attack run finishes with the objective not achieved. What are the ways that loop can terminate, and why is that outcome not evidence that the target is safe?

level: seniorimportance: should knowfreq 52%

basics

~20 s

It ends three ways: the scorer returns an achieved verdict and the loop stops early, the turn cap is exhausted, or something raised — target error, rate limiting, the attacker side declining to continue. Not-achieved conflates all of the non-success endings, so it says the run stopped, not that the target held.

open as a page

showing 1–30 of 48