skip to content

Prompt Converters

A converter rewrites the payload on the way out, so a chain tests encodings not the technique underneath, and a hit can mean the detector missed. Interviewers ask why one alone is not an attack.

on this pageshow

explore

questions

5

In PyRIT, what does a prompt converter do to a prompt before it reaches the target, and why is adding a converter on its own not an attack?

level: juniorimportance: must knowfreq 72%

answer

  1. rewrite on the outgoing path
  2. form, not objective
  3. cargo not payload
  4. scorer sees the converted reply
  5. needs a plain-form control

basics

~20 s

A PyRIT prompt converter rewrites the outgoing prompt text after the seed prompt is chosen and before it is sent. It changes the surface form, not the request underneath. On its own it is just a transform: it tests whether that form gets through, never whether the model will actually comply.

solid answer

~50 s

A converter sits between the seed prompt and the prompt target. PyRIT hands it the outgoing text, it returns rewritten text, and only the rewritten text is sent. That is how one seed corpus is fanned out into many surface forms without editing the corpus. That is also its limit. The converter **carries** the request; it does not make one. If the underlying prompt asks for nothing a model would refuse, no amount of rewriting turns the run into a finding. If the model refuses the plain version, the converter only tells you whether that particular form is handled differently by whatever sits in front of the model. So a converter changes the question the run asks — *does this form get through?* — not the objective. Deciding whether the reply counts as compliance is still the scorer's job, and the scorer sees the converted exchange, not your intent.

go deeper

for a junior

Say a converter rewrites the prompt text before it is sent, and that the target only ever sees the rewritten version.

for a middle

Add that the converter does not change the objective, and that a suspiciously good result may be the transform confusing the scorer rather than the model complying.

for a senior

Frame it as a change to what the run measures — input handling versus model behaviour — and insist on an unconverted control run before any number is reported.

for a principal

Talk about which claim the organisation is allowed to make from a converter run, and how you keep input-handling results from being reported as model-alignment results.

### The object and where it sits in the send path A PyRIT prompt converter is a subclass of `PromptConverter` with one required method, `convert_async(prompt, input_type)`, which returns a `ConverterResult` carrying `output_text` and an `output_type`. That is the whole contract: a value goes in, a rewritten value comes out. You attach converters to the thing that drives the run, never to the target — older PyRIT takes `prompt_converters=[...]` on the orchestrator, newer versions take a list of `PromptConverterConfiguration` on the attack's converter config. PyRIT's `PromptNormalizer` applies them on the outbound path immediately before the target's `send_prompt_async`, so the endpoint is handed a finished string and has no idea a chain existed. ### Two values are recorded, and that distinction is the whole leaf Every turn lands in PyRIT's memory as a `PromptRequestPiece` that stores `original_value` and `converted_value` as separate columns, alongside `original_value_data_type`, `converted_value_data_type` and `converter_identifiers`. The target received `converted_value` and nothing else. Any sentence you write afterwards that says "the prompt" is ambiguous until you say which of the two you mean, and that ambiguity is where most bad converter reports begin. ### Why a converter on its own is not an attack The objective — the thing the run declares success against — is configured on the attack and the scorer, not on the converter. A converter *carries* a request; it does not *make* one. Three consequences follow directly: - **Benign seeds stay benign.** If the underlying request asks for nothing a model should refuse, no chain, however exotic, turns the run into a finding. A hit there is by construction an artefact of scoring. - **The unit under test moves.** If the plain form is refused and the converted form is not, you have measured the *input path* — tokenisation, any input classifier in front of the model, the model's ability to read that form — not the model's behaviour on the request. Those are different claims with different owners, and conflating them makes a report wrong even when its counts are right. - **The technique is cargo here.** Whatever encoding or obfuscation a converter applies, its efficacy as an evasion belongs to the guardrail-evasion topic. What this leaf owns is that PyRIT composes such a transform in series, bills for it, and stores both sides of it. ### What it costs Local converters are microseconds of string work and cost nothing. A converter configured with a `converter_target` — anything that asks a model to paraphrase, translate, restyle or expand the prompt — is one full inference call per prompt per link, spent before the system under test is touched at all. Over a 200-seed corpus that is 200 extra calls for a single such link, several minutes of extra wall clock, and quota burned on an endpoint that is usually the same provider account as your scorer and your attacker model. The target call and a model-backed scorer's call are metered separately on top. So the honest unit of a converter run's cost is **calls per attempt**, never seeds. ### Where the number misleads Four readings go wrong, in ascending order of embarrassment. First, an input-handling result published as a model-alignment result: "we jailbroke the model" when what you showed is that a filter does not normalise a surface form. Second, the scorer is downstream of the transform — if the converted prompt nudges the model to answer in the same form, a substring or refusal-detecting scorer is reading text it was never calibrated on, and it fails in both directions: unreadable output flagged as compliance, genuine compliance in the converted form missed. Third, an aggressive rewrite can drift the request off the objective entirely, and the low success rate that follows is your chain's failure, not the target's robustness. Fourth, and purely mechanical: newer PyRIT accepts converters on the response path as well as the request path, and a chain configured into the response list changes nothing the target ever saw while still looking, in the config, like a tested chain. ### What I would check before believing anything Pull a dozen `PromptRequestPiece` rows out of memory and read `original_value` against `converted_value` side by side, confirming the converted string still asks for the objective at all. Check that `converter_identifiers` lists the chain you intended, in the order you intended — the cheapest available proof that the chain actually ran, on the request path. Keep one unconverted control run over the same seeds, scored the same way, so the converter's contribution is a measured difference rather than an assumption. And price a single prompt end to end before committing to the sweep.

  • If the seed prompt is benign, can a converter chain produce a real finding?
    No. The converter only changes form. With nothing objectionable in the request there is nothing for the target to refuse, so any 'hit' is a scoring artefact.
  • Why do you keep the pre-conversion prompt at all?
    Because triage needs both: the stored exchange lets you check that the rewrite preserved the objective and lets you re-score the reply later without re-running the target.
  • Where does a converter sit relative to the scorer?
    Strictly upstream on the outgoing path. The scorer judges the reply to the converted prompt, so any distortion the transform causes in the reply lands in the scorer's input.

saying these in an interview costs you the question

  • Calling a converter an attack, or saying it 'jailbreaks' the model by itself.
  • Believing the target or the scorer sees the original pre-conversion prompt.
  • Assuming all converters are free local string operations.
  • Reporting a converted run's hit count with no unconverted baseline.

context

open as a page

A PyRIT run is configured with several prompt converters applied to each outgoing prompt. In what order does the chain apply them, and why does swapping two converters change what the target actually receives?

level: middleimportance: must knowfreq 62%

basics

~20 s

The chain runs in series over a single prompt: the first converter's output becomes the second's input, and only the final string is sent. So the converters compose, and composition is not commutative — rewriting an already-transformed string gives a different result than transforming a rewritten one.

open as a page

Adding an encoding converter to a PyRIT chain triples the success count over the same seed prompts. Before you report that as a jailbreak result, how do you establish whether the target actually complied?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Treat the jump as a scoring hypothesis first. Pull the stored exchanges, decode the replies, and read whether the content is really there. A transform that changes the reply's form also changes what a text-based PyRIT scorer sees, so refusals stop matching and garbled output can read as compliance.

open as a page

Your team's standing PyRIT suite runs every seed prompt through a matrix of prompt-converter variants. How do you decide how many variants the suite carries, and what does a per-converter success number actually measure when you publish it?

level: principalimportance: should knowfreq 33%

basics

~20 s

Size the matrix by distinct transform classes, not by count: variants multiply target and scorer calls linearly and mostly re-test the same input path. A per-converter number measures which surface form got past the input handling and the judge — not distinct model weaknesses, and not attack-surface coverage.

open as a page