skip to content

A PyRIT run is configured with several prompt converters applied to each outgoing prompt. In what order does the chain apply them, and why does swapping two converters change what the target actually receives?

level: middleimportance: must knowfreq 62%

answer

  1. series, output feeds input
  2. composition is not commutative
  3. unreadable transforms go last
  4. model-backed links can eat the transform
  5. print the final sent string

basics

~20 s

The chain runs in series over a single prompt: the first converter's output becomes the second's input, and only the final string is sent. So the converters compose, and composition is not commutative — rewriting an already-transformed string gives a different result than transforming a rewritten one.

solid answer

~50 s

A converter chain is a pipeline, not a fan-out. Each link takes the previous link's output, so the target sees one string: the accumulated result of every rewrite in order. That makes ordering a real design decision rather than cosmetic. Put a structural transform first and a model-backed rewrite second, and the rewriting model is now paraphrasing an already-mangled string — it may 'fix' it, drop it, or refuse to touch it, and you have silently lost the transform. Put the model-backed rewrite first and the structural transform second, and the structural step is applied to text you can still read and reason about. The practical rule: anything that makes the prompt unreadable to another model belongs last, and anything that needs comprehensible input belongs earlier. When a chain produces surprising results, print the final converted string from the stored exchange before blaming the target — most 'the attack stopped working' reports are one link eating another's output.

go deeper

for a junior

Know that the converters run one after another and that the target sees only the final string.

for a middle

Explain composition explicitly, and give the concrete failure: a model-backed link asked to rewrite already-mangled text can normalise the earlier transform away.

for a senior

Add attribution — a long chain yields one number and no idea which link mattered — and the habit of reading the final sent string out of the stored exchange before diagnosing the target.

for a principal

Argue for a small, ordered, documented basis of chains the team actually understands, over an ad-hoc pile whose composed behaviour nobody can explain in a report.

### The mechanic, stated plainly PyRIT's `PromptNormalizer` walks the configured converter list in order and feeds each link the previous link's `ConverterResult.output_text`. Two links A then B produce `B(A(x))`. Only the final string reaches the target's `send_prompt_async`, and only it is stored as the piece's `converted_value`. A chain is a **pipeline over one prompt**, not a fan-out that sends several variants — a fan-out is a different configuration (several separate chains, or a variant matrix), and confusing the two is the most common misreading of a converter config. ### Order changes the artefact `B(A(x))` and `A(B(x))` are different strings. So any claim of the form "this chain tests X" is a claim about the composed output, and it can only be checked against the composed output. Nothing about a link in isolation survives into the run's numbers. ### Order changes what survives — the concrete failure Links are not all pure functions. A link configured with a `converter_target` hands the current string to a model and asks for a rewrite. Give that model an already-mangled string and it can do any of four unhelpful things: normalise the earlier transform away while tidying the text up, refuse the rewrite, truncate it, or return something semantically adrift. You then pay for the call, the earlier link's effect is gone, and the run shows a clean drop in hits that reads exactly like the target getting stronger. The rule that falls out: anything that renders the text unreadable to another model belongs at the **end** of the chain; anything that needs comprehensible input belongs before it. ### Order can also break the chain outright — and that is the good case `ConverterResult` carries an `output_type`, and each converter declares which input types it accepts. Put a link that emits a non-text type (an image or audio path) ahead of a text-only link and the chain fails loudly. Loud is the good failure: you see it immediately. The dangerous ordering failure is the semantic one above, because it produces a plausible number instead of an error. ### What it costs Each model-backed link bills one call per prompt regardless of its position, but position changes the *yield* of those calls: a rewrite link placed after a mangling transform has a much higher refusal and retry rate, so you pay for calls that return nothing usable. Attribution costs far more. A five-link chain gives you one number and no idea which link carried it; getting that answer means ablation — five runs with one link removed each, or five singleton runs — which multiplies target and scorer calls by roughly the number of links, on top of the converter calls. Over a 200-seed corpus that is the difference between about 200 target calls and about 1,000. That is the honest price of the sentence "this transform is what worked", and it is worth paying only for results you intend to publish. ### Where the number misleads | Observation | The tempting reading | The likelier one | |---|---|---| | Hits drop after inserting a link | the target got more robust | a later link normalised or destroyed the earlier transform | | Chain C reports 40% | transform C works 40% of the time | 40% belongs to the composition; no link owns it, and it is not additive with other chains | | Chain ending in a heavy transform looks strongest | that transform is the most effective | the transform also changed the reply's form and defeated the scorer | The last row is the one to say out loud: order affects the judge as much as it affects the target, because the final link shapes the reply that the scorer reads. ### What I would check Reconstruct the sequence from memory rather than re-running the target: ```text original_value -> [link 1 output] -> [link 2 output] -> converted_value (sent) ``` Read all four. `converter_identifiers` on the stored piece tells you which links ran and in what order — confirm it matches the config you believe you shipped, because a chain edited in one code path and read in another is a real and silent failure. Then ask two questions of the final string: does it still ask for the objective, and is it in a form the scorer has ever been calibrated on? A "no" to the first means stop and fix the chain; a "no" to the second means every success number from the run is a hypothesis. Finally, keep the chains you publish from short and documented: a three-link chain whose composed behaviour you can explain in one sentence is worth more in a report than a nine-link chain nobody can account for.

  • How do you find out which link in a five-converter chain is carrying the result?
    Run each converter as a chain of one against the same seeds and compare, or ablate one link at a time. A single composed run gives no attribution.
  • A chain that used to produce hits now produces none after a converter was inserted. First check?
    Read the final converted prompt from the stored exchange. Most often an added link normalised or destroyed the earlier transform, or drifted the prompt off the objective.

saying these in an interview costs you the question

  • Thinking each converter gets the original prompt and all outputs are sent.
  • Treating chain order as cosmetic.
  • Never inspecting the final converted string before blaming the target.
  • Reporting a five-link chain's result as evidence about one specific transform.

context