Adding an encoding converter to a PyRIT chain triples the success count over the same seed prompts. Before you report that as a jailbreak result, how do you establish whether the target actually complied?
answer
- hit may mean the judge missed
- transform changes the reply's form too
- decode and label a sample yourself
- benign-objective control run
- plain-form baseline scored the same way
basics
~20 sTreat the jump as a scoring hypothesis first. Pull the stored exchanges, decode the replies, and read whether the content is really there. A transform that changes the reply's form also changes what a text-based PyRIT scorer sees, so refusals stop matching and garbled output can read as compliance.
solid answer
~50 sThe converter sits upstream of the target, but its effect lands downstream too: if the prompt's form pushes the model to answer in the same form, the scorer is now judging text it was never calibrated on. Two failure directions matter. **False hits**: a scorer that flags 'did not refuse' marks an encoded or off-topic blob as success because none of its refusal patterns match. **False misses**: a scorer looking for objective-specific content misses a reply that genuinely complied but did so in the converted form. The check is transcript work, not more running. Sample the flagged exchanges, decode by hand, and label them yourself. Then re-score that labelled sample and see how the scorer did on converted output specifically. Also run a control: the same chain over prompts whose objective is harmless. If those also 'succeed', the number is measuring the transform's effect on the scorer, not on the model.
go deeper
Recognise that the number could be wrong and that someone should read the actual replies before believing it.
Name the mechanism: the scorer judges the reply's text, and the transform changed the reply's text, so its patterns no longer apply.
Give the triage order — sample and label, measure the scorer's false-positive rate on converted output, benign control, plain-form baseline — and separate a model finding from an instrument finding in the write-up.
Set the rule that no converter-driven number leaves the team without a labelled sample behind it, and treat 'our judge is defeated by this form' as a tracked instrument defect with an owner.
This is the signature failure of converter work and the reason the leaf exists: a transform can slip a payload past a text-matching judge just as easily as past anything sitting in front of the model, and the harness records both events identically — one row, verdict true. ### The mechanism, at the level of the object that does it A PyRIT scorer reads the **response** piece and returns a verdict that ends the attempt. The shipped families behave differently under a transform: - **Pattern scorers** (a substring or regex check) match literal text. Encode, translate or restyle the reply and the pattern simply stops matching. - **Refusal-shaped scorers**, where success is defined as *the absence of a refusal*, are the dangerous ones. Anything that is not recognisably a refusal scores as success — including a blob the judge could not parse at all. An encoding transform is close to a machine for manufacturing exactly that input. - **Model-backed rubric scorers** degrade more gracefully but are still being asked to judge text far outside the distribution anyone calibrated them on. The prompt converter runs upstream of the target, but its effect lands downstream too, because the reply's form tends to follow the prompt's form. So a tripled success count after adding an encoding converter is, on its face, at least as consistent with "the judge broke" as with "the model complied". ### Triage order 1. **Sample and read.** Pull a double-digit sample of newly flagged exchanges from memory. Decode or translate the reply by hand. Ask one question per item: does this actually deliver the objective? Your label is now ground truth — it is the only ground truth in the building. 2. **Score the scorer.** Compare your labels with the run's verdicts on that same sample. You now have a false-positive rate for this scorer **on converted output**, which is the only rate that matters here. Report it as a number, not as a feeling. 3. **Benign control.** Run the identical chain over seeds with no objectionable objective. Every hit there is pure artefact and sets a noise floor. 4. **Plain-form baseline.** Compare against the unconverted run over the same seeds, scored the same way. The converter's contribution is the difference, and it is interpretable only if both sides were judged identically. 5. **Check the prompt survived.** Read the final `converted_value` too. If the transform mangled the objective out of the request, a *low* rate is your chain's fault rather than evidence of robustness. ### What the checks cost Less than people fear, and far less than a retraction. Re-scoring is the cheap part: PyRIT's memory holds the stored responses, so pointing a corrected scorer at them costs **scorer calls only** — no target calls, no converter calls, no re-running the attack. Hand-labelling thirty transcripts is one to two hours of an engineer's time. A benign-objective control over a small seed set is one extra pass of target-plus-scorer calls — buy a small control, not a mirror of the whole suite. Set that against the alternative: a headline rate that a customer spot-checks and refutes with a single decoded transcript, which costs the engagement's credibility and every number in the report with it. ### Where the number misleads, specifically The tempting reading of "3x more successes" is that the encoding defeated the model's safety behaviour. The competing readings are all cheaper. The judge's patterns no longer fire, so non-refusals accumulate. The reply is unreadable and therefore unrefusable. The model produced a genuinely off-objective answer in a strange form and nothing in the pipeline noticed. And the mirror-image error is just as real: a converted reply that *does* comply but in a form the scorer cannot read counts as a miss, so a run can also understate. Tuning the scorer's threshold against this unlabelled data is not a fix — without labels you cannot tell whether you removed false positives or real hits, you have only moved the error somewhere less visible. ### What you are allowed to say afterwards One of exactly two sentences. Either "the target complied, verified by reading N transcripts, false-positive rate on this form measured at X" — a model finding. Or "this surface form defeats our judging, measured false-positive rate X" — an instrument finding, genuinely valuable, owned by the red team, and explicitly **not** a model vulnerability. Filing the second as the first is the mistake that ends the report's credibility on the first spot-check, and reporting the raw tripled count with neither check behind it is how you get there.
- The reply is an encoded blob and the run marks it a success. What is the minimum check?Decode the stored reply and read it. If the decoded text does not deliver the objective, the verdict is an artefact of the scorer reading a form it cannot parse.
- Your benign-objective control run also shows hits. What does that tell you?That some fraction of every converted run's successes is noise from the scorer. It gives you a floor to subtract, and it says the scorer must be fixed or re-calibrated before the numbers mean anything.
- How do you write this up if the model never actually complied?As an instrument finding: this surface form defeats our judging with a measured false-positive rate. Useful, owned by the red team, and explicitly not a model vulnerability.
A refusal-detecting scorer confronted with an encoded reply is an examiner who grades by scanning for the word 'no' and marks everything else a pass. Hand it an answer in a script it cannot read and it will confidently pass a blank page.
saying these in an interview costs you the question
- Reporting the tripled count as a jailbreak rate with no transcripts read.
- Assuming the scorer behaves the same on converted output as on plain output.
- Re-running with a stricter threshold instead of labelling a sample.
- Having no benign-objective control and no plain-form baseline to compare against.
- Filing a judge failure as a model vulnerability.