skip to content

In a ticket-triage workflow, what must hold for text in a filed ticket to reach the parser reading the assistant's answer?

level: middleimportance: should knowfreq 45%

answer

  1. three hops, all fragile
  2. the model must reproduce, not paraphrase
  3. position decides, wording does not
  4. anything in between can dissolve it
  5. no feedback in an unattended workflow

basics

~20 s

Three conditions. The model must reproduce the filed span in its answer rather than paraphrase it away; the span must land where the parser's grammar gives it meaning; and nothing between the model and the parser may rewrite or re-encode the answer.

solid answer

~50 s

The construction is a relay with three fragile hops. First, carriage: the assistant summarises a customer ticket, and a summariser is free to paraphrase — the span only counts if it survives close to intact, so wording that reads as something worth quoting carries better than wording that reads as noise. Second, position: the parser lifting routing fields out of the answer assigns meaning by place and marker, so a span that appears in the wrong part of the answer is inert. Third, an unbroken path: any normalising, re-serialising or templating step between the model and the parser can strip the property without anyone intending to. The cost sits in the fact that none of this is observable from outside. The attacker never sees the parser's grammar, and in an unattended workflow the reply channel returns nothing about how the answer was read.

go deeper

for a junior

Recall that text from a filed ticket ends up inside the assistant's context and can be echoed into its answer. Know that a summariser may or may not reproduce a given fragment.

for a middle

Explain all three conditions — carriage through summarisation, position within the parser's grammar, and an unbroken path — and say which one fails most often and why.

for a senior

Demonstrate that you can state the attempt economics: no feedback, a blind guess at the grammar, variance per generation, and near-zero cost per attempt against a consumer that never tires.

for a principal

Own the framing that the boundary lives at the consumer, and be ready to defend spending on that side against the cheaper instinct to constrain the producer.

## What is actually being attempted A ticket-triage workflow takes a customer-filed ticket, has an assistant read and summarise it, and hands the assistant's answer to a parser that lifts out the values the workflow acts on: a queue, a priority, an owner. Nobody reads the answer at the time; the only human sight of it is a digest of routing decisions the next morning. The construction under discussion does not try to make the model say something forbidden. It tries to get a span from the ticket body to arrive in the assistant's answer *in a place where the parser's grammar gives it force*. The harm classifier in front of the answer is looking for abuse and violence; a span that looks like an ordinary fragment of a support ticket scores nothing. That sounds easy and it is not, because three separate things must hold at once. ## Condition 1 — carriage The model is not a pipe. Between the ticket body and the answer sits a summarisation step, and summarisers paraphrase. A span survives when the model has a reason to reproduce it rather than restate it: material that presents itself as a quotation, an identifier, a code, a verbatim line the summary would naturally carry through. Material that reads as prose the model can compress is compressed. **This is the single largest source of failure in the class**, and it is not deterministic — the same ticket, the same model, and the span arrives in one attempt and not the next. Anyone reproducing such a finding runs into this immediately. ## Condition 2 — position The parser assigns meaning by *place and marker*. A field value is a field value because of where it sits, not because of what it says. So a span can survive carriage perfectly and still do nothing, because it landed in the narrative half of the answer rather than the half the parser reads. Getting the position right requires knowing something about the grammar — and the attacker is guessing. Common shapes get guessed successfully far more often than one would like, precisely because these consumers are usually simple and conventional. ## Condition 3 — an unbroken path Even with carriage and position, anything between the model and the parser can dissolve the effect: a normaliser that collapses whitespace or strips characters, a templating step that wraps the answer, a service that re-serialises it, a truncation budget that cuts the answer short. None of these were put there for this reason, which is why the class often stops working for reasons no one records. It also means the finding can evaporate on a deployment change that nobody associated with security. ## What it costs the attacker - **No feedback.** In an unattended workflow the ticket author sees an acknowledgement, not a parse trace. Whether the answer was read as structure is invisible from outside. - **Blind on the grammar.** The consumer's field names, markers and positions are not published; attempts are guesses against an unseen target. - **Probabilistic on every attempt.** Carriage varies by generation. A construction that works one time in five is genuinely working, but that is not the same as a reliable finding. - **Cheap per attempt, though.** Filing a ticket costs nothing, and the workflow will process every one of them. That asymmetry is the honest summary: high uncertainty per attempt, near-zero cost per attempt, and a consumer that never gets tired. ## Where it stops working It stops when the answer stops being the thing the workflow reads for its values — when what the consumer acts on no longer originates as generated prose. Note that this is a statement about *where the property lives*, not a recommendation: the point for an interview is that the boundary sits at the consumer, and every attempt to stop it at the producer is fighting a probability rather than a mechanism. Naming that boundary correctly is what separates a candidate who has thought about this class from one who has read a list of attack names.

  • Why is the summarisation step the most common reason an attempt fails?
    Because a summariser rewrites by design. The span only counts if it arrives close to intact, and whether the model quotes or compresses a given fragment varies between generations. That makes carriage the noisiest link in the chain and the reason the same ticket can work once and then not again — which matters enormously when someone tries to reproduce a filed finding.
  • Does the attacker need the model to disobey its instructions for this to work?
    No, and that is the point. Nothing about the assistant's behaviour has to be abnormal — it summarised a ticket, which is its job. There is no refusal to get past and no trained boundary being pushed on. The consequence is manufactured entirely at the consumer, which is why framing this as a jailbreak leads people to look in the wrong place.
  • If the finding reproduces once in five attempts, is it a finding?
    Yes, with the rate stated. A probabilistic producer means success rate is part of the finding, not a caveat that discounts it. One in five against a workflow that processes every filed ticket is a real exposure; the honest write-up gives attempts, successes and the conditions, and resists both rounding it up to "reliable" and dismissing it as a fluke.

saying these in an interview costs you the question

  • Assumes the model reproduces input text verbatim by default
  • Thinks the span works wherever it appears in the answer
  • Calls it a jailbreak because a ticket body was the origin
  • Ignores normalising or templating steps between model and parser
  • Treats one-in-five reproduction as proof the finding is invalid

context