Why does an output screen that blocks secrets and raw tool JSON let a review bot describe its own operations in prose?
answer
- what does the screen compare against?
- shape, not meaning
- a paraphrase has no braces
- an inventory written in sentences
- passing is a score, not a verdict
basics
~20 sAn output screen matches on shape - credential patterns, key formats, blocks of JSON. A plain-English sentence about which operations a bot has matches none of those, so capability talk leaves as ordinary help text while carrying an inventory.
solid answer
~50 sAn output screen applied after generation is a shape check on the reply: things that look like credentials, key formats, or a block of JSON that looks like a tool definition. A sentence such as 'I can read the diff and the CI output, and I leave one review comment' has none of that shape, so it goes out looking like ordinary help text - which is what it is, and also an inventory of what this installation can do. Two directions matter. The screen passing proves the text scored below whatever the screen matches on, not that the text is harmless. And the reply proves what the model said about itself, not what the deployment actually holds - it can name an operation that was never wired up. An attacker who asks a public pull-request bot what it can do gets both, for one comment.
go deeper
Be ready to say what an output screen compares against - shapes and patterns in the emitted text - and why a sentence describing capabilities has none of those shapes. Then say plainly that passing a screen is not the same as being harmless.
Explain the form mismatch precisely: structural detectors and key-format patterns against a paraphrase that carries neither. Note the boundary too - push the same content back into a JSON or fenced shape and the screen has something to match.
Show you can read what the observation supports. A pass means no match; a named operation means the model said so. Do not write a report that treats the bot's self-description as an inventory of record without confirming anything.
The angle to own is what a shape-anchored screen can be claimed to cover at all. Anything whose risk lives in meaning rather than form sits outside its stated scope, and that gap should be stated openly rather than discovered in a report.
## The setting A pull-request review bot reads a diff, the issue thread and the CI output for a change, and posts its reply publicly in the same thread. Anyone who can open a pull request - including a first-time external contributor - can put text in front of it and read what comes back. On its replies sits an output screen: a check applied after generation, looking for secrets, credential-shaped strings, and raw tool payloads such as a JSON object that looks like a capability definition being pasted out. The construction is unglamorous. The contributor asks, in an ordinary help-seeking register, what the bot is able to do here and roughly what it needs to be told to do it. The bot answers in prose. The screen passes it. ## What the screen actually compares against This is the whole of the mechanism, and it is why the question is a first-encounter one rather than an advanced one. A screen of this kind is a matcher over the *form* of the emitted text: regular expressions for key formats, entropy heuristics, structural detectors for JSON or code fences, sometimes a classifier trained on examples of leaked secrets. Every one of those is anchored to a shape. A paraphrase has a different shape from the artefact it paraphrases. `{"name": ..., "parameters": {...}}` is structurally distinctive; the sentence *I can fetch the build log for a run and comment once on the pull request* is structurally indistinguishable from every other helpful sentence the bot emits all day. There is nothing for the matcher to catch, and a screen built to stop leaked credentials has no notion that an inventory of operations is a thing worth stopping. Note what is **not** the explanation: the screen did not understand the text and decide it was benign. It compared and found nothing. Passing a screen is a statement about a score or a match, never a statement about consequence. ## What the attacker got Not secrets. A map. Specifically: - Which operations exist **in this deployment**, as opposed to what the product is documented to support generally. Installations enable different subsets. - Roughly what the arguments are called, and roughly what they accept. - Which of them the bot describes as read-only and which produce an effect somebody sees. That is reconnaissance, and its value is that the next thing the attacker writes is aimed rather than speculative. Before, they would have to guess at a surface; after, they are working from a description the system gave them, in public, for free. ## What the pass does and does not prove Three directions are easy to get backwards: | Observation | What it proves | What it does not prove | |---|---|---| | The screen passed the reply | The text did not match what the screen matches on | That the content is harmless | | The bot named four operations | The model produced those four names | That the installation holds exactly those four | | The bot said an operation is read-only | The model described it that way | That the operation has no effect | The second row is the one candidates miss. A language model asked about itself produces a plausible account, and plausible is not the same as accurate: it can omit an operation it holds, and it can name one it does not. The paraphrase is a lead, not a record. ## Where it stops working The route depends on the reply staying prose. Push the bot toward emitting the actual definition - a pasted JSON object, a fenced block - and the reply reacquires the shape the screen was built for, and the screen fires. That is the honest boundary: the attacker is not defeating the screen, they are staying outside the class of thing it was built to recognise. It is also why a candidate who says *the screen was bypassed* is describing it wrongly. Nothing was bypassed. The request produced an artefact the screen has no opinion about. A second boundary is fidelity: everything gained is a description, so anything acted on later rests on the model having described its own surface correctly. ## How to say it in an interview Name the screen's matching basis, name the shape mismatch, and then say what the disclosure is worth - a map, not a secret. Candidates who stop at *it does not match the regex* have the mechanism; candidates who add *and the pass tells you nothing about harm* have the reasoning an interviewer is scoring.
- The bot's reply names an operation that does not exist in this installation. What does that tell you?That the reply is the model's account of itself, not a record of the deployment. A model asked what it can do produces a plausible description, and plausible includes invented. Treat the list as leads to be confirmed by what the system actually does with a later request, and never report it as an inventory of record.
- Would the same screen behave differently if the bot pasted an actual capability definition instead?Very likely yes - a JSON block or a code fence is exactly the structural shape such a screen exists to catch, so it would fire. That contrast is the point: the screen is anchored to form. The prose route exists because a description of the same information carries none of the form the screen recognises.
- Does it matter that this arrived on a public pull-request comment rather than a logged-in session?It matters a great deal for cost. The requester needed no account with the project, no review and no approval, and the reply is posted publicly in the same thread where anyone can read it. The disclosure is therefore repeatable by strangers and durable in a public record.
A metal detector at a door finds metal. Somebody walking through it and describing the building's floor plan aloud sets nothing off - not because the detector approved, but because a floor plan is not metal.
saying these in an interview costs you the question
- Says the screen read the reply and judged it harmless
- Treats anything the screen passes as safe by definition
- Calls it a bypass when nothing matched in the first place
- Assumes the operations the bot names are certainly the ones it holds
- Confuses the model declining with the screen blocking