skip to content

The owner adds a preamble line telling a review bot never to discuss its tooling - what did that buy them?

level: seniorimportance: should knowfreq 40%

answer

  1. grant that the line always holds
  2. the surface did not move
  3. the work product describes the work
  4. citations, refusals, failure text
  5. cost changed, exposure did not close

basics

~20 s

A preamble line buys the cheapest read - a stranger no longer gets the surface described on request. The surface itself is unchanged, and the bot's ordinary replies still show which sources it reads and what it cannot do.

solid answer

~50 s

Two separate things are worth saying, and candidates usually say only one. The first is the familiar one: a preamble line is a preference the model was trained to weigh, not a boundary anything enforces, so it holds unevenly. The second matters more here. Even where it holds perfectly, it only removes the direct answer to a direct question. The operation surface is still visible in the bot's ordinary output: it declines a task by explaining what it is not able to do, it says which material it read for this review, and a failed step surfaces as an explanation in the same public thread. So the attacker loses a one-comment read and gains a slower inferential one at lower fidelity. That is a real change in cost and worth acknowledging - it is simply not the same as closing the finding.

go deeper

for a junior

Know that a line in a preamble tells the model what to prefer saying; it does not change what the system is wired to do. The capabilities are still there whether or not they are discussed.

for a middle

Explain the residual concretely: source attribution, refusal wording, failure messages and visible effects all describe the surface without any discussion of tooling. Then say what the change did buy - the one-interaction read.

for a senior

Demonstrate the reporting discipline: state precisely which route closed and which remains, price the cost change honestly, and back the residual with the bot's ordinary public output rather than with a phrasing that may stop working.

for a principal

The tradeoff to own is what a text-level change may be claimed to have achieved when the finding is reported as resolved, and how much of that claim survives a model or configuration change nobody re-tests.

## The situation A finding says a public review bot will describe its own operations to whoever asks in a pull-request comment. The owner adds a line to the preamble telling it not to discuss its tooling, re-runs the original request, sees a polite non-answer, and asks for the finding to be closed. The question is what that line actually changed. ## The part everybody says A line in a preamble is an instruction the model was trained to prefer, not a rule anything enforces. It holds most of the time and not all of the time, and the same phrasing behaves differently as the model or the surrounding context changes. Say it in one sentence and move on - it is assumed knowledge in this domain and it is not the interesting half of the answer. ## The part that decides the question Grant the strongest possible version of the fix: assume the line is honoured every single time. What is still true? **The surface did not move.** The bot still holds the same operations and still exercises them on every review. Nothing about what it can reach changed; one sentence about what it may say changed. **The bot's ordinary work product describes the surface anyway.** Consider what a public review reply contains: - *Attribution of what it read.* A review that refers to the failing job's output, the linked issue, and the diff has told the reader which sources this installation reaches. - *Refusal-by-explanation.* Asked to do something outside its wiring, an assistant naturally explains it cannot do that here - which is a statement about the surface, arrived at without discussing tooling in the abstract. - *Failure text.* When a step does not succeed, the reply commonly says something went wrong retrieving or posting something, which names the operation by its effect. - *Effects that are simply visible.* If it labels, links or comments, the reader watches it happen. None of that is the model discussing its tools. It is the model doing its job in public, and doing the job is itself an account of what the job includes. **The attacker's cost changed, and that is the honest finding.** Before, one comment produced a list. Now, the same map has to be assembled from behaviour across several interactions, at lower fidelity and with more inference. That is a genuine improvement in cost and should be reported as one. It is not an elimination. ## Reporting this without overclaiming or underclaiming Two failure modes, both common: | Failure mode | What it sounds like | Why it is wrong | |---|---|---| | Underclaiming | The finding is closed, the bot no longer answers | Tests one phrasing against one behaviour and reads a preference as a boundary | | Overclaiming | The fix is worthless | Ignores that the cheapest, most repeatable read really did go away | The accurate statement is narrow and defensible: *the direct disclosure route no longer returns the surface in one interaction; the surface remains inferable from the bot's normal public output, at higher cost and lower fidelity.* That sentence survives review because every clause is something you can point at. ## Verifying rather than assuming If you assert the residual, show it. The demonstration is not a cleverer request - it is reading what the bot already posts, unprompted, on ordinary pull requests: which sources it cites, how it phrases what it cannot do, what its failures say. That evidence is durable, it does not depend on a phrasing that may stop working next week, and it is what makes the difference between an opinion about preambles and a finding about this deployment. ## Where the residual genuinely shrinks Be fair about the other direction too. If the bot's replies stop attributing sources, stop explaining failures in the thread, and produce no visible effects, the behavioural route narrows sharply - because there is less work product to read. The residual exists in proportion to how much the bot says about what it did, which is a property of the deployment rather than of the preamble line. ## The interview signal An interviewer is checking whether a candidate can hold two things at once: that the change bought something real, and that it did not touch the mechanism the finding was about. Candidates who only debunk the preamble sound reflexive; candidates who accept the closure sound untested. The good answer prices the change and names the residual.

  • How would you demonstrate the residual without relying on a cleverer request?
    By reading what the bot already posts on ordinary pull requests: which sources it cites for its review, how it phrases things it cannot do, and what its failure messages name. That evidence does not depend on a phrasing that may stop working, and it is specific to this deployment rather than to a general claim about preambles.
  • Is there a version of this where you would agree the residual is small?
    Yes - where the bot's replies carry no source attribution, explain no failures in the thread, and produce no visible effects, there is little work product to read and the behavioural route mostly closes. That is a property of how much the deployment narrates, not of the preamble line, and it is worth saying so explicitly.
  • Why is testing the original request once a weak basis for closing the finding?
    It tests one phrasing on one occasion against a preference that holds unevenly, and it only covers the direct route. A single non-answer establishes that this request returned nothing this time, which is a much smaller claim than the surface no longer being obtainable.

Telling a courier not to talk about their route does not change the route. Anyone watching which doors they knock on learns most of it anyway, just more slowly.

saying these in an interview costs you the question

  • Treats a preamble line as an enforced boundary
  • Says the fix is worthless without pricing what it did buy
  • Closes the finding on one non-answer to one phrasing
  • Forgets the operations still exist and still run
  • Overlooks the bot's own replies as a description of its surface

context