skip to content

A wiki owner says every diff is reviewed and the assistant is told to ignore page instructions. What do those two buy?

level: principalimportance: should knowfreq 38%

answer

  1. price, not coverage
  2. narrowest true sentence each claim supports
  3. reviewed artefact is not the consumed one
  4. the strong version contains no instruction
  5. sign a bound and name the residual

basics

~10 s

The review buys attacker cost, not coverage: the whole-page reading must stay unremarkable, but the reviewed and consumed artefacts differ. The prompt line buys less, because the construction need not read as an instruction.

solid answer

~50 s

Take them separately and state the bound on each. The diff review buys a real cost: the whole-page reading has to survive an attentive reader, which prices out the crude plant. It buys nothing about the retrieved unit, because the reviewed artefact and the consumed artefact are different objects — the honest claim is "no edit looked wrong as a page", not "no retrieved unit reads as guidance". The prompt line buys less than it appears to: precedence over retrieved content is a trained preference, and the stronger version of this construction contains no imperative at all — a fragment that reads as the current answer has nothing to ignore. What a lead can sign is a bounded claim plus a named residual: wiki-sourced answers inherit the wiki's trust model, which is "anyone edits, one person skims".

go deeper

for a junior

Recall the core fact you would need before joining this conversation: an approval covers the page a human rendered, and the model answers from a fragment produced afterwards.

for a middle

Be able to explain why a prompt line aimed at instructions misses an assertive fragment, and why review and retrieval consume different artefacts.

for a senior

Show that you can price a control instead of scoring it pass or fail, and report an intermittent result without inflating it into a rate.

for a principal

Own the wording of the assurance claim and the placement of the residual: state the narrowest true sentence, name who holds what is left, and leave the remedy to the owner who funds it.

## Why this conversation goes wrong Both statements are offered as if they were coverage, and both are actually prices. The job in the room is to convert each claim into the narrowest true sentence it supports, name what is left over, and say who has to hold it. Doing that badly in either direction loses the room: telling an owner their review is worthless is false and gets you ignored; letting "we review every diff" stand as an assurance statement puts a claim into a security document that the pipeline does not support. ## Claim one: every diff is reviewed What it buys, honestly: - **A floor on plausibility.** The paragraph must read as ordinary prose in place, framed by its heading and its neighbours. That eliminates the blatant plant and forces an attacker into a construction that is weaker in the fragment, because force on one reading costs plausibility on the other. - **A record and an owner.** Somebody looked, and there is a timestamp. What it cannot buy: - **Any statement about the unit the model receives.** The reviewer renders a diff hunk inside a whole page; ingestion emits chunks; retrieval returns one. Approval is evidence about the first artefact only. "We reviewed the document" is the wrong answer here, and the correcting fact is that the model never sees a document. - **Scaling.** As pages and edits grow, the number of independently retrievable units grows with them, and every one of those units is an artefact no review workflow renders. More reviewers read the same set harder; they do not converge on the consumed set. - **Trust in its own attestation.** The approval record becomes a laundering surface: when an answer is later questioned, the named approver is what people cite, and it points at an artefact that was never in the context. ## Claim two: the assistant is told to ignore instructions in pages Assume the well-known part rather than lecturing it: instruction precedence between a system message and retrieved content is a trained preference, not an enforced boundary. The more useful point for this class is narrower and often missed — **the construction does not have to be an instruction.** The strong version of a dual-reading paragraph is an assertion: read alone, the fragment is simply the current answer to the question that retrieved it. There is no imperative for a policy line to override. Aiming a prompt line at imperatives leaves the assertive family untouched, and the payoff — a false answer somebody acts on, cited to a page with an approver's name on it — is unaffected. ## What you can honestly sign A statement of this shape survives review: - Every page edit is read by a person as a rendered page, and no edit in the period was found objectionable in that form. - Content the assistant answers from is not the same artefact: it is a fragment produced after the review, and no human-facing view of that fragment exists in the workflow. - Therefore the corpus carries the wiki's trust model into every answer: any employee may write text that reaches the model, subject to one skim of a different artefact. That is a bound, an explanation, and a residual. It is not a plan, and it is deliberately not one — choosing what to do about the residual is an owner's decision with a budget attached, and the numbers that decide it (what wiki-sourced answers are used for, who acts on them, what a wrong answer costs downstream) are not the red-teamer's to assume. ## The reporting honesty that comes with it This class reproduces badly. Boundaries drift when pages are edited above the paragraph and re-indexed, so a construction can work once and never again without anyone touching it. A single reproduction demonstrates that one construction worked once against one index state and one query phrasing — it is not a success rate, and reporting it as one invites the owner to dismiss the whole class when the second attempt fails. Report the class, the conditions that produced the observation, and the fact that instability here is a property of the pipeline rather than evidence of weakness in the finding. ## Ownership The most common failure is to file this against the reviewer. The reviewer performed the task correctly on the artefact they were given; nothing in their workflow renders the consumed unit. Filing it as "reviewer missed it" produces a demand for more careful reading of the same object, which changes nothing and burns the relationship. The finding belongs with whoever owns the boundary between an editable corpus and an assistant that answers from it, and it should be written so that the owner's options are theirs to weigh. ## In an interview Separate the two claims, price each one, name the residual, and decline to pretend a prompt line closes a class. The senior answer explains the mechanism; the principal answer states the claim somebody can sign and says who holds what is left.

  • The owner asks for a number. What can you honestly give?
    Not a success rate. A single reproduction shows one construction working once against one index state, and boundary drift means a second attempt can fail for reasons unrelated to the construction. What generalises is the class and the structural fact behind it: the reviewed artefact and the consumed artefact differ, which holds for every page in the corpus regardless of how any individual attempt scored.
  • Who should the finding be filed against?
    Not the reviewer. They executed correctly on the artefact their workflow renders, and no view of the consumed fragment exists for them to have missed. It belongs with whoever owns the boundary between an editable corpus and an assistant that answers from it, because that is where the residual has to be accepted or funded.
  • The owner proposes tightening the review instead. What do you say?
    That it raises the attacker's cost and is worth having, but it operates on the page, and the fragment's second reading is created after the review by a stage the workflow does not render. More attention to the same artefact moves the price, not the class. Say that plainly, then leave the choice of what else to do where it belongs.

saying these in an interview costs you the question

  • Review coverage is 100 percent, so the risk is closed
  • The system prompt handles injection, so the corpus is fine
  • It reproduced once, so we can quote a success rate
  • File it against the reviewer who approved the edit
  • Assign more reviewers to the same diff queue

context