skip to content

A report says any forum user can outrank the official docs in a support assistant's index — is that a finding?

level: seniorimportance: should knowfreq 42%

answer

  1. who can write into that index?
  2. one query is one data point
  3. steerable, or merely relevant?
  4. could they nominate the question first?
  5. severity lives downstream, not in the score

basics

~20 s

It depends on steerability, not on the score. Ranking above the docs is ordinary retrieval behaviour; a finding needs evidence that an untrusted author can choose which question their passage wins, and a consequence downstream when it does.

solid answer

~50 s

Start by separating what the report shows from what it claims. One passage ranking above one docs chunk for one phrasing shows that the passage matched that query better — the retriever working as designed on a corpus the vendor deliberately chose to index. Three things turn it into a security finding. **Authorship:** the reporter wrote the passage, so an untrusted party can put text into the grounding set. **Steerability:** they picked the question in advance and hit it, and can do it again for another question. **Consequence:** something downstream acts on the grounded text — an answer a developer will paste, a step they will run, a decision made without checking the official page. If all three hold, this is not a relevance bug; the severity comes from the consequence, never from the similarity number. If only the ranking is demonstrated, it is a retrieval-quality report against the ingest decision.

go deeper

for a junior

Know that a passage ranking above official documentation is ordinary retrieval behaviour on its own, and that it is not evidence anything was broken into.

for a middle

Explain what would have to be added to the observation to make it security-relevant: an untrusted author, a target question chosen in advance, and something downstream that acts on the grounded text.

for a senior

Demonstrate the triage itself — bound the claim to what was shown, ask for a nominated second question, report the reproduction rate honestly, and route it to the owner of the ingest decision when nothing more is demonstrated.

for a principal

Be ready to defend a classification when both a plausible-sounding escalation and a quick close are on the table, and to say plainly which part of the system the report is actually about.

## The triage seat Somebody else filed this. You are deciding what it is. The assistant answers developer questions for a cloud vendor and its index holds two sources: the vendor's own documentation, and a public community forum. The report says a forum answer now outranks the docs for a support question, with a screenshot of the assistant using it. The temptation is to resolve it in one move — either 'that is just how retrieval works, close it' or 'external content is in the context, that is critical'. Both are wrong for the same reason: they answer without asking what the report actually demonstrated. ## What the report has shown Be exact about the claim you are entitled to. The evidence is one query, one phrasing, one snapshot of a corpus that changes daily. It establishes that a passage scored inside the retrieval budget for that query and outscored a documentation chunk. It does **not** establish that the store was tampered with, that the docs were removed, that the retriever now favours the forum, or that any other question behaves the same way. Ranking is a distance computation over text; it has no notion of who wrote a passage, so an outsider outranking the vendor is not by itself evidence of anything having gone wrong. That is also why the score in the report is nearly useless as a severity input. A high similarity number means the passage matched the query. It says nothing about impact. ## What turns it into a finding Three questions, in order. **1. Who wrote the passage, and can anyone?** If the reporter authored it themselves through the ordinary public route — an account, a post, wait for the crawl — then an untrusted party can place text into the grounding set. That is the property with security content. If the passage is a genuine community answer that happens to be good, there is no attacker in the picture at all and this is a relevance report. **2. Could they choose the question?** This is the one that decides most triage calls. A passage that happens to win a query is retrieval noise. A passage written *for* a nominated question, which then wins that question, is a steerable capability — and steerability is what makes it repeatable against a question of the attacker's choosing rather than a lottery. Ask the reporter to nominate a second question in advance and hit it. If they can, you have a method; if they cannot, you have an anecdote, and the honest write-up says so rather than inflating it. **3. What happens downstream?** Displacement is control of the grounding set, and its worth depends entirely on what the assistant does with grounded text. A read-only chat answer a developer sanity-checks is one thing; an answer containing a command a developer will run, or one feeding a step that acts without a human reading it, is another. Severity lives here, not in the ranking. ## Reproduction, honestly reported Retrieval over a live corpus is not stable: the index grows, budgets and thresholds get tuned, new documentation lands. A result that reproduced on Tuesday and not on Thursday is normal, and it is neither proof nor disproof. Report the rate and the conditions rather than rounding it to 'works' or 'does not work'. A method that wins a nominated question three times in five is a stronger result than a single lucky top-1, and saying so keeps your credibility when the next report is the reverse. ## Where it genuinely is ordinary behaviour Sometimes the correct outcome really is 'this is the system as designed'. The vendor chose to index a surface anyone can write to, in exchange for coverage the documentation does not have. That decision fused a relevance choice with a trust boundary, and every consequence of it — including outsiders competing on equal terms for the top of a result set — follows from it. When the report demonstrates nothing beyond that, the useful triage output is not a fix but a routing: this is a property of what the index admits, and the owner is whoever owns the ingest decision. Filing it against the retriever wastes everyone's time, because the retriever is not misbehaving. ## How to write the verdict A good triage note on this report has four lines: what was demonstrated, whether an untrusted author could place the text, whether the target question was chosen in advance, and what the assistant does with a grounded passage. Those four determine both the classification and the owner. What must never appear is a severity derived from the similarity score, or a conclusion that the store was compromised — the retriever did exactly what it says on the label, and that is the uncomfortable part.

  • The reporter can only reproduce it on their own phrasing. Does that sink the report?
    No, it bounds it. The useful next step is to ask them to nominate a target question in advance and hit it — that separates a steerable method from a lucky match. Report the reproduction rate and the conditions rather than rounding to 'works'; over a live, changing corpus an intermittent result is expected and is still evidence.
  • What raises the severity here without changing the ranking at all?
    What the assistant does with the grounded passage. The same displacement is worth little when a developer reads an answer and checks it, and a great deal when the answer contains a step they will run, or feeds a workflow that acts without anyone reading the text. Severity is a downstream property, not a retrieval property.
  • Whose problem is it when the report demonstrates nothing beyond ordinary ranking?
    The owner of the ingest decision. Indexing a surface that outsiders can write to was a deliberate trade of coverage against trust, and outsiders competing for the top of a result set is a direct consequence of it. Routing it there is a more useful triage outcome than filing a defect against a retriever that is behaving exactly as specified.

saying these in an interview costs you the question

  • Closes it as a relevance bug without asking who wrote the passage
  • Calls one top-1 result on one phrasing a reliable finding
  • Derives severity from the similarity score itself
  • Concludes the vector store must have been compromised
  • Assumes indexing a public surface was never a deliberate choice

context