A research agent's brief carries a claim planted on a page it fetched — which stage owns the finding?
answer
- read the per-stage record first
- every stage met its own specification
- no stage ever made a trust call
- the citation proves fetch, not vetting
- separate survival from obedience in the report
basics
~20 sNo stage does. The scraper, sanitiser and converter each met their own specification, so this is not a stage defect: page text chosen at run time entered the context with the same standing as the analyst's request.
solid answer
~50 sStart by reading the per-stage record rather than arguing ownership. It will show the sanitiser removing exactly the markup it exists to remove, the boilerplate remover keeping the main content region, and the converter emitting text byte-identical to what it received. Every stage did what it was written to do, so filing against any of them asks a text-cleaning component to do semantic work it never claimed. Filing nothing is worse: an analyst acted on a stranger's claim delivered in the agent's own voice with a citation that proves retrieval, not vetting. The honest write-up puts the defect in the trust assignment — prose from a page chosen at run time reaches the model carrying no marker separating it from the operator's instruction — and attaches the stage record as evidence that no stage is at fault.
code
text · 9 linesfetch out: text/html, 38 KB
dom-parse out: 741 element nodes
boilerplate out: 88 nodes kept (nav, footer, aside dropped)
sanitize out: 88 nodes; removed 3 script elements, 11 event-handler attrs;
text node contents unchanged
html-to-text out: 4.1 KB plain text; [directive span elided] byte-identical to
the text node it came from
chunk out: 6 passages; the span lands whole inside passage 4
...go deeper
Know that a citation in an agent's output means the agent read that page and nothing more, and that a text-cleaning stage passing cleanly is not evidence that the text is trustworthy.
Be able to walk the per-stage record and say, for each stage, what its specification was and why it met it. That walk is what rules out the three obvious owners before any ownership argument starts.
Show the triage: lead with the record, rule out the stage-level owners on their contracts, and state the finding at the level of what standing fetched prose acquires. Separate the deterministic survival claim from the probabilistic obedience claim in writing.
Own the bug-versus-design-limit framing explicitly and be willing to say both halves aloud, including what a proposed change would leave untouched. A report whose credibility survives the first failed re-run is worth more than a confident one that does not.
## The situation A market-research assistant follows links off a search step, reads pages chosen at run time, and returns a cited brief. An analyst acts on a claim in that brief. The claim came from a page somebody published for the agent to find, and it arrived wearing the agent's phrasing and a citation to a real URL — indistinguishable, on the page, from what the agent read for itself. Now somebody has to file it, and the first three suggestions are the scraper, the converter, and nothing at all. ## Read the record before assigning blame The artefact that settles this is a per-stage record of one fetch: what each stage received and what it emitted. In this family it reads the same way every time. The parser re-represented the bytes. The boilerplate remover kept the main content region and dropped the furniture. The sanitiser removed the executable markup it is defined to remove and left text node contents untouched. The converter emitted prose byte-identical to those text nodes. The chunker split on its policy and the span landed whole inside one passage. That record is not evidence of a malfunction. It is evidence of the *opposite*, and that is precisely why it belongs in the report: it forecloses the three cheap resolutions before anyone spends a week on them. ## Why each candidate owner is the wrong one **The sanitiser.** Its contract is markup removal for a browser consumer. Asking it to notice a directive sentence is asking for a different component with a different threat model. A change here would also be unfalsifiable as a fix — the same span written slightly differently survives, because prose is what the stage is required to keep. **The converter or the scraper.** Both are defined as lossy transforms in the direction that preserves readable text. If they stopped preserving it, the agent would stop working. There is no version of `fix the converter` that is not `stop reading pages`. **Nothing at all.** This is the resolution that gets chosen most often and it is the one that costs the most later, because the fact that survives is real: an analyst made a decision on an assertion authored by whoever wanted to be found. Closing it as no-fault turns a design property into folklore that gets rediscovered every time somebody new sees it. ## What the finding actually asserts Strip the stages away and the claim is about *standing*. At the moment page text is placed into the context window, it acquires the same standing as the operator's own instruction, and nothing in the chain ever re-derived a trust decision about the source — the fetch step selected the page at run time, and every stage downstream was written to move text, not to adjudicate it. That is a property of how the pipeline was composed, not a bug inside any component of it, and it is the level at which the finding is both true and actionable by somebody. The second half of the finding concerns the payoff. The brief cites a URL. A citation establishes that the agent retrieved that page; it does not establish that the page is authoritative, that a second source agreed, or that a human looked. Readers of briefs read citations as vetting, and the laundering effect — a stranger's assertion re-voiced in the assistant's register — is what makes this worth filing at all, independent of the injection mechanics. ## Bug or design limit Both answers are defensible and the interviewer is scoring how you distinguish them, not which you pick. It is a design limit in the sense that no single component is behaving wrongly and no local change removes it. It is a bug in the sense that the system delivers attacker-authored assertions to a decision-maker under its own byline, which nobody chose and nobody documented. The productive write-up says both: name the property, name who could own a change to it, and be explicit that a change inside any text-cleaning stage does not touch it. Saying honestly what a proposed fix does not fix is the whole value of the report. ## Reproducibility, and the trap in it Expect the reproduction to be uneven. Which pages a run-time selection step reaches varies, and whether the model prefers the planted sentence on a given turn is probabilistic. A finding that fires two times in five is still a finding here, because the survival half — the stage record — is deterministic and demonstrable even when the obedience half is not. Separate the two claims in the report: *this text reaches the context unchanged, always* and *the model acted on it in N of M runs*. Merging them into one confident sentence is how these reports lose credibility with the team that has to act on them. ## How to say it in a loop Lead with the record, not the opinion. `Every stage met its spec, and here is the evidence` is a stronger opening than any ownership argument, and it forces the conversation up to the level where the finding is actually true — what standing fetched prose has when it lands in a context window, and what a citation in the delivered brief does and does not prove.
- The reproduction fires in two runs out of five. Does that weaken the report?Only if the two claims are merged. Survival through the chain is deterministic and the stage record demonstrates it every time; whether the model preferred the span on a given turn is probabilistic and belongs in a separate sentence with the count attached. Reported that way, the variance is data. Reported as one confident claim, the first failed re-run discredits the whole finding.
- The team proposes tightening the sanitiser. What do you tell them that buys?It buys what it already bought: markup a browser could execute stays out. It does not touch this finding, because the span is text and the stage is contractually required to preserve text. Say so plainly and say what would still be true afterwards — the same claim, written slightly differently, arrives in the next brief. A fix nobody can falsify is worse than an open finding.
- What does the citation in the delivered brief prove to the analyst reading it?That the agent retrieved that page. Not that the page is authoritative, not that another source agreed, not that anyone reviewed it. Readers treat citations as vetting because in human-written work they usually imply it, and that mismatch is half the payoff of this construction: a stranger's assertion arrives in the assistant's voice with a real URL beside it.
saying these in an interview costs you the question
- Files the finding against the sanitiser and closes it
- Closes it as no-fault because every stage passed
- Treats a citation in the brief as evidence of vetting
- Reports survival and obedience as a single claim
- Demands a fix without saying what it fails to fix
- Assumes an unreliable reproduction means no finding