A research assistant renders an attacker's page-declared byline and publisher in a source chip: what does glancing at it verify?
answer
- ask where each displayed field came from
- pages describe themselves
- the chip renders, it does not attest
- claim and credentials, one author
basics
~10 sNothing about authorship. The title, byline and publisher in a source chip are values the fetched page declared about itself, so a page the attacker wrote supplies its own credentials alongside its claim.
solid answer
~50 sA source chip is a rendering of metadata, not an attestation. The assistant fetched a page, read the title, byline, publisher and date that the page declares about itself, and drew those strings beside the answer. If the attacker published the page, they wrote both halves: the claim and the credentials shown next to it. A reader's glance feels like a verification step, because an institution's name is doing the work of a signature, but the only thing established is that some retrieved text carried that name. Note what this is not. It is not the model inventing an attribution, and it is not evidence that anything was breached: the page was fetched normally and rendered faithfully. That is what makes the construction cheap. The attacker never has to defeat the interface, only to fill in the fields it will display.
go deeper
Be ready to say, in one sentence, where a source chip's title, byline and publisher come from: the fetched page declares them about itself. Then say what follows - the same person wrote the claim and the credentials next to it.
Explain the pipeline path the values take, from page metadata to extraction to rendering, and why nothing on that path attests to authorship. Be able to separate this from model invention and from prompt injection without hesitating.
An interviewer expects you to route the finding correctly. Nothing here is a generation defect or a store compromise, so a fix aimed at the model is aimed at the wrong layer, and you should be able to say so plainly.
Own the framing question: is displaying self-declared provenance in a position readers read as verification a defect or an accepted product limit? That call decides who is accountable when a third party's name ends up on a claim they never made.
## What a source chip is actually made of A research assistant that answers from fetched web pages usually renders a small provenance strip beside each answer: a title, a publisher name, often a byline and a date, sometimes a favicon. Those values are not the product of the assistant judging who wrote the page. They are read off the page itself, from the title, author, publisher and date fields the document declares about itself in its markup and structured data, plus the host it was served from. That makes every field in the strip **content**. It is under the control of whoever wrote the page, in exactly the way the sentences in the body are. ## Why that matters once the writer is the attacker The construction in this leaf is the whole of it: the attacker publishes a page, writes a claim in it, and *also writes the provenance the interface will display beside that claim*. One author supplies the assertion and the evidence for it. A reader who glances at the chip and sees a familiar institution has performed a check whose entire input the attacker controlled. Get the direction of the claim right, because interviewers listen for it: | What the reader believes the chip shows | What it actually shows | | --- | --- | | A named institution stands behind the claim | Retrieved text carried that institution's name in a metadata field | | The assistant confirmed the source | The assistant rendered strings it read | | Someone published this | A page declared itself published by someone | The chip is faithful. The pipeline is not malfunctioning. Faithful rendering of a self-declared string is precisely the property being used. ## Why this is not a hallucination, and not injection either Two classifications get reached for and both are wrong here. - **Model invention.** If the assistant had produced a publisher name with nothing retrieved to support it, that would be the generation making something up. Here the name was in the fetched artefact; the model repeated a value that was really there. - **Prompt injection.** Injection turns the application's own instructions around using data it feeds the model, so there is a directive somewhere in the text. This construction contains no directive at all. Nothing tells the assistant to do anything. The planted page just carries a claim and a costume, and the pipeline does its ordinary job with both. Controls that look for instruction-shaped text have nothing to find. ## What the attacker is actually buying The payoff is **attribution**. The point is not to make the assistant misbehave; it is to fix a recognisable name to a sentence that name never said, so that a person who reads the answer forwards it onward under that name. In a consumer product where text is the only thing that leaves, that is the whole of the damage, and it lands on a third party who was never in the conversation and may never learn where the quote came from. That also explains the target selection. The construction only pays when the impersonated name is one the reader recognises without looking it up. An obscure publisher buys nothing, because there is no borrowed credibility to spend. ## Where it is weak Worth knowing even at a first encounter, because an interviewer will push here. The construction depends on the reader's check staying inside the attacker's artefact. Any move that leaves it - going to the institution's own site instead of following the chip, or asking whether they published such a thing - is outside the attacker's control and returns nothing. It also depends on the page remaining reachable; once it is pulled, the claim can still be circulating in screenshots and forwarded text while the artefact that produced it no longer exists. ## How to answer this in an interview Say where the fields come from, in one sentence, before saying anything else: they are declared by the fetched page. Then name the consequence: the attacker wrote the evidence as well as the claim, so checking the citation checks nothing they did not supply. Then separate this from invention and from injection. A candidate who reaches straight for 'the answer cites its source, so we can check it' has skipped the only step that matters.
- If the assistant repeated a publisher name that really was on the page, is this a model failure at all?Not in the generation. The model reproduced a value present in the retrieved text, and the interface rendered it faithfully. The failure is that a self-declared string is being shown in a position readers treat as verification. Calling it a hallucination sends the finding to the wrong owner and guarantees the wrong fix gets tried.
- Why does an attacker choose a well-known publisher name rather than an invented one?Because the payoff is borrowed credibility, and borrowed credibility requires recognition. A name the reader has to look up buys nothing; a name they recognise instantly converts the glance into belief and makes the answer worth forwarding. It also narrows the useful target set to a handful of names per audience, which is a real constraint on the construction.
- Does the attacker need the assistant to obey anything?No, and that is the point. There is no instruction in the artefact. The page is fetched, read and cited by a pipeline working exactly as designed. Anything watching for directive text, tool misuse or refused content sees a completely ordinary turn.
A parcel arrives with a return address printed on it. The return address tells you what the sender typed, not who they are.
saying these in an interview costs you the question
- Says the answer cites a source, so it can be checked
- Treats a rendered publisher name as an identity check
- Calls it a hallucination when the value was really on the page
- Assumes the index or the store must have been compromised
- Looks for a hidden instruction in a page that contains none