How do you stop model-authored URLs and images from leaking data at render?
answer
- the payload looks like ordinary text
- constrain the destination, not the string
- parse to nodes, never inject markup
- same-origin proxy with an approved host list
- browser policy as an independent backstop
basics
~20 sAllowlist the destinations instead of inspecting payloads: strip or rewrite every model-authored image source through a same-origin proxy that only fetches approved hosts, restrict link targets to an approved list, never render model-authored raw HTML, and add a browser-enforced content policy as a backstop.
solid answer
~50 sThe payload in a URL beacon is ordinary text, so denylisting content or known-bad domains cannot work. Constrain the sink instead. At render time, parse the model's output into a structured document rather than injecting a string, then apply a destination allowlist: image sources are rewritten to a same-origin proxy that will fetch only approved hosts, or dropped; hyperlinks whose host is not approved are rendered as inert text; raw HTML, inline styles and iframes from model output are never rendered at all. A Content-Security-Policy with a narrow img-src and connect-src gives a second, browser-enforced layer if the sanitizer is bypassed. The remaining judgement call is usability: a real-estate assistant that legitimately shows listing photos needs those hosts on the allowlist, which is exactly why the list should be explicit, small and reviewed, rather than inferred at runtime from what the model produced.
code
typescript · 13 linesconst ALLOWED_IMAGE_HOSTS = new Set(["images.internal.example", "cdn.partner.example"]);
export function rewriteImageSrc(src: string): string | null {
let url: URL;
try {
url = new URL(src);
} catch {
return null;
}
if (url.protocol !== "https:") return null;
if (!ALLOWED_IMAGE_HOSTS.has(url.hostname)) return null;
return "/img-proxy?u=" + encodeURIComponent(url.toString());
}go deeper
Know that a model-written image URL is fetched automatically when a message renders, and that the fix is to restrict which hosts are allowed rather than to look for bad content.
Explain parsing output into nodes and applying a host allowlist, why raw HTML and inline styles must be dropped, and how a same-origin proxy changes who contacts the third party.
Demonstrate layered thinking: server-side sanitizing plus a browser policy that fails independently, plus the discipline of auditing every render surface including mobile, email and link unfurling.
Be ready to own the product tradeoff: which external hosts the assistant may ever reference, who approves additions, and why an allowlist inferred from model or document content is not an allowlist at all.
## Why this is a sink problem, not a content problem When a model writes a URL that carries private data in its query string, nothing about the string is anomalous. The host may be a domain registered an hour ago or a legitimate analytics endpoint; the data may be base64, or plain text, or split across path segments. Any detector you build is deciding whether an ordinary-looking URL was intended, which is undecidable from the string alone. The tractable question is the inverse: is this destination one that somebody approved for this surface? That question has a finite answer, which is why output-sink allowlisting is the control that actually holds. ## The rendering pipeline Start by refusing to treat model output as markup. Parse it into a document model, walk the nodes, and emit only node types you have decided to support. Concretely: **Images.** The highest-risk node because it dereferences with no user action. Two workable policies. Either drop image nodes entirely, which is the right default for an assistant that has no reason to display pictures; or rewrite every source to a same-origin proxy endpoint that resolves the requested host against an allowlist server-side and refuses anything else. The proxy is worth the effort in a product that must display real images, because it also removes the client's direct contact with third-party hosts, strips referrer leakage and gives you one place to log and rate-limit. **Links.** Lower risk because they need a click, but a plausible-looking link in an assistant's answer gets clicked. Render links to non-allowlisted hosts as plain text showing the full URL, so the user sees where it would go, rather than as an active anchor. Never render javascript or data scheme targets. **Everything else.** Raw HTML, inline style attributes, iframes, object and embed elements, SVG with external references, and web-font declarations are all fetch-capable and should not survive the pass. A markdown renderer configured to pass through HTML is the single most common way this control is silently disabled. ## Defence in depth in the browser A Content-Security-Policy is not a substitute for sanitizing, because it is enforced by the client and only constrains what the browser will load. But as a second layer it is cheap and effective: a narrow img-src limits which hosts images may come from even if a beacon slips through the sanitizer, connect-src limits script-initiated requests, and frame-src blocks embedded documents. The value is that these two layers fail independently: one is your code, the other is the user agent. ## The same control applies outside the browser Any surface that dereferences model output needs the same treatment. A mobile client with its own renderer needs its own allowlist and typically does not inherit the web CSP. A chat integration that unfurls link previews will fetch a URL server-side on your behalf, which moves the beacon from the user's browser to your infrastructure, and often makes it invisible to the sanitizer entirely. Email delivery of model-authored content is worse, because mail clients fetch remote images and you no longer control the renderer; the usual answer is to inline or drop images before sending. ## Where the tradeoff bites The honest difficulty is that some products genuinely need model output to reference external resources: a research assistant that cites sources, a shopping assistant that shows product photos, a CRM assistant that surfaces listing images from a partner CDN. The allowlist then has to be a real list, maintained by humans, with hosts added deliberately. Two failure patterns are common. The first is inferring the allowlist from context, for instance permitting any host that appeared in a retrieved document, which is precisely the untrusted input the control exists to contain. The second is allowlisting a host that itself forwards requests, such as a URL shortener, an open image proxy or a documentation site with an open redirect, which hands the whole channel back. That second pattern generalises to a rule worth stating explicitly: an allowlist entry is a promise that the destination does not relay. Check for open redirects on approved hosts, and make the proxy refuse to follow cross-host redirects rather than transparently chasing them. ## What good looks like in review A reviewer should be able to point at one function through which every rendered model output passes, read the node allowlist and the host allowlist from it, and see tests that assert an image node with an unapproved host is dropped rather than proxied. If the answer is that the renderer is configured safely by default, ask what happens on the mobile client, in the email path and in the link-unfurling integration, because those are usually three different renderers.
- Your sanitizer is solid, but the assistant's answers are also posted into a team chat integration that unfurls links. What changed?The egress moved off the browser and onto a server you do not control, which fetches the URL to build a preview. Your renderer never sees it, so the sanitizer and the CSP are both bypassed. Either strip URLs from content destined for that integration, or post through a formatting mode that suppresses unfurling, and treat every downstream surface as its own sink with its own allowlist.
- Why can allowlisting a URL shortener or a documentation host with an open redirect undo the whole control?An allowlist entry asserts that the destination will not relay the request elsewhere. A shortener or an open redirect does exactly that, so the model can encode data in a path the approved host forwards to an arbitrary destination. Verify approved hosts do not redirect off-host, and configure the proxy to reject cross-host redirects instead of following them.
- Is a Content-Security-Policy enough on its own?No. It is enforced by the browser, applies only to that client, and covers only the resource types its directives name. It is valuable precisely because it fails independently of your sanitizer, but a native app, an email client or a server-side unfurler ignores it entirely. Treat it as the second layer behind server-side sanitizing, never as the first.
saying these in an interview costs you the question
- Denylisting known attacker domains instead of allowlisting destinations
- Configuring the markdown renderer to pass raw HTML through
- Trusting a Content-Security-Policy as the only defence
- Building the allowlist from hosts seen in retrieved documents
- Forgetting mobile, email and link-unfurling render paths