skip to content

A page is served from https://shop.example.com/cart. In browser terms, what is that page's origin, and which of https://shop.example.com:443/help, http://shop.example.com/cart and https://api.example.com/cart are same-origin with it?

level: juniorimportance: must knowfreq 80%

answer

  1. three parts, not four
  2. path and query never count
  3. scheme, host, port
  4. default port 443 normalises away
  5. a subdomain is a different host

basics

~20 s

An origin is scheme plus host plus port: here https://shop.example.com on port 443. Only the :443/help URL is same-origin; switching the scheme to http or the host to api.example.com produces a different origin. Path is never part of it.

solid answer

~40 s

An origin is the triple **scheme + host + port**, so this page's origin is `https://shop.example.com` with the implicit port 443. Comparing the three URLs: `https://shop.example.com:443/help` is **same-origin** — 443 is the default port for https, and the path is not part of an origin at all. `http://shop.example.com/cart` is **cross-origin** because the scheme differs; an http page can be tampered with on the wire, so the browser refuses to treat it as the same trust domain as the https one. `https://api.example.com/cart` is **cross-origin** because the host is compared as a whole string — there is no "same parent domain, therefore same origin" rule. At runtime you can read the value as `location.origin`, and `new URL(href).origin` computes it for any URL.

code

javascript · 15 lines
javascript
const base = new URL("https://shop.example.com/cart");

for (const href of [
  "https://shop.example.com:443/help",
  "http://shop.example.com/cart",
  "https://api.example.com/cart",
  "https://shop.example.com:8443/cart",
]) {
  const other = new URL(href);
  console.log(other.origin, other.origin === base.origin);
}
// https://shop.example.com       true
// http://shop.example.com        false
// https://api.example.com        false
// https://shop.example.com:8443  false

go deeper

for a junior

Be able to say "scheme, host, port" without hesitating, and to judge a pair of URLs on the spot. Remember that the path never matters and that a different subdomain is a different origin.

for a middle

Explain why each component is in the tuple — especially why the scheme is, since a tamperable http page must not inherit an https page's storage — and know that default ports normalise away in the serialized form.

for a senior

Show you reason about origins as trust boundaries in a real deployment: which hosts you split user content onto, why storage keyed by origin and cookies keyed by host diverge, and what that means for staging and dev subdomains.

for a principal

Own the origin layout of a product. Decide how many hosts the system needs, what gets its own origin because it must never script the main app, and what the cost of that split is in cookies, sessions and deployment complexity.

## The origin tuple The unit of trust on the web is the *origin*, and an origin is exactly three things: the **scheme**, the **host**, and the **port**. Everything else in a URL — path, query string, fragment, credentials — sits outside the tuple and has no effect on origin comparisons. Two URLs are same-origin when all three components match; if any one differs they are cross-origin, and the browser applies the same-origin policy between them. Serialized, an origin looks like `https://shop.example.com`: scheme, `://`, host, and the port *only* when it is not the default for that scheme. `https` defaults to 443 and `http` to 80, so `https://shop.example.com` and `https://shop.example.com:443` are the identical origin written two ways. `https://shop.example.com:8443` is a different one. ## Walking the example The page at `https://shop.example.com/cart` has origin `https://shop.example.com` — scheme `https`, host `shop.example.com`, port 443. - **`https://shop.example.com:443/help` — same origin.** The explicit `:443` is the default port for https and normalises away, and `/help` versus `/cart` is path, which the comparison ignores entirely. - **`http://shop.example.com/cart` — different origin.** Only the scheme changed, and that is enough. A plaintext http response can be rewritten by anyone on the network path, so if the browser let it share a trust domain with the https page, an attacker who could tamper with one request would inherit the secure page's DOM and storage. The scheme is in the tuple precisely to stop that. - **`https://api.example.com/cart` — different origin.** Hosts are compared as whole strings. `api.example.com` and `shop.example.com` are as different to this check as `example.com` and `attacker.example`. Sharing a registrable domain makes them the same *site*, which is a looser notion some other browser features use, but it does not make them the same origin. ## Reading the origin at runtime `location.origin` gives the current document's serialized origin. `window.origin` (and `self.origin`, which also works inside workers) gives the same string. The neighbouring `location` properties are easy to confuse: `location.protocol` returns `"https:"` *with* the trailing colon, `location.hostname` is the host alone, `location.host` is host plus port when the port is non-default, and `location.port` is the empty string when the default port is in use. ```js const u = new URL("https://shop.example.com:443/cart?id=7#top"); u.origin; // "https://shop.example.com" u.hostname; // "shop.example.com" u.port; // "" (default port normalised away) ``` ## Where the tuple actually shows up The same three components decide a surprising amount of platform behaviour. `localStorage`, `sessionStorage` and IndexedDB are partitioned by origin, so an https page and an http page on the same host see two separate stores. The `Origin` request header carries the serialized origin. Cross-document messaging takes an explicit target origin string that must match. Every "blocked by the same-origin policy" message in DevTools is ultimately a tuple mismatch. Cookies are the notable exception: they are keyed by host and path rather than by origin, ignore the port completely, and historically ignored the scheme too. That mismatch — storage keyed by origin, cookies keyed by host — is the source of a lot of confusion when people assume one model covers both. ## Why three components, and why not four Each component earns its place. The host stops an unrelated site from claiming your data. The scheme stops a downgraded, tamperable connection from inheriting a secure page's privileges. The port is included because two servers on one host can be operated by different people — shared hosting, a dev server on `:3000`, a debug endpoint on `:8080`. The path is deliberately *excluded*, and that omission has real consequences. Everything on one origin is a single trust domain: `https://host/~alice` and `https://host/~bob` share cookies and storage, and script on either page can fully script the other. That is why hosting untrusted user content under a path of your main domain gives it your whole origin, and why platforms that host user content put it on a separate host instead.

  • Where can a script read its own origin, and how does that differ from location.host?
    `location.origin` (and `window.origin`/`self.origin`, which also works in workers) gives the serialized scheme+host+port, for example `https://shop.example.com`. `location.host` is only host plus non-default port — no scheme — and `location.hostname` is the host alone. For an arbitrary URL string, `new URL(href).origin` computes the same serialization.
  • Does the Origin request header always equal location.origin?
    Usually, but not always. It carries the serialized origin, browsers omit it on many same-origin GET and HEAD requests, and documents with an opaque origin — a sandboxed frame or a `data:` URL document — send the literal string `null`. So a server must never treat `Origin: null` as "one of my own pages".
  • Why does the origin comparison exclude the path?
    Because an origin is one trust domain end to end. Everything on it shares cookies and storage and can script everything else, so `/~alice` and `/~bob` on one host cannot be isolated from each other by the browser. That is why untrusted user content belongs on a separate host, not a separate path.

saying these in an interview costs you the question

  • Says two subdomains of one domain are the same origin
  • Thinks the path or query string is part of the origin
  • Treats http and https on one host as one origin
  • Believes an explicit :443 makes a different origin than the default
  • Confuses origin with the domain name alone

context