An `Origin` header arrives as `https://board.example.org` — why does a configured entry written `https://board.example.org/` never byte-match it?
answer
- compare bytes, not URLs
- the sender decides the canonical spelling
- an origin has no path component
- one byte too many is a stranger
- a default port must be omitted
basics
~20 sThe arriving value is a serialized origin, which has no path component and therefore no trailing slash. The entry is one byte longer, so an equality test over bytes fails even though a human reads the two strings as the same site.
solid answer
~40 sWhat arrives is the 25-byte sequence `https://board.example.org` — scheme, `://`, host, and no port because 443 is the https default. A serialized origin has no path component, so there is no trailing slash to match; the configured string is 26 bytes and the comparison is an equality test over bytes, not a URL-aware comparison that normalises one side into the other. The same failure has three common spellings: a trailing slash, a host written in mixed case, and an explicit `:443` on an https entry. All three are legal URLs and none of them is a serialized origin. The fix is to write the configured side in exactly the form a browser produces.
code
http · 3 linesGET /v1/departures?stop=42 HTTP/1.1
Host: api.departures.example
Origin: https://board.example.orggo deeper
Remember that the value is compared as raw bytes. If the string someone wrote down differs anywhere — a slash, a capital letter, a port — it is simply not equal, and there is no partial credit.
Explain the grammar behind it: an origin has scheme, host and a conditional port and no path component, so a trailing slash cannot appear. Name the three near-miss spellings and why each one is wrong.
Demonstrate the diagnosis. Read the raw arriving field, compare its byte length to the configured entry, and identify a one-byte difference as a slash and a four-byte one as an explicit default port before touching any code.
Treat the canonical spelling as a data-quality problem: origins written by hand in free text will drift, so the durable answer is to derive the entry from the same serialization rule rather than to trust what someone pasted.
A departures host receives calls from many embedded arrivals boards, and every one of those calls arrives carrying a value that the browser produced mechanically. Somewhere on the receiving side that value is compared against a written-down list. The comparison is where a correct-looking deployment silently stops working. ## The arriving value is a serialized origin, not a URL ```http GET /v1/departures?stop=42 HTTP/1.1 Host: api.departures.example Origin: https://board.example.org ``` `serialized-origin = serialized-scheme "://" serialized-host [ ":" serialized-port ]`. That is the whole grammar. It has three components and the third is emitted only when the port differs from the scheme's default (`80` for http, `443` for https). It has **no path component at all**, which is why there is no trailing slash: a slash is the start of a path, and there is no path here to start. So `https://board.example.org` is exactly 25 bytes, and that is what the comparison has to be against. ## The four written forms that look right and are not | What someone wrote | Why it is not the arriving value | |---|---| | `https://board.example.org/` | 26 bytes — a trailing slash is an empty path, and an origin has no path component | | `https://Board.Example.org` | the arriving host is ASCII-lower-cased during serialization | | `https://board.example.org:443` | a port is serialized only when it is **not** the scheme's default, and 443 is | | `https://board.example.org/embed` | a path, which is never serialized into an origin | Every one of these is a perfectly valid URL, and a person reading any of them aloud says "the board host". None of them is a serialized origin, and the comparison is not done by a person. ## Why the comparison is over bytes The value is a short, canonical byte sequence precisely so that a decision can be made on it without parsing. A URL-aware comparison would have to decide, for every pair, whether an empty path equals no path, whether an explicit default port equals an omitted one, and whether case folding applies to the host but not to a path — the same normalisation questions that make URL equality famously subtle. The serialization moves all of that work to the sending side: the browser produces one canonical spelling, and the receiving side does an equality test. The consequence is blunt, and it is the whole point of the leaf: - **The canonical form is decided by the sender.** Nothing on the receiving side gets a vote on how the value is spelled. - **A near-miss fails exactly like a stranger.** An entry that differs by one byte is not "nearly matched" — it is simply not equal, and the outcome is identical to an origin nobody ever heard of. - **The failure is invisible in the calling code.** The embed makes the same call it always made; nothing in the board's own source mentions the byte that broke it. - **Both sides must be written the same way.** Since the sending side's spelling is fixed by the specification, the configured side is the one that has to be written as a serialized origin. ## How this shows up in practice Someone adds a new host page for the arrivals board and records the new origin by copying it out of an address bar, which renders a root document with a trailing slash. Or an operator types the host in the capitalisation used in a company document. Or a careful engineer writes `:443` because being explicit feels safer than relying on a default. Each of those produces a string that no browser will ever send, and a board that loads its markup perfectly and then shows no departures. The diagnostic that settles it in seconds is to read the raw request and copy the value of the field as literal bytes, then compare its length to the configured entry's length. A difference of one is a trailing slash. A difference of four on an https entry is `:443`. A difference of zero with an inequality is a case fold. ## Stating the rule so it survives The arriving value is scheme, `://`, host, and a port only when non-default, ASCII-lower-cased, with nothing after it. Whatever is written on the other side should be constructible by that same rule, and if it cannot be, it is not an origin and will not match.
- Why does an entry written `https://board.example.org:443` fail against an https board?Because the port is serialized only when it differs from the scheme's default, and `443` is the default for `https`. The browser therefore sends `https://board.example.org` with no port, and the explicit entry is four bytes longer. Written against an `http` board, `:443` would be non-default and would be sent — but then the scheme would not match either.
- The board moves to `http://board.example.org:8080` for a local trial. What arrives now?`http://board.example.org:8080`. The scheme is different, and 8080 is not the default for `http`, so the port is serialized too. It is a different origin from the https one in two components at once, and any entry written for the production board is simply a different byte sequence.
- Is there any spelling of a path that would match?No. A serialized origin has no path component, so no value carrying a path — not `/`, not `/embed`, not `/.` — can ever be produced by the serialization. An entry containing any path is a string a browser will never send, whatever the path is.
saying these in an interview costs you the question
- Says a trailing slash is cosmetic and gets ignored
- Assumes a URL-aware comparison normalises the two forms
- Writes an explicit :443 into the entry to be safe
- Believes a mixed-case host still matches because names are case-insensitive
- Blames the browser for stripping something from the header