Your validator calls url.Parse on a callback URL and compares u.Host to an allowlist — what slips through, and how does url.ParseRequestURI differ?
answer
- parsing succeeded says almost nothing
- a reference need not be absolute
- an empty field matches nothing on your list
- two leading slashes still name a host
- the port and the case are yours to normalise
basics
~20 surl.Parse accepts relative references, so a string with no scheme and no host parses without error and leaves u.Host empty — the allowlist comparison then never runs against anything. url.ParseRequestURI accepts only an absolute URI or an absolute path.
solid answer
~50 s`url.Parse` parses a *reference*, not necessarily an absolute URL: `callbacks/return` returns successfully with `Scheme` and `Host` both empty and the whole string in `Path`, and `//other.test/x` returns with `Host` set but `Scheme` empty. So a check shaped like "parse succeeded and the host is not on the deny list" passes for inputs that point nowhere you approved. A correct check requires the scheme explicitly (`u.Scheme != "https"` is a rejection), requires a non-empty host, and compares `u.Hostname()` — which strips the port and the brackets around an IPv6 literal — case-insensitively, since Go lowercases the scheme but not the host. `url.ParseRequestURI` is the stricter parser: it interprets the string only as an absolute URI or an absolute path, so `example.test/x` is an error rather than a relative path, and it assumes there is no `#fragment` because a request target does not carry one.
code
go · 7 linesfor _, s := range []string{"callbacks/return", "/callbacks/return", "//other.test/x"} {
u, err := url.Parse(s)
fmt.Printf("%q err=%v scheme=%q host=%q path=%q\n", s, err, u.Scheme, u.Host, u.Path)
}
// "callbacks/return" err=<nil> scheme="" host="" path="callbacks/return"
// "/callbacks/return" err=<nil> scheme="" host="" path="/callbacks/return"
// "//other.test/x" err=<nil> scheme="" host="other.test" path="/x"go deeper
Know that url.Parse returns a *url.URL with separate Scheme, Host and Path fields, and that a string with no host still parses without an error.
Explain what each parse result contains for a relative path, an absolute path and a scheme-relative reference, and say what url.ParseRequestURI rejects that url.Parse accepts.
Demonstrate a positive, field-by-field check: require the scheme, require a host, compare the hostname without its port and case-insensitively, and rebuild the outgoing URL from validated parts rather than echoing the input.
Own where this check lives. Decide whether callback targets are validated once in a shared helper or re-implemented per service, and prefer registered identifiers over free-form URLs so no team has to get the parsing subtleties right at all.
## `url.Parse` is a parser, not a validator The crucial sentence about `url.Parse` is that it accepts a URL **in the context of a reference**: an absolute URL, a scheme-relative reference, an absolute path, a relative path. All of those are legal, so almost any string parses without error. `err == nil` tells you the text was well-formed; it tells you nothing about where it points. Three results are worth memorising, because they are exactly the ones a host allowlist mishandles: - `url.Parse("callbacks/return")` → `Scheme: ""`, `Host: ""`, `Path: "callbacks/return"`, `err == nil`. A relative path. - `url.Parse("/callbacks/return")` → `Scheme: ""`, `Host: ""`, `Path: "/callbacks/return"`. An absolute path, still no host. - `url.Parse("//other.test/x")` → `Scheme: ""`, `Host: "other.test"`, `Path: "/x"`. A scheme-relative reference: there *is* a host, and it is not yours. A validator written as `if deny[u.Host] { reject }`, or as `if u.Host != "" && !allow[u.Host] { reject }`, waves the first two through on an empty host and treats the third as an ordinary URL. The empty host is the trap: an absent value compares equal to nothing on the list, and "nothing on the list matched" is not the same as "this is fine". ## What a correct check looks like Validate positively, field by field, and reject anything you did not explicitly allow: 1. **Require the scheme.** `u.Scheme != "https"` rejects the relative forms (their scheme is `""`), the scheme-relative form, and anything using another scheme entirely. Go lowercases the scheme during parsing, so a plain equality comparison is right here. 2. **Require a host, and compare the right thing.** `u.Host` includes the port and, for an IPv6 literal, the square brackets. `u.Hostname()` strips both; `u.Port()` gives the port or `""`. Compare `u.Hostname()` case-insensitively — the host is not case-normalised by parsing, so `strings.EqualFold` or `strings.ToLower` is on you. 3. **Be aware of userinfo.** `https://[email protected]/x` has `Host: "other.test"` and `User: "api.example.test"`. The parser gets this right; humans skimming the string do not, which is why the check must read fields rather than eyeball text. 4. **Rebuild rather than echo.** If the target is going into a response, construct the outgoing URL from the parts you validated, rather than passing the caller's original string through untouched. Avoid substituting string tests for parsing. `strings.HasPrefix(raw, "https://api.example.test")` is satisfied by `https://api.example.test.other.test/`, and a prefix test on `/` is satisfied by `//other.test/x`. ## Where `url.ParseRequestURI` fits `url.ParseRequestURI` parses a string that was received **in an HTTP request**, so it interprets the input only as an absolute URI or an absolute path, and it assumes the text has no `#fragment` suffix — browsers strip fragments before sending, so a request target never has one, and it is not split off into `Fragment` the way `url.Parse` splits it. That makes it stricter in a useful way: `url.ParseRequestURI("example.test/x")` returns an error, where `url.Parse` happily calls the same string a relative path. If your input is supposed to be a request target — you are about to put it on a request line, or you are checking something that arrived as one — `ParseRequestURI` is the parser whose contract matches. It is not, however, a general "is this an absolute URL" test either, because an absolute path satisfies it by design. When you need an absolute URL, the honest test is still: parse it, then require the scheme and host you want. ## Composing a URL you were given a piece of When a base comes from configuration and a reference comes from input, `base.ResolveReference(ref)` produces a new `*url.URL` resolved per the reference-resolution rules; it mutates neither operand. It is the right function, but note the ordering it imposes on your checks: a reference carrying its own scheme and host wins outright, and a scheme-relative reference replaces the host. So validate the **result** of the resolution, not the input to it. ## The diagnostic habit When a URL check misbehaves, print the parsed fields rather than the string: `fmt.Printf("%q scheme=%q host=%q hostname=%q path=%q\n", raw, u.Scheme, u.Host, u.Hostname(), u.Path)`. Nearly every one of these bugs is instantly visible as an empty `Scheme`, an empty `Host`, or a `Host` that includes a port you did not expect — and none of them is visible in the string you started with.
- What does u.Hostname() give you that the Host field does not?`Host` is the authority as written, including the port and, for an IPv6 literal, the surrounding square brackets. `Hostname()` strips both, and `Port()` returns the port or an empty string. Comparing the raw `Host` against an allowlist fails the moment a caller appends `:443`, so compare hostnames and treat the port as a separate decision.
- A base URL comes from config and a relative reference from input. How do you combine them in Go?`base.ResolveReference(ref)` returns a new `*url.URL`, leaving both operands untouched. Remember what resolution allows: a reference with its own scheme and host replaces both, and a `//host/path` reference replaces the host. So run your scheme and host checks on the resolved result, not on the reference you were handed.
- When is url.ParseRequestURI the right parser to reach for?When the string is meant to be a request target — an absolute URI or an absolute path — rather than an arbitrary reference. It rejects `example.test/x`, which `url.Parse` accepts as a relative path, and it assumes no `#fragment`, since a request target never carries one. It still accepts a bare absolute path, so it is not by itself a test for absoluteness.
- Why is strings.HasPrefix on the raw string a poor substitute for parsing?A prefix test on `https://api.example.test` is satisfied by `https://api.example.test.other.test/`, because the boundary between host and the rest of the string is a parsing question, not a character-count question. Likewise a prefix test for `/` is satisfied by `//other.test/x`. Parse the string and compare fields.
url.Parse is a spell-checker, not a bouncer: it reports that the string is well-formed, never that it points somewhere you approve of.
saying these in an interview costs you the question
- Assumes a successful url.Parse means the string is absolute
- Checks u.Host but never checks u.Scheme
- Uses strings.HasPrefix on the raw string instead of parsing
- Compares u.Host to an allowlist with the port attached
- Compares host names case-sensitively
- Expects url.Parse to return an error for a relative reference