skip to content

Your client-side request cache stores responses keyed by URL string, and identical requests keep missing the cache. Using the URL and URLSearchParams APIs, how would you canonicalise a URL into a stable key, and what does the parser normalise for you already?

level: seniorimportance: should knowfreq 35%

answer

  1. parse first, never key on the raw string
  2. the parser handles case, port, dot segments
  3. the fragment never leaves the browser
  4. make parameter order irrelevant
  5. trailing slash is your decision, not the spec's

basics

~20 s

Parse with the URL constructor, which already lower-cases the scheme and host, drops a default port and resolves dot segments, then finish the job yourself: drop the fragment, sort the query with URLSearchParams.sort(), and remove parameters that do not affect the response.

solid answer

~50 s

Build the key from a parsed `URL`, never from the raw string. The parser gives you a lot for free — it lower-cases the scheme and hostname, drops `:443` on `https`, resolves `.` and `..` segments, and percent-encodes consistently — so `https://API.Test:443/v1/../v1/items` and `https://api.test/v1/items` already agree. What it will not do is decide semantics for you. Strip `url.hash`, since the fragment is never sent to the server and cannot change the response. Call `url.searchParams.sort()` so `?b=2&a=1` and `?a=1&b=2` produce the same string. Delete parameters you know are irrelevant — `utm_*`, cache-busters, a random `_t`. Then decide explicitly whether `/items` and `/items/` are the same resource, because the parser will not merge them. Serialise the result with `url.origin + url.pathname + (url.search)`, and remember the key must also cover anything else that varies the response, such as the request method or an auth scope.

code

javascript · 17 lines
javascript
function cacheKey(input, base) {
  const url = new URL(input, base);
  url.hash = '';
  for (const name of [...url.searchParams.keys()]) {
    if (name.startsWith('utm_') || name === '_t') url.searchParams.delete(name);
  }
  url.searchParams.sort();
  if (url.pathname.length > 1 && url.pathname.endsWith('/')) {
    url.pathname = url.pathname.slice(0, -1);
  }
  return url.origin + url.pathname + url.search;
}

const a = cacheKey('HTTPS://API.Test:443/v1/items/?b=2&a=1&utm_source=x#top');
const b = cacheKey('https://api.test/v1/./items?a=1&b=2');
console.log(a);
console.log(a === b);

go deeper

for a junior

Know that the URL constructor already lower-cases the host and cleans up the path, and that sort() on the search parameters makes two differently ordered query strings match.

for a middle

Explain precisely which normalisations the parser performs — scheme and host casing, default port removal, dot-segment resolution — and which ones, like parameter order and trailing slash, are left to you.

for a senior

Show the operational reasoning: strip the fragment and known noise parameters, pin the rules with collision tests, keep the client's trailing-slash rule aligned with the server's so you do not trade misses for redirects, and watch the hit rate as the signal.

for a principal

Own it as a cross-boundary contract — the same canonical form must be used by the client cache, any CDN key, and the server's own caching, and the key must encode auth scope so a performance optimisation never becomes a data-leak between users.

## Why identical requests miss the cache A URL is a string, but two different strings can name the same resource. If the cache key is the raw string the caller happened to write, then `?a=1&b=2` and `?b=2&a=1`, `https://API.test/x` and `https://api.test/x`, and `/x#top` and `/x` are four keys for two resources. Canonicalisation is the step that maps every spelling of a request onto one key. ## What the URL parser already normalises Parsing with `new URL(input)` runs the WHATWG algorithm, which performs several normalisations before you touch anything: ```js new URL('HTTPS://API.Test:443/v1/./sub/../items?q=1').href; // 'https://api.test/v1/items?q=1' ``` Specifically: the scheme is ASCII-lower-cased; the host is lower-cased and, if it contains non-ASCII characters, converted to its Punycode form; a port equal to the scheme's default (`80` for `http`, `443` for `https`) is removed and `url.port` becomes `''`; `.` and `..` path segments are resolved; an empty path becomes `/`; and characters outside each component's safe set are percent-encoded. That is the free half of the work, and it is already more than most hand-written normalisers get right. What the parser deliberately does not do is make semantic judgements. It will not decide that `/items` and `/items/` are the same resource, that `?utm_source=x` is noise, or that the order of query parameters is irrelevant. Those are application decisions, and the spec cannot make them for you. It also leaves existing percent-escapes as written rather than re-casing them, so `%7E` and `%7e` stay distinct in the string even though they denote the same byte. ## The application half ```js function cacheKey(input, base = location.href) { const url = new URL(input, base); // 1. the fragment never reaches the server url.hash = ''; // 2. drop parameters that do not vary the response for (const name of [...url.searchParams.keys()]) { if (name.startsWith('utm_') || name === '_t') url.searchParams.delete(name); } // 3. make parameter order irrelevant url.searchParams.sort(); // 4. decide on trailing slashes explicitly if (url.pathname.length > 1 && url.pathname.endsWith('/')) { url.pathname = url.pathname.slice(0, -1); } return url.origin + url.pathname + url.search; } ``` Four things are worth calling out. **The fragment.** `url.hash` is client-only — the browser strips it before sending the request — so it can never distinguish two responses. Clearing it also removes the `#` from the serialisation. **Deleting while iterating.** `url.searchParams.keys()` is a live iterator over the same list you are mutating, so deleting during the loop skips entries. The `[...url.searchParams.keys()]` snapshot avoids that, and note that repeated names make `keys()` yield the same name more than once, which is harmless here because `delete(name)` is already all-or-nothing. **`sort()`.** It orders pairs by name using code-unit comparison, and it is stable for pairs that share a name — so `tag=b&tag=a` keeps `b` before `a`. That is correct: for repeated keys, order can be meaningful to the server, and sorting values would change the request. If your server treats repeated values as a set, sort them yourself before appending. **Trailing slash.** There is no universally right answer; `/items` and `/items/` are different URLs and *may* be different resources. Pick a rule, apply it in one place, and make sure it matches whatever the server and any CDN do — a mismatch turns into a redirect on every request, which is worse than the cache miss you started with. ## What the key must also include A URL alone under-specifies a response. Two requests to the same URL can legitimately return different bodies when they differ in method, in headers the server varies on (`Accept`, `Accept-Language`, `Authorization`), or in credentials. If your cache stores anything user-specific, the key must incorporate an identity or scope component, or one user's response will be served to another — a correctness and privacy bug, not a performance one. Conversely, over-keying (including a timestamp, a request id, a nonce) guarantees a permanent miss, which is usually the actual cause when "identical" requests never hit. ## Verifying it Canonicalisation logic is cheap to unit-test and expensive to get wrong, so pin it with a table of pairs that must collide and pairs that must not: `?a=1&b=2` versus `?b=2&a=1` collide; `/x#top` and `/x` collide; `/x` and `/x/` collide or not according to your stated rule; `?tag=a&tag=b` and `?tag=b&tag=a` must **not** silently collide unless you decided they are equivalent. Instrument the hit rate too — a cache whose hit rate is near zero is almost always a keying bug rather than a workload that genuinely never repeats. ## Why interviewers ask The question looks like trivia about `sort()` but is really about knowing where the platform's guarantees stop and your product decisions begin. A candidate who says "parse it, then decide about fragment, order, noise parameters and trailing slash — and make sure the key covers everything that varies the response" has clearly operated a cache rather than configured one.

  • Does URLSearchParams.sort() also reorder the values of a repeated key?
    No. It sorts pairs by name using code-unit comparison and is stable, so `tag=b&tag=a` keeps `b` first. That is deliberate — for a repeated name the order can carry meaning for the server. If your API treats repeated values as an unordered set, sort those values yourself and rebuild the parameter before serialising.
  • Why must you snapshot the keys before deleting parameters in a loop?
    `keys()` returns a live iterator over the same pair list you are mutating, so removing an entry mid-iteration shifts the remaining pairs and skips one. Spreading into an array first — `[...params.keys()]` — detaches the traversal from the mutation. It is the same hazard as editing any list while walking it.
  • Is including the fragment in a cache key ever justified?
    Not for a network cache: the fragment is never sent, so it cannot change the response. It can be legitimate for a purely client-side view cache, where `#tab=billing` selects which rendered state you memoise. Keep the two caches separate rather than smuggling client-only state into a key that is supposed to model a server resource.
  • What besides the URL belongs in the key for a request cache?
    Anything that varies the response: the HTTP method, the headers the server declares it varies on such as `Accept` or `Accept-Language`, and — critically — the authenticated identity or scope. A cache keyed on URL alone will happily serve one user's personalised response to another, which is a privacy defect rather than a slow path.

saying these in an interview costs you the question

  • Keying the cache on the raw, unparsed URL string
  • Assuming the parser sorts query parameters for you
  • Including the fragment, which never reaches the server
  • Treating /x and /x/ as automatically the same resource
  • Ignoring auth scope so responses leak between users

context