skip to content

When designing resource URIs for an HTTP API, how do you choose between an opaque surrogate identifier such as `/users/8f3c...` and a natural key such as `/users/[email protected]` or `/products/SKU-1234`? What are the consequences of each?

level: seniorimportance: should knowfreq 45%

answer

  1. natural keys mutate, get reused, need escaping, leak data
  2. surrogate = stable + neutral + uniform
  3. prefix ids (cus_, ord_) for legible logs and wrong-type 400s
  4. lookup by natural key = filtered collection or 301 alias
  5. opaque ≠ authorization; still check object-level access

basics

~20 s

Natural keys are readable but change, collide across scopes, and leak data into URLs and logs. Surrogate ids are stable and neutral, so make them canonical. Support natural-key lookup as a filtered query or a documented alias that redirects to the canonical URI.

solid answer

~60 s

**Natural keys** (email, SKU, slug, username) read well and let clients construct URLs from data they already have. Their costs are real: they *change* — an email or slug edits, and every stored URI, cache entry and log reference now points at nothing or, worse, at someone else after reuse; they need escaping (`@`, `/`, `+`, unicode, case rules); they leak personal or commercial data into URLs, proxy logs and `Referer` headers; and uniqueness is often scoped, not global. **Surrogate ids** (UUID, ULID, prefixed opaque strings) are stable for the resource's lifetime, uniform to route and validate, and carry no domain meaning. Costs: unreadable in logs, an extra lookup when a client only has the natural key, and sequential ones enable enumeration. My default: **surrogate id is canonical in the URI path.** Expose natural-key access as `GET /[email protected]` returning a collection, or as an alias path that answers `301`/`308` to the canonical URI. Never let a mutable natural key be the only address, and never let a reusable one be an address at all.

code

http · 9 lines
http
GET /users?email=alice%40example.com HTTP/1.1

200 OK
{"data":[{"id":"usr_01HX","email":"[email protected]"}]}

GET /users/by-email/alice%40example.com HTTP/1.1

308 Permanent Redirect
Location: /users/usr_01HX

go deeper

for a junior

Say that ids in URLs should be stable and that emails or usernames can change, so an opaque id is safer.

for a middle

Add the encoding, scoping and lookup-by-filter points, and note that natural keys are fine when an outside authority guarantees immutability.

for a senior

Lead with reuse and retargeting, disclosure through logs and referrers, enumeration versus authorization, and the canonical-plus-alias pattern with permanent redirects.

for a principal

Set the identifier policy for the platform: prefixed, sortable opaque ids as canonical, natural keys as filtered lookups, documented redirect behaviour, and the migration story when a key must change.

## Why the URI identifier is a durable decision A URI is a name clients store: in databases, webhooks, logs, bookmarks, cached entries, and other systems' foreign references. Whatever you put in the path becomes part of your contract for as long as anyone holds it. So the question is not just readability; it is whether the identifier can be relied upon to mean the same thing forever. ## Natural keys A natural key is a domain value that already identifies the thing: an email, a username, a SKU, an ISBN, a slug. **Advantages.** Readable and self-describing in logs and support conversations. Clients can build a URL from data they already hold, avoiding a lookup round trip. For genuinely immutable, externally-governed identifiers (ISBN, ISO country code, a currency code), they are excellent — the identifier is stable because an outside authority guarantees it. **Mutability.** Most business keys are editable. A user changes their email; a product's SKU is corrected; a slug is renamed for SEO. Every stored URI is now stale. You can mitigate with permanent redirects from old keys to new, but that means keeping a forever-growing alias table and deciding what to do about a key someone else later takes. **Reuse and collision.** The dangerous case: a natural key that can be released and re-assigned. If `@alice` is freed and taken by someone else, a URI stored six months ago now resolves to a different person. Every reference — an audit log, a webhook payload, an ACL entry — silently retargets. This alone disqualifies usernames and emails as canonical identifiers in any system with real security requirements. **Encoding.** Natural keys contain characters that fight with URIs: `@`, `+`, `/`, `.`, spaces, unicode, case. `[email protected]` needs percent-encoding, and proxies, frameworks and clients disagree about decoding, especially for an encoded `%2F`. Case sensitivity is another trap: emails are commonly case-insensitive in the local part in practice, but URIs are case-sensitive, so `/users/[email protected]` and `/users/[email protected]` are two URIs for one resource unless you normalise and redirect. **Disclosure.** Putting an email or a customer name in a path leaks it into access logs, proxy logs, browser history, and `Referer` headers on outbound links. That is a privacy exposure that surrogate ids simply do not have, and it is often the point that ends the debate in a regulated environment. **Scope.** Many natural keys are unique only within a tenant or a catalogue. `/products/SKU-1234` is ambiguous the day a second seller uses the same SKU, and retrofitting scope into the path is a breaking change. ## Surrogate identifiers **Stability** is the whole point: the id is assigned at creation and never changes, so a URI stored today is valid for the resource's life regardless of any domain edit. **Neutrality** means no leakage and no encoding hazards — a UUID or a base32 ULID is URL-safe by construction. **Uniformity** simplifies routing, validation, and caching: one format to parse and reject. **Enumeration** is the risk to manage. Sequential integers advertise volume (`/orders/1041` tells a competitor your order count) and invite id-guessing. Random UUIDv4 or ULIDs remove practical enumeration, but *never* treat unguessability as authorization — object-level access control is still required on every request; an opaque id is defence in depth, not a control. **Readability** is the real cost. Mitigate with a type prefix (`cus_8f3c…`, `ord_2b19…`): it makes logs and support tickets legible, catches "customer id passed where an order id was expected" as a `400` rather than a mysterious `404`, and gives you a namespace for future migration. **Sortability** matters at scale: ULID or UUIDv7 keep creation-time ordering, which is friendlier to database index locality than UUIDv4. ## The pattern that works Make the surrogate canonical and give natural keys a first-class lookup path: ``` GET /users/usr_01HX... # canonical GET /[email protected] # filtered collection, 200 with 0 or 1 items GET /by-email/alice%40example.com # optional alias -> 301/308 to canonical ``` Using a *collection with a filter* for lookup is often cleanest: it sidesteps encoding-in-path issues (query values escape more predictably), returns `200` with an empty list rather than a `404` you must interpret, and generalises to multiple matches when the key turns out not to be unique after all. If you offer an alias path, answer with a permanent redirect so clients learn and store the canonical URI. Human-friendly slugs can coexist: `/articles/usr_01HX.../my-title` or `/articles/my-title-1041` where the numeric part is authoritative and the slug is decorative — a mismatched slug redirects to the canonical form. That gives readable URLs without making a mutable string load-bearing. ## Deciding quickly Use a natural key in the path only when it is (a) immutable, (b) never reused, (c) globally unique in your scope, (d) URL-safe, and (e) not sensitive. ISO codes and ISBNs pass. Emails, usernames, SKUs and slugs fail at least one, usually three.

  • A username can be released and later claimed by someone else. Why does that rule it out as a canonical URI segment?
    Because every stored reference silently retargets. An audit entry, ACL, webhook payload or bookmark captured months ago now resolves to a different person, with no error to detect. Identifiers used in security or audit contexts must never be reusable, so the canonical URI has to use a surrogate id.
  • Are opaque random ids a security control?
    No. Unguessability is defence in depth at best: ids leak through logs, referrers, screenshots, and shared links. Every request must still perform object-level authorization for the caller against that specific resource. Random ids only remove trivial enumeration; they do not replace an access check.
  • Why prefix identifiers with a short type token like `cus_` or `ord_`?
    It makes logs and support conversations legible without exposing domain data, and it lets the server reject an id of the wrong type with a clear `400` instead of a puzzling `404`. It also gives you a namespace, so you can change the encoding behind the prefix later without ambiguity.

saying these in an interview costs you the question

  • Using an email or username as the canonical path segment because it is readable
  • Treating an unguessable UUID as if it were an authorization control
  • Ignoring that natural keys in paths leak into access logs and Referer headers
  • Assuming a SKU or slug is globally unique when it is only unique per tenant or catalogue
  • Exposing sequential integer ids without considering enumeration and volume disclosure

context