skip to content

What are the risks of exposing a table's internal primary key values in public URLs and API responses, and what are the alternatives?

level: seniorimportance: should knowfreq 45%

answer

  1. Enumeration widens a missing authz check into bulk extraction
  2. Sequential ids leak counts and growth rate
  3. Internal bigint + indexed public id
  4. Hashids are encoding, not a secret
  5. 404 vs 403 distinction leaks existence

basics

~20 s

Sequential ids are guessable, so they invite enumeration of your data and leak business volume and growth rates. The fix for access is authorization checks on every request; the fix for guessability is a separate opaque public identifier stored alongside the internal key. Obscured ids are not authorization.

solid answer

~60 s

Two distinct risks. First, **enumeration**: consecutive integers let anyone walk your id space, and if any endpoint fails to check ownership, that missing check becomes a bulk extraction rather than a single leaked record. Second, **information leakage**: id values disclose how many customers or orders you have, and two ids sampled a week apart disclose your growth rate. The non-negotiable control is authorization on every request - the object must be checked against the caller's permissions regardless of how the id looks. Unguessable ids reduce blast radius; they never replace that check. Given that, the practical pattern is a second column: keep the narrow internal key for joins and foreign keys, and add an indexed public identifier - a UUID, or a random token - that is the only value that ever leaves the system. It also decouples the external contract from storage, so you can change the internal key without breaking published URLs. Note that a time-ordered UUID still reveals creation time, and reversible obfuscation such as id-hashing libraries is encoding, not a secret.

go deeper

for a junior

Say that sequential ids can be guessed and walked, and that the server must check the caller is allowed to see each object regardless of the id.

for a middle

Separate enumeration from information leakage, and propose an internal integer key plus an indexed public identifier that is the only value exposed.

for a senior

Add blast-radius reasoning, uniform error responses, rate limiting, where identifiers leak besides the API, and the cost profile of each alternative including index locality.

for a principal

Treat the exposed identifier as a long-lived external contract; decide it before launch, define what may be inferred from it, and plan the transition path since published identifiers cannot be recalled.

## The two risks, kept separate **Enumeration and broken object-level authorization.** If `/orders/1041` works, someone will try `/orders/1042`. The vulnerability is never the id itself - it is an endpoint that returns the object without checking that the caller may see it. But the id design decides the blast radius. With sequential integers, a single missing check becomes a script that extracts the entire table in minutes. With a 128-bit random identifier, the same missing check leaks only objects whose ids the attacker already knows. **Information disclosure.** Sequential identifiers publish counts. Sign up twice a week apart, subtract the two ids, and you have the customer growth rate; the highest order id you have seen approximates total orders. For some businesses this is genuinely sensitive competitive information, and it leaks with no vulnerability at all - just from normal, correct use of the product. A lesser third risk is coupling: once an internal key is published, it is in customer bookmarks, partner integrations and search indexes, so it becomes an external contract you cannot change without breaking people. ## What actually protects data Authorization on every object access, evaluated server-side from the authenticated principal, scoped by tenant or owner, on every endpoint including the ones nobody thinks about - exports, attachments, webhooks, PDF renderers, admin tooling. The most reliable form is structural: make the data-access layer require an owner or tenant scope, so a query without one does not compile or does not run, rather than relying on each handler to remember. Unguessable identifiers are defence in depth. They meaningfully reduce the damage of an inevitable missed check, and they are cheap. They are not a control you can claim as protection. ## The alternatives **Expose the primary key directly.** Simplest, and acceptable for objects where neither enumeration nor counting matters - reference data, public content, small internal tools. Be honest that you have accepted both risks. **Random UUIDv4 as the primary key.** Removes enumerability at the source, at the index-locality and width cost that random keys carry on hot tables. **Separate public identifier column.** Keep a `bigint` primary key for joins and foreign keys, add an indexed unique `public_id`. Reads by public id cost one extra index lookup; everything internal keeps the narrow key. This is usually the best balance and it also decouples the published contract from storage. Choose UUIDv4 if you do not want to leak timing; UUIDv7 if roughly sortable public ids are useful and creation time is not sensitive. **Per-resource unguessable tokens.** For share links and similar capabilities, the identifier *is* the permission, so it must be high-entropy, revocable, expiring, and treated as a secret in logs and referrer headers. This is a different design from an identifier that merely names an object. **Reversible obfuscation.** Libraries that turn integers into short strings are encoding with a public algorithm and a non-secret salt. They stop casual incrementing and nothing more. They also make ids stable-looking while remaining trivially reversible if the salt leaks, so do not present them as a security control. **Composite scoped identifiers.** Exposing `/tenants/acme/invoices/2024-0007` keeps ids meaningful and forces the tenant scope into the route, which pairs well with structurally enforced tenant filtering. The number is still sequential within a tenant, which may be exactly what customers want on an invoice, so scope your leakage assessment accordingly. ## Practical details that get missed - **The identifier must be looked up efficiently**: index the public identifier uniquely, and beware that a random public id has the same index-locality characteristics as a random primary key, though usually on a much less hot index. - **Error responses leak too.** Returning 404 for "not yours" and 403 for "exists but forbidden" tells an attacker which ids exist. Prefer a uniform response for both. - **Timing and rate limits.** Enumeration is a volume activity; rate limiting and anomaly detection on object reads matter as much as id design. - **Do not publish internal ids in secondary places** - error messages, ETags, filenames, exported CSVs and email subjects have all leaked internal keys after the API was cleaned up. - **Migration is expensive.** Adding a public identifier later means backfilling, dual-reading during a transition, and honouring old URLs indefinitely. Deciding early is much cheaper than deciding correctly later. ## What a strong answer sounds like Separate the concerns out loud: authorization prevents unauthorized access; identifier design limits blast radius and prevents counting; a distinct public identifier decouples the external contract from internal storage. Then pick: internal `bigint` plus an indexed random public identifier for anything customer-facing, direct exposure for non-sensitive reference data, and capability tokens for share links.

  • If every endpoint checks authorization correctly, is there still a reason not to expose sequential ids?
    Yes - information disclosure. Sequential values reveal how many rows exist and, sampled over time, the rate at which they are created, which can be commercially sensitive with no vulnerability involved. Perfect authorization also depends on every current and future endpoint being correct, so unguessable ids remain worthwhile as blast-radius reduction against the check that eventually gets missed.
  • Are libraries that encode integer ids into short opaque-looking strings a security measure?
    No. They are reversible encodings with a published algorithm and a salt that is not a cryptographic key, so anyone who obtains or infers the salt can decode and re-encode ids at will. They raise the effort of casual incrementing and are fine as a cosmetic or contract-decoupling choice, but they must never be described as preventing enumeration or as a substitute for authorization.
  • You need to add public identifiers to a live table that currently exposes its integer primary key. How do you roll that out?
    Add the new column nullable with a unique index, backfill in batches to avoid a long lock, then make it NOT NULL and generate it by default for new rows. Accept both identifier forms on read during a transition period while emitting only the new one, so existing bookmarks, partner integrations and search results keep working. Keep the old form resolvable more or less indefinitely, since you cannot recall URLs you have already published.

saying these in an interview costs you the question

  • Describing unguessable identifiers as a replacement for authorization checks
  • Presenting reversible id-obfuscation libraries as a security control
  • Assuming random ids are unnecessary because 'the API requires a token' - tokens still belong to some user with limited rights
  • Returning distinguishable 404 and 403 responses that reveal which ids exist
  • Forgetting internal ids leak through error messages, filenames, exports and ETags

context