skip to content

Why does a global object id pack the type name with the local key, and stay opaque?

level: middleimportance: should knowfreq 47%

answer

  1. One field resolves every type
  2. Two rows can share a number
  3. Uniqueness and refetchability are the requirements
  4. Encoding is not encryption
  5. Opacity is a contract, not a lock

basics

~20 s

Because one lookup field has to resolve identifiers for every type, the value must say what it points at as well as which row. Opacity keeps the encoding the server's to change, since no caller is entitled to read it.

solid answer

~50 s

The **Relay server specification** requires two properties of the value: it is unique across the whole schema, and the server can refetch the object from it. It does not prescribe an encoding. The near-universal convention is to base64 the string `"TypeName:localKey"` — for instance `"SolarPanel:8412"` — because the type tag is what lets one polymorphic lookup field decide which backend to ask, and what stops two tables that both number a row 8412 from colliding in one identifier space. Opacity is the other half. base64 is an encoding, not encryption: anything inside is one decode away, so a sequential primary key inside an identifier is still a sequential primary key. Opacity is a *contract*, not a protection — callers must neither parse an identifier nor construct one, and the moment a client does, the server has frozen its encoding and cannot change it without breaking that client.

code

pseudocode · 17 lines
pseudocode
function encodeGlobalId(typeName, localKey):
    return base64(typeName + ":" + localKey)

function decodeGlobalId(value):
    raw = base64Decode(value)          # may fail -> not found
    if raw is null: return NOT_FOUND

    parts = splitOnFirst(raw, ":")
    if parts.count != 2: return NOT_FOUND

    typeName = parts[0]
    localKey = parts[1]

    if not isKnownNodeType(typeName): return NOT_FOUND
    if not matchesKeyFormat(typeName, localKey): return NOT_FOUND

    return (typeName, localKey)        # caller dispatches on typeName

go deeper

for a junior

Know that the identifier is a single opaque string containing more than a row number, that you read it from a response and pass it back unchanged, and that you never build one yourself.

for a middle

Explain both halves: the type tag lets one polymorphic lookup dispatch and prevents collisions across types, and opacity is what preserves the server's freedom to change the encoding later.

for a senior

Be ready to describe the failure a missing type tag actually causes — a lookup keyed on the bare local key returning another type's row — and to treat decoding as an input-validation path that fails closed.

for a principal

Own the encoding as a long-lived contract: what goes in, whether the local key is guessable, who is allowed to depend on the format, and how you would notice a caller that started parsing it.

## What the specification actually asks for Two things, and only two. The value must be **unique across the whole schema**, and the server must be able to **refetch the object** from it alone. Everything else — base64, the colon, the type name in front — is a widespread convention that grew up around the specification, not a rule inside it. Saying "the spec requires base64 of TypeName:id" is a small mistake with a large tell: it means the candidate learned the convention from one library's output rather than from the contract. ## Why the type name is inside A single lookup field has to resolve identifiers for every implementing type. The identifier is the only input it gets. So the identifier has to answer two questions at once: *what kind of thing is this*, and *which one*. In a solar-array telemetry graph, panels and inverters are separate tables and both number rows from one. Panel 8412 and inverter 8412 both exist. If the identifier were the bare local key, three things break in ascending order of nastiness: 1. The lookup cannot dispatch. There is no way to know whether to ask the panel store or the inverter store. 2. The identifier stops being globally unique, so anything keyed on it — a store, a log line, an audit row — conflates two distinct objects. 3. Worst, it can silently succeed with the wrong object. On one telemetry service a shared per-request lookup cache was keyed on the bare local key rather than the whole identifier; at a 1,200-request-per-minute peak an inverter lookup warmed key 8412 and the next panel lookup in the same batch was handed that cached row straight back. Nothing errored. The dashboard rendered an inverter's firmware string under a panel's name, and the bug survived review because every individual resolver was correct. The type tag is what makes that class of collision impossible rather than merely unlikely. ## Why the local key is inside The other half is the object's own key within its type — whatever the backing store uses to find one row. It must be **stable for the lifetime of the object**. That rules out anything a user or an operator can edit: a serial number that gets re-stamped after a repair, a slug derived from a name, a position in a list. If the local key can change, the identifier can change, and an identifier that changes is not an identifier. ## What opacity means, and what it does not Opaque means *the caller is not entitled to look inside*. It does not mean the caller cannot. base64 is an encoding; a decode is one call away, and any competent caller who wants to read your identifiers will. So opacity buys exactly one thing, and it is worth a lot: **freedom to change the encoding later**. As long as no caller parses identifiers, the server can switch encodings, add a version marker, salt the key or move a type to a different store, and no caller notices. The day one client library starts splitting on the colon to recover the row number, that freedom is gone — not by policy but by fact, because the next change breaks that client. The symmetric rule binds the other direction: a caller must never **construct** an identifier. Client code that base64s `"SolarPanel:" + serial` to save a round trip has hard-coded your internal format into someone else's release cycle. If a caller needs to reach an object from a natural key, give it a root field that takes that natural key. ## What leaks, and what to do about it Since the contents are readable, treat them as published. A sequential integer inside an identifier tells a caller roughly how many rows you have and lets them walk the neighbours by incrementing and re-encoding. Using a random or opaque local key — a UUID, a hashed key — removes that inference. It is a defence against *enumeration by guessing*, and nothing more: what a caller is allowed to see once it holds a valid identifier is a separate concern with its own answer. ## Decoding as untrusted input An identifier arriving on the wire is caller-supplied bytes, so the decode path is an input-validation path. It should fail closed when the value is not valid base64, when it does not split into the expected parts, when the type tag is not a known type implementing the interface, or when the local key does not match that type's key format. And the failure should surface as an unmatched lookup, not as an error message quoting the decoded parts — an error that helpfully reports `expected "Type:key", got "8412"` has just documented your encoding for anyone probing. ## The short version for an interview Type name plus local key, so one polymorphic lookup can dispatch and identifiers cannot collide. Opaque, so the encoding stays yours to change. base64 by convention, not by specification, and never mistaken for secrecy.

  • Is base64 of "TypeName:localKey" required by the Relay server specification?
    No. The specification asks that the value be unique across the schema and sufficient for the server to refetch the object, and treats it as opaque to callers. base64 of `"TypeName:localKey"` is a very widespread convention that grew up around it, largely because several toolchains emit exactly that. Any scheme meeting the two properties is compliant, including a random string mapped to the object in a table.
  • How should a server decode an identifier that a caller may have tampered with?
    As untrusted input, failing closed at every step: not valid base64, does not split into the expected parts, the type tag is not a known type implementing the interface, or the local key does not fit that type's format. Surface all of those as an unmatched lookup — null — rather than as an error quoting the decoded fragments, which simply documents your encoding for whoever is probing it.
  • What is the practical downside of putting a sequential primary key inside the identifier?
    It is one decode away from being read. A sequential key discloses roughly how many rows exist and lets a caller enumerate neighbours by incrementing and re-encoding. A random or hashed local key removes that inference cheaply. It only closes guessing; what a caller may see once it holds a genuine identifier is a separate question with a separate answer.

Packing the type name in is like printing the ward on a hospital wristband as well as the bed number: bed 12 exists on every ward, and the bed number alone sends a nurse to the wrong floor.

saying these in an interview costs you the question

  • Treats base64 in an identifier as a security measure
  • Parses identifiers client-side to recover the row key
  • Uses the bare database primary key as the global identifier
  • Builds identifiers in client code instead of reading them
  • Derives the local key from an editable business value
  • Says the specification mandates the base64 type-prefix format

context