skip to content

Some languages give certain types no stable reference identity at all: structs are copied on assignment, or values are moved. How do you model an entity whose identity must persist in such a language, and what changes once the entity crosses a process or storage boundary?

level: principalimportance: should knowfreq 33%

answer

  1. Address identity, value equality, domain identity are three different questions
  2. Go copies structs; Swift gives === only to classes; Rust moves and makes Rc::ptr_eq opt-in
  3. Free identity in Java, C#, Python is process-scoped and CPython reuses id()
  4. Entities compare by key, values by content
  5. Identity map in Hibernate, EF Core, SQLAlchemy; unsaved-entity trap

basics

~20 s

Reify identity as a field. Go copies structs on assignment, Swift gives === only to classes, Rust moves values and offers Rc::ptr_eq only on request. No address survives serialization anywhere, so entities carry an explicit key and compare by it.

solid answer

~50 s

Three kinds of identity exist: address identity, value equality, and domain identity (a key you can write down). Only the third survives. - **Go**: `b := a` copies the struct, and `==` compares field by field, so two distinct entities with equal fields compare equal and one entity in two variables is two addresses. Identity must be a field. - **Swift** forces the modelling decision at declaration: `===` exists only for classes, so choosing struct versus class *is* choosing value versus entity. That is a rare case of the language making you say which you meant. - **Rust** moves values and gives no universal identity; `Rc::ptr_eq` is opt-in, so sharing is explicit in the type. - **Java, C#, Python, Ruby** hand every object a free address-flavoured identity, which is why it gets used by accident. CPython even reuses `id()` values after collection. ORMs make the cost visible: Hibernate, EF Core and SQLAlchemy all need an identity map because one row loaded twice is two objects.

code

go · 9 lines
go
type User struct { ID int; Name string }

a := User{ID: 7, Name: "ann"}
b := a            // full copy, different storage
b == a            // true - field-by-field, though these are two values

c := User{ID: 7, Name: "ann"}   // a different record entirely?
c == a            // also true - structural comparison cannot tell
// idiomatic answer: compare a.ID == c.ID and treat ID as the identity

go deeper

for a junior

Know the distinction between 'same object' and 'same content', and that an entity such as a user is normally identified by an ID field rather than by either of those.

for a middle

Be able to explain that some languages copy or move values so no stable address exists, and that entity equality should therefore compare a key while value equality compares content.

for a senior

Discuss the boundary: serialization drops addresses, ORM identity maps only hold within a session, and unsaved entities have no stable key yet. Name the mitigations.

for a principal

Own the modelling call: decide entity versus value per type, pick natural or surrogate keys and where they are generated (client-side UUID versus database sequence) so identity exists before persistence, and state the consequences for deduplication, caching and cross-service references.

## Three identities, only one portable Separate three questions that are easy to conflate. **Address identity**: do these two expressions denote the same storage right now, in this process? **Value equality**: is the content equivalent? **Domain identity**: do these two representations stand for the same thing in the world, an order, a user, a device? Only domain identity is a property of the model. The other two are properties of a runtime, and both evaporate the moment a value is copied, moved, serialized, cached in another process or reconstructed from a database row. ## Languages that hand out address identity for free Java, C#, Python, Ruby and JavaScript give every object a distinguishable reference. Because it is free, it gets used unintentionally: a set of objects deduplicates by reference until someone defines otherwise, and two loads of the same database row produce two members. Even where it exists it is narrower than people assume. CPython's `id()` is an address that is reused after the object is collected, so two objects with equal ids can exist at different times in one run. On a compacting collector the address itself moves, which is why identity hash values are stored rather than derived from location. ## Languages where it is absent or deliberately withheld **Go**: assignment copies a struct. After `b := a` there are two independent values, and `==` on comparable structs compares field by field, so two genuinely different entities that happen to share field values compare equal, while one entity held in two variables occupies two addresses. Neither address nor structural comparison answers "same entity?", so idiomatic Go carries an ID field and compares that. **Swift**: `===` is defined for class references only. Structs have value semantics and no identity operator, so the choice between `struct` and `class` at declaration time *is* the choice between value and entity. Few languages force the modeller to make that call explicitly; Swift's guidance to prefer structs unless you need identity is exactly this distinction stated as a rule. **Rust**: values move, the compiler may relocate them, and there is no universal identity accessor. Comparing shared ownership requires `Rc::ptr_eq`, which you can only call if you already chose a shared pointer type. Identity is opt-in and visible in the type signature rather than ambient. **C and C++**: an object has an address while it lives, but copies, moves and reallocation invalidate any assumption that the address is the thing. A `std::vector` reallocation relocates every element. ## Modelling identity when the language will not The answer is the same everywhere and is the reason it is worth learning: **reify identity as data**. Give the entity a key: a natural key drawn from the domain (ISBN, IATA code) or a surrogate one (UUID, ULID, database sequence). Compare entities by that key alone, and compare values by content. This yields three concrete rules. First, an entity's equality must not depend on mutable content, or the same entity stops equalling itself after an edit. Second, a value type's equality must not depend on a key, or two equal values become distinct for no domain reason. Third, an entity with no key yet assigned has no stable identity, which is the classic transient-instance trap: Hibernate, EF Core and SQLAlchemy all let you place an unsaved object into a hash-based set before the key exists, after which assigning the key silently changes what the object claims to be. The standard mitigations are to assign identity client-side (UUID at construction, so it exists before persistence) or to keep unsaved entities out of key-based containers entirely. ## The boundary Once an entity crosses a boundary, address identity is simply not in the payload. Serialization writes fields; the receiver reconstructs a new object. This is why every mature ORM maintains an **identity map**: within one session, one row maps to one object, so reference comparison happens to work; across sessions it does not, and code that relied on reference comparison breaks in exactly the situation you cannot test locally. It is also why distributed systems generate identity that is meaningful without coordination, UUIDv4, UUIDv7 or ULID for time-ordered locality, rather than relying on a central sequence. A single systems note, worth one sentence: content-addressed stores invert the relation deliberately, deriving identity from a hash of the value, which is sound only because the value is immutable, and which by design makes two independently produced identical values the same entity. ## The judgement calls Deciding entity versus value is the primary modelling act, and languages differ in whether they make you state it. A money amount, a coordinate and a date range are values: equal content means equal, and copying is harmless. A user, an order and a device are entities: they have a life, they change, and two snapshots at different times are the same thing. Getting this wrong is expensive in both directions, an entity treated as a value silently merges distinct records during deduplication, and a value treated as an entity fragments caches and prevents sharing. Ask, for each type, "if every field changed, would this still be the same thing?" A yes means you need a key, and the key must be data.

  • Why do ORMs such as Hibernate, EF Core and SQLAlchemy maintain an identity map, and what breaks without one?
    The identity map guarantees that within one session a given row is represented by exactly one object, so reference comparison and object caches behave. Without it, two loads of the same row produce two objects, and a set of loaded entities silently contains duplicates. Across sessions the guarantee is gone entirely, which is why entity equality must be based on a key rather than a reference.
  • What is the transient-instance problem, and how do teams avoid it?
    An entity whose key is assigned by the database has no stable identity before it is saved, so putting it into a key-based container beforehand means its identity changes underneath the container. Teams avoid it either by generating identity client-side at construction, typically a UUID, or by keeping unsaved entities out of hash-based collections. Basing equality on mutable business fields instead just relocates the same failure.
  • When should identity be derived from the value instead of assigned?
    Only for immutable values, which is the premise of content addressing: the key is a hash of the content, so two independently created identical values are deliberately the same entity. It is the right choice for artifacts, blobs and configuration snapshots, and the wrong choice for anything that changes, since a mutation would change the identity of the thing that is supposed to persist through change.

An address identity is a seat number: useful inside one theatre for one performance, meaningless once the audience goes home. A domain key is the name on the ticket, which still identifies the person tomorrow, in another city, on paper.

saying these in an interview costs you the question

  • Assuming every object in every language has a reference identity available to compare
  • Using reference comparison for domain identity and discovering the gap only when objects cross a session or process boundary
  • Basing entity equality on mutable business fields, so an edit makes the entity stop equalling itself
  • Treating a Go struct copy as a reference to the same entity, or expecting Swift's === to work on structs
  • Treating CPython's id() as globally unique for the life of the program, when values are reused after collection

context