What does a mapper demand of a composite key class used as a mapped object's identifier?
answer
- the key is a value, not a field
- equality and hash over every part
- no part may be empty
- frozen once the row exists
- every link carries all parts
basics
~20 sThat it behave as a value: all parts populated, equality and hashing over every part, and no mutation once the row exists. The layer files objects under that whole value, so a changed key loses the row it addressed.
solid answer
~40 sA multi-column key stops being a scalar field and becomes a mapped type of its own, and the layer leans on it much harder than on ordinary columns. It needs a construction path it can use while materialising a row; **value equality and a hash derived from every part**, because the identity map and the tracked set are keyed by that value and two loads of one row must produce equal, equally-hashing keys; and every part populated, since a missing part makes a lookup ambiguous. The value must also be stable once the row exists — changing a part does not rename the row, it points the object at a different one. Finally, every link that targets such a row has to carry all the parts, not one of them.
go deeper
Know that a multi-column identifier is expressed as a small class holding all the parts, and that the layer compares whole key values rather than individual columns.
Explain what the layer needs from that class: a way to build it from column values, equality and hashing over all parts, no empty parts, and no change after the row exists.
Show the live failures: a mutated part silently re-points an object at another row, an empty part turns a lookup into a miss, and a stale key sitting in a cache or a hash container is never found again.
Weigh what a multi-part identifier imposes system-wide — every link carries every part, every cache and message serialises the whole value, and any correction becomes a row move — and decide whether the schema, not the mapping, is what should change.
When a row is identified by more than one column, the mapped object cannot hold its identifier in a single scalar field. The key becomes a small **value type** — an object whose whole purpose is to carry the parts together — and the mapping layer places demands on it that it never places on ordinary mapped fields. ## Why the layer cares so much The identifier is the layer's addressing scheme. It is what the **identity map** files loaded objects under so that one row yields one instance; it is what the tracked set uses to match a loaded object against the statement it will emit; it is what a lookup, a link, a cache entry and a message payload all carry. Every one of those uses compares key values and most of them hash them. A scalar key gets those properties for free. A composite key only has them if you give it them. ## What the layer requires 1. **A construction path from column values.** The layer materialises the key while reading a row, so it needs a way to build the value from the raw parts without running application logic. 2. **Equality over every part.** Two key values built from the same row must compare equal. Reference equality is fatal here: each load builds a fresh key object, so an identity-based comparison files the same row twice. 3. **A hash derived from the same parts as equality.** Equal values must hash alike, or the identity map will hold duplicates that never find each other. 4. **No null parts.** A partially populated key makes both the comparison and the lookup undecidable, and the store cannot match a key column against nothing in the usual way. 5. **Stability after the write.** The key is the row's address. Once the row exists, the value must not change. ## Mutation is not a rename Changing a part of a stored object's key does not correct the row; it re-points the object at a different row, one that probably does not exist. The layer had filed the object under the old value, so lookups by the new value miss, and any statement computed from the object addresses the wrong row. Layers respond differently — some refuse the change, some emit a statement whose predicate matches nothing — but none of them treat it as a rename. A genuine correction to a key part is a new row plus the removal of the old one, with every referring row updated, which is exactly why teams treat key parts as frozen. ## Nulls, and what they really signal A key part that is sometimes absent is a modelling error dressed as an edge case. Comparison, hashing and lookup all need the whole value; an optional part means the row is really identified by a smaller set of columns, or by something else entirely. The right response is to shrink the key or replace it, not to teach the key class to tolerate a hole. ## Where the demands ripple outward | aspect | scalar key | composite key | |---|---|---| | the key in memory | a field | a value type with equality and hash | | a link to the row | one column | one column per part, on every referring row | | a cache or message key | the value itself | the whole value, written out and read back consistently | | ordering a page of results | one column | all parts, in an agreed order | The second column is the real cost of a multi-part identifier, and it is paid by every part of the system that refers to the row, not only by the class that owns it. ## Where layers differ Data-access layers differ in how the composite value is expressed: some want a separate key class held as a field, some let the parts stay on the object and be nominated as the key, and some accept both. The requirements above are the same in every case, because they come from what the layer does with the value — file it, match it, hash it, send it — not from how it is declared. ## What to say in an interview Say that the key becomes a value: constructible, compared and hashed on all parts, complete, and frozen after the write. Then say what it costs the rest of the model — every link carries every part — and note that this cost, not the class itself, is why multi-column identifiers are usually a decision about the schema rather than about the mapping.
- What happens if a part of a stored object's key is changed in memory?The object stops matching the row it came from. It was filed under the old value, so lookups by the new value miss and any statement built from it addresses a row that does not exist. Correcting a key part is a new row plus a removal, not an update in place.
- Why must every part of a composite key be populated?Because the layer must compare, hash and look up by the whole value, and a missing part makes all three undecidable. A part that is genuinely optional means the key includes a column it should not, so the fix is to change the key rather than to tolerate the hole.
- Does the key value need to survive being written out and read back?Usually yes. A key travels beyond the object — into a shared cache entry, into a message, into a page cursor — and must compare the same after the round trip as before it. That is another reason the parts should be simple, immutable values rather than rich objects.
saying these in an interview costs you the question
- Leaves the key class with reference equality
- Allows one part to be empty and expects lookups to match
- Edits a key part in place to correct a value
- Hashes only the most selective part of the key
- Assumes a link to such a row needs a single column