When a data-access layer maps a class hierarchy, how does it decide which concrete subclass to instantiate for a row?
answer
- a row does not carry its class
- two mechanisms, one is free
- marker column read with the row
- or probe which child table has the key
- marker values are data, not code
basics
~20 sEither a type-marker column stored with the row names the class, or the layer discovers which subtype table holds a row for that key. The marker is free with the read; discovery costs a join or a probe per subtype.
solid answer
~50 sA row carries values, not a class, so a mapped hierarchy needs one extra decision before hydration. The common mechanism is a **type marker**: an extra column holding a code that the layer looks up in a registry built from the mapping metadata at startup. It is read in the same statement as the rest of the row, so deciding is free, and narrowing a query to some subtypes becomes a plain predicate on that column. The alternative, used when no marker is stored, is to infer the class from which subtype table holds a row for the key — an outer join to every subtype table, or a probe of each. That grows with the hierarchy: one more subtype is one more join in every base-type read. Whichever mechanism is used, the class is fixed at hydration and cannot be changed afterwards without reloading the object.
go deeper
Recall the two mechanisms and which is cheap: a marker column comes along with the row for free, discovery by probing subtype tables does not.
Explain how each mechanism behaves when a subtype is added, and how each one narrows a query to a subset of subtypes.
Show the operational consequences: marker values are stored data shared by every reader, so renames, deploy order and drift between code and data all become your problem.
Frame the marker set as part of the schema's published contract, and weigh that coupling against the flexibility a hierarchy is supposed to buy.
## What the mapper has to decide A row is only values. Nothing in it says the record is a card payment rather than a bank transfer; the class is a fact the **mapping metadata** adds on top of the data. So for any mapped type that has subclasses, every read gains a step before hydration: **choose the concrete class**, then fill an instance of that class. The choice is made once, at hydration, and everything afterwards follows it — type tests, virtual dispatch, narrowing, even how the object serializes. Two mechanisms cover nearly all of what data-access layers do here, and the table layout decides which one is available. Some layers additionally let application code register its own resolver, but that is the exception rather than the norm. ## Mechanism one: a type marker stored with the row An extra column holds a value naming the concrete class — a short code, a small integer, or the class name itself. At startup the layer builds a registry from the mapping metadata, and each read looks the row's marker up in that registry. - Deciding costs **nothing extra**: the marker rides along in a SELECT the layer was issuing anyway. - It works whether the subtype's own columns sit in the same wide table or in a joined child table. In the joined case the marker also tells the layer which child table to expect, so it can skip the rest. - Narrowing a query to a few subtypes becomes an ordinary predicate on the marker column, which an index can serve: `... WHERE kind IN ('CARD', 'VOUCHER')`. - Marker values are **data**. They outlive the code that wrote them, so a marker derived from a class name turns a rename or a package move into a data migration. - A marker value the running build cannot resolve is a hard failure at hydration, not a quietly skipped row. ## Mechanism two: discovering which subtype row exists When no marker is stored, the concrete class is inferred from **where a matching row lives**: the base key appears in exactly one subtype table, and the layer finds out either by outer-joining every subtype table in a single statement and seeing which side came back non-null, or by probing the subtype tables until one answers. - The decision now costs joins or extra round trips, and the cost is proportional to the size of the hierarchy: one more subtype means one more outer join in every base-type read. - Correctness rests on an invariant no single row enforces — exactly one child row per base key. A base key with no child row, or with two, is data the layer can only report as broken. - Reading a **known** concrete type is unaffected: if the caller asks for one subtype by name, the layer goes straight to that subtype's table (plus the base table when the shared columns live there). ## The two side by side | | type marker column | probing subtype rows | |---|---|---| | Cost of the decision | one more column in a read already happening | an outer join or probe per subtype | | Cost of adding a subtype | a new marker value | a wider statement in every base-type read | | Restricting to some subtypes | a predicate on the marker | choosing which tables to join | | Correctness depends on | the marker matching the class registry | exactly one child row per base key | | Typical failure | an unknown marker value | a base row with no child row | ## Consequences worth carrying into review 1. **The decision is per row, not per query.** One base-type result set can mix classes freely; each row is instantiated according to its own marker or its own child row. 2. **It happens before your code sees anything.** By the time results are handed over the classes are fixed, so a wrongly typed object cannot be re-resolved in place — it has to be evicted and read again. 3. **Markers are a deployment coupling.** Every instance that reads the table must know every marker value written to it, which constrains rolling deploys and makes the marker set part of the schema's published contract. 4. **Neither mechanism escapes the hierarchy.** Both make the shape of a base-type read depend on how many subtypes exist; the marker makes that dependency cheap, probing makes it structural. A practical default: use an explicit, stable marker code that is not controlled by the class name, index it if base-type queries are ever narrowed, and treat adding a subtype as a change that touches deployment order as much as it touches schema.
- Why is a marker value derived from the class name riskier than an explicit short code?The marker is stored data that outlives the build that wrote it. If the value is the class name, renaming or moving the class makes every existing row unreadable until they are rewritten, so a code refactor silently becomes a data migration. An explicit stable code decouples the two.
- Where the class is inferred from which subtype table holds the key, what invariant must the data uphold?Exactly one child row per base key. Nothing in a single row enforces it, so a base key with no child row, or with two, is unresolvable: the layer cannot pick a class and can only report the data as broken. A marker column has no equivalent structural invariant.
- Does the decision happen per query or per row?Per row. One base-type result set can mix classes freely, each row instantiated according to its own marker or its own child row. That is why a base-type read can hand back several different concrete classes from a single statement.
A parcel with no label has to be identified by opening every shelf to find its slot; a label printed on the box tells you in one glance. The type marker is the printed label.
saying these in an interview costs you the question
- Thinks the layer infers the class from the fields that are non-null
- Believes a type marker costs an extra query per row
- Treats marker values as code details that can be renamed freely
- Assumes every mapped hierarchy stores a marker column
- Thinks a wrongly typed loaded object can be re-resolved in place