skip to content

You own a JPA domain model where a core entity hierarchy is about to grow from three subtypes to a dozen, several of which add many mandatory columns. How would you decide which @Inheritance strategy the hierarchy should use, and what would you weigh beyond raw query speed?

level: principalimportance: nice to knowfreq 28%

answer

  1. Four axes: read profile, constraints, storage/ops, change cost
  2. SINGLE_TABLE: no join growth, no NOT NULL on subtype columns
  3. JOINED: constraints kept, each subtype taxes every base query
  4. CHECK conditioned on discriminator = SINGLE_TABLE integrity patch
  5. Twelve subtypes = ask whether it is one hierarchy at all

basics

~20 s

Weigh read shape against constraint enforcement and change cost. SINGLE_TABLE is fastest but forfeits NOT NULL on subclass columns and grows sparse; JOINED keeps constraints and normalization at a join and INSERT per level, and every new subtype taxes existing polymorphic queries. Also ask whether the hierarchy should exist at all.

solid answer

~1 min

I would drive the decision from four axes, not one. 1. **How the data is read.** If most queries are polymorphic list screens over base columns, SINGLE_TABLE wins outright — no joins, no growth in cost per subtype. If most reads are of one concrete subtype, JOINED's single-branch join is cheap and the polymorphic case is rare. 2. **Whether the database must enforce mandatory-ness.** A dozen subtypes with many mandatory columns is the strongest argument for JOINED: under SINGLE_TABLE every one of those columns must be nullable and the constraint moves into application code or into CHECK constraints conditioned on the discriminator, which are easy to forget. 3. **Change cost per new subtype.** Under SINGLE_TABLE each subtype widens the shared table; under JOINED each adds a join to every existing base-type query. Both compound — you are choosing which way. 4. **Whether the hierarchy is real.** A dozen subtypes each with many unique mandatory columns often means these are distinct aggregates that share only technical fields; `@MappedSuperclass` or separate entities may be the honest model. My default: SINGLE_TABLE for shallow, behaviourally-uniform hierarchies; JOINED when subtypes carry substantial mandatory state — and I would push hard on whether all twelve belong in one hierarchy.

code

sql · 6 lines
sql
ALTER TABLE payment ADD CONSTRAINT card_number_required
  CHECK (payment_type <> 'CARD' OR card_number IS NOT NULL);

ALTER TABLE payment ADD CONSTRAINT iban_required
  CHECK (payment_type <> 'TRANSFER' OR iban IS NOT NULL);
-- one per mandatory subtype column; easy to forget when subtype #13 arrives

go deeper

for a junior

Recall the headline trade: one wide nullable table and fast reads, versus normalized tables with joins and real constraints.

for a middle

Lay out read profile, constraints and write cost, and pick a strategy with a stated reason rather than a default.

for a senior

Add operational realities — ALTER on a hot wide table, per-level inserts, CHECK-constraint upkeep — and show how you would measure the actual query mix before deciding.

for a principal

Lead with whether the hierarchy should exist, treat reversibility and per-subtype change cost as first-class, and leave behind a written rule for how the thirteenth subtype gets added.

## Why this is a judgment question There is no strategy that is right in general. The three JPA strategies trade **query cost**, **constraint enforcement**, **storage shape** and **change cost** against one another, and a dozen subtypes with heavy mandatory state sits exactly where the trade flips. A strong answer names the axes before naming a strategy. ## Axis 1 — the read profile What do the actual queries look like? - Mostly polymorphic reads of base columns (a feed, an audit list, a search screen): **SINGLE_TABLE**. One table, one index lookup, no join growth as subtypes multiply. This is the only strategy whose polymorphic cost does not degrade with hierarchy width. - Mostly reads of a known concrete type: **JOINED** costs one join per level, which is cheap and index-driven, and the rare polymorphic query can be handled with `TYPE()` pruning or DTO projections. - Essentially never polymorphic: this is a hint that the hierarchy may not be earning its keep at all (see Axis 4). Measure rather than assume: log the queries the application actually issues and classify them. ## Axis 2 — constraint enforcement This is the axis candidates most often miss and the one the scenario emphasises. Under SINGLE_TABLE, a column belonging to one subtype must be NULL for rows of every sibling, so it **cannot** be `NOT NULL`. With one or two optional extras that is a shrug; with a dozen subtypes each adding several mandatory fields, you have moved a large amount of integrity out of the database. The recovery options are real but imperfect: - Bean Validation (`@NotNull`) — enforced only by paths that go through the application. - `CHECK (payment_type <> 'CARD' OR card_number IS NOT NULL)` per column — enforced by the database, but verbose, easy to omit for a new subtype, and awkward to migrate. If the data is written by more than one system, or if it is financially or legally sensitive, database-enforced constraints usually decide the question in favour of **JOINED**. ## Axis 3 — storage and operational shape - SINGLE_TABLE with twelve subtypes becomes a very wide, sparse table. Row width affects how many rows fit per page and therefore scan efficiency; NULLs are cheap to store in most engines but the width is not free. Adding a column is an `ALTER TABLE` on the hierarchy's hottest table, which on very large tables is an operational event. - JOINED spreads writes across levels: an insert of a leaf is N statements, a delete likewise, and a bulk load pays that multiplier. Foreign keys can point at the base table, which is often a genuine requirement. - TABLE_PER_CLASS is rarely the right answer at this size: base-type queries become a twelve-way `UNION ALL`, identifiers must be unique across the hierarchy (no `IDENTITY`), and nothing can hold a foreign key to the base type. ## Axis 4 — is this really one hierarchy? A dozen subtypes, each with many unique mandatory columns and little shared behaviour, is the classic signature of a hierarchy created for schema reuse rather than for polymorphism. Ask: - Does any code hold a variable of the base type and behave uniformly on it? - Does any association reference the base type? - Do any queries genuinely need the mixed result? If the answers are no, then `@MappedSuperclass` (shared columns, no shared type, no join or width cost) or plain independent entities is the better model, and the strategy question evaporates. If the answers are yes for only three of the twelve, consider a narrower hierarchy with the others outside it. A fifth possibility deserves a mention: model the variable part as a mapped **JSON column** or a side table of attributes when subtypes are numerous, sparsely used and evolve constantly. It sacrifices per-attribute constraints and typed queries, so it is a deliberate trade, not a default. ## Axis 5 — reversibility Every one of these is a schema migration to change later. SINGLE_TABLE → JOINED means extracting N tables and backfilling; JOINED → SINGLE_TABLE means merging tables and relaxing constraints. Neither is catastrophic, both are a project. So weigh which mistake is cheaper to correct: widening a shared table is easier to live with than discovering you cannot enforce integrity on regulated data. ## How I would decide, concretely Given "a dozen subtypes, several adding many mandatory columns": 1. Split the hierarchy first — keep in it only the subtypes that are genuinely used polymorphically, and move the rest to `@MappedSuperclass`-backed independent entities. 2. For what remains, if the mandatory columns are substantial and the polymorphic reads are a minority, choose **JOINED**, and manage its cost with concrete-type queries, `TYPE()` pruning and DTO projections for the list screens. 3. If the remaining hierarchy is shallow and dominated by polymorphic reads, choose **SINGLE_TABLE** and pay for integrity with discriminator-conditioned CHECK constraints, written as part of the same migration that adds each subtype so they cannot be forgotten. 4. Either way, write down the rule for adding subtype number thirteen, because the compounding cost is what actually hurts. The answer an interviewer is listening for is that you know SINGLE_TABLE trades constraints for speed, JOINED trades speed for constraints, TABLE_PER_CLASS is rarely right at scale, and that the best move is often to make the hierarchy smaller.

  • How do you keep integrity if you pick SINGLE_TABLE for a hierarchy with many mandatory subtype columns?
    Bean Validation covers application writes but nothing else, so the durable answer is CHECK constraints conditioned on the discriminator — one per mandatory column, asserting that the column is non-null whenever the row is of that subtype. The risk is procedural: they must be added in the same migration that introduces each subtype, or they quietly never appear. If that discipline is unrealistic for a dozen subtypes, that is itself an argument for JOINED.
  • Why is TABLE_PER_CLASS rarely the right choice for a twelve-subtype hierarchy?
    Base-type queries become a twelve-way UNION ALL with NULL padding, so polymorphic cost scales with hierarchy width just as JOINED does, without JOINED's normalization benefits. Identifiers must also be unique across the whole hierarchy, ruling out IDENTITY generation, and no foreign key can reference the base type because it has no table. It fits only isolated, never-queried-together types — which usually means @MappedSuperclass fits better.
  • What signals tell you the hierarchy itself is the wrong model?
    No variable, parameter, collection or association in the codebase is typed as the base class; no query genuinely needs mixed results; and each subtype adds many unique mandatory columns while sharing almost no behaviour. Those together mean the base exists for schema reuse, which @MappedSuperclass provides without polymorphic cost or a shared schema decision.

SINGLE_TABLE is one enormous form everyone fills in, with most boxes blank and no way to mark a box required-for-some-people. JOINED is a short common form plus a mandatory type-specific annex — stricter, but the clerk must fetch every annex when asked for 'any submission'.

saying these in an interview costs you the question

  • Answering with a single strategy and no axes — "JOINED is normalized so it is correct"
  • Overlooking that SINGLE_TABLE forfeits NOT NULL on subtype columns
  • Treating the choice as purely a read-performance question
  • Proposing TABLE_PER_CLASS at this size without mentioning UNION cost, identifier uniqueness or the missing base-table foreign key
  • Never questioning whether all twelve subtypes belong in one hierarchy

context