skip to content

Entity Mapping & Annotations

How a POJO becomes a row: identifiers and generation strategies, columns and converters, embeddables, enums and temporals, and inheritance hierarchies. Interviewers start here because mapping choices — IDENTITY vs SEQUENCE, SINGLE_TABLE vs JOINED — lock in performance behavior for the life of the schema.

part ofHibernateoverview, primer and where to startread it →
on this pageshow

explore

questions

page 2 of 2

What does the JPA @Lob annotation do to a String or byte[] entity attribute, and what problems do teams run into with it in production?

level: seniorimportance: should knowfreq 34%

basics

~20 s

@Lob tells the provider to map the attribute to a large-object type — CLOB for String, BLOB for byte[] — instead of varchar or varbinary. In production it costs memory: the whole value is loaded with the row and rewritten on update unless you split it out.

open as a page

Why does Hibernate insist that a composite identifier class implement equals() and hashCode(), and what concretely goes wrong at runtime when the implementation is missing, incomplete, or based on mutable state?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Hibernate keys the persistence context and the second-level cache by identifier value, comparing keys with equals/hashCode. Without a correct implementation, lookups miss: the same row loads as multiple instances, find() re-queries, merge misbehaves, and cache hits never happen. Keys must also be immutable, or the hash changes under the map.

open as a page

A link entity's primary key is (order_id, product_id) and the same entity also needs @ManyToOne references to Order and Product. How do you map that in JPA/Hibernate without mapping the same two columns twice?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Use derived identity: keep an @EmbeddedId holding the two id values and annotate each association with @MapsId("orderId") / @MapsId("productId"). Hibernate then owns the columns once through the associations and fills the key fields from them, so you set only the associations.

open as a page

What does the Hibernate configuration property hibernate.jdbc.time_zone do, why do teams set it to UTC, and what must you be careful about when introducing it into a system that already has data?

level: seniorimportance: should knowfreq 32%

basics

~20 s

It tells Hibernate which time zone to pass to JDBC when binding and reading timestamps, instead of the JVM default. Setting it to UTC makes stored values independent of the server's zone. Introducing it later re-interprets existing rows, so it needs a data migration, not just a config flip.

open as a page

A production table stores an enum attribute as ordinal integers and the team now needs to insert a new constant in the middle of the enum and move the column to string storage. How do you carry that out safely, and how would you detect damage if someone had already reordered the constants?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Freeze the current ordinal-to-name mapping from the old source, add a new string column, backfill it by mapping each stored ordinal through that frozen table in SQL, switch the mapping to EnumType.STRING, then drop the integer column. To detect prior damage, reconstruct the mapping from git history and cross-check row counts and timelines against business expectations.

open as a page

A field is annotated with Hibernate's @ColumnDefault("0") and the table was created by Hibernate's schema generation, yet new rows still arrive with NULL in that column. Why, and what actually makes the database default apply?

level: seniorimportance: should knowfreq 30%

basics

~20 s

@ColumnDefault only adds DEFAULT to generated DDL. Hibernate still lists every mapped column in the INSERT, so it sends an explicit NULL and an explicit value always beats a column default. Use @DynamicInsert so null columns are omitted, plus @Generated to read the value back — or just set it in Java.

open as a page

Explain why a Java equals() implementation that starts with getClass() != other.getClass(), or that reads the other object's fields directly, can misbehave when Hibernate hands you a lazily-loaded proxy — and what to write instead.

level: seniorimportance: should knowfreq 42%

basics

~20 s

A lazy reference is a generated subclass of your entity, so getClass() returns Order$HibernateProxy, not Order, and a getClass() check rejects the same row. Direct field access on a proxy reads the uninitialized subclass fields (null), bypassing the interception. Use instanceof and call getters.

open as a page

You assign entity identifiers in application code — for example a UUID set in the constructor — instead of letting the database generate them. What does that change for Hibernate's handling of the entity, and what is the storage and index cost of random UUIDs?

level: seniorimportance: should knowfreq 35%

basics

~20 s

An id that is never null removes Hibernate's usual new-versus-detached signal, so merge() must SELECT before it knows whether to insert. Random UUIDs are 16 bytes and insert at random points in the primary-key index, causing page splits and poor cache locality; time-ordered UUIDs and a native 16-byte column fix most of that.

open as a page

An application maps a deep entity hierarchy with JPA's JOINED inheritance strategy and its polymorphic queries have become slow. Explain the SQL these queries generate, what makes them expensive, and the levers available to reduce the cost without abandoning the hierarchy.

level: seniorimportance: should knowfreq 38%

basics

~20 s

A base-type query under JOINED outer-joins every subclass table so Hibernate can build each concrete instance, so cost grows with hierarchy width. Levers: query the concrete subtype instead, restrict with TYPE()/TREAT(), select a DTO projection of base columns only, index the join keys, or flatten to SINGLE_TABLE.

open as a page

One of your Hibernate entities maps to a table called order and has a column called user, both reserved words in the database. What options does Hibernate give you to make the generated SQL valid, and what is the risk of switching on the setting hibernate.globally_quoted_identifiers?

level: seniorimportance: should knowfreq 30%

basics

~10 s

Quote the offending names individually - backticks or escaped double quotes in @Table/@Column - or enable hibernate.auto_quote_keyword. hibernate.globally_quoted_identifiers quotes everything, which makes all identifiers case-sensitive and locks the schema to their exact spelling.

open as a page

How do you make one attribute of a JPA entity use property access while the rest of the class uses field access, and when is that worth doing?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Declare the class default explicitly with @Access(AccessType.FIELD), mark the backing field @Transient so it is not mapped twice, then put @Access(AccessType.PROPERTY) plus the mapping annotations on the getter. Hibernate then calls that getter and setter for that one column.

open as a page

Hibernate's sequence-backed id generators can use different optimizers — hi/lo, pooled, and pooled-lo. What does the value returned by the database sequence mean in each, and which are safe when another application inserts into the same table?

level: seniorimportance: nice to knowfreq 25%

basics

~20 s

With legacy hi/lo the sequence value is a multiplier, so it bears no relation to stored ids and outside writers collide. With pooled, nextval returns the top of the reserved block; with pooled-lo it returns the bottom. Both keep sequence values inside the id space, so external callers using the same sequence stay safe.

open as a page

For an association table entity such as order_line, how would you decide between a composite primary key built from the two foreign keys and a single surrogate key with a unique constraint on the pair? What drives the call, and what does each choice cost you later?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Composite keys model the truth (a pair is unique) with no extra column, but every child table, cache key and query carries both columns and the key must never change. A surrogate key gives one stable, cheap handle to reference and evolve, at the cost of an extra column plus a unique constraint you must not forget.

open as a page

You are setting the convention for how a multi-service system persists dates and times through JPA/Hibernate entities. How do you decide what each attribute stores — UTC instants, local wall-clock plus a zone, or plain calendar dates — and what do you standardise across the codebase?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Classify each attribute first: a past moment stores an Instant in UTC; a calendar fact stores LocalDate; a future local commitment stores wall-clock plus an explicit zone id so later zone-rule changes stay correct. Then standardise UTC everywhere in storage, conversion only at edges, and one mapping style per meaning.

open as a page

You are designing the identity model for a domain whose objects routinely live outside an open persistence context — cached DTO graphs, HTTP sessions, sets nested inside aggregates. How would you decide between assigning identifiers in the constructor (for example UUIDs), relying on a business/natural key, and using database-generated sequence identifiers as the basis for object equality?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Pick by whether an identifier exists at construction time. Assigned UUIDs always do, so equality is simple everywhere, at the cost of index locality — mitigate with time-ordered UUIDs. A true immutable business key is best when one genuinely exists. Generated sequence ids are cheapest in the database but need the constant-hashCode workaround.

open as a page

You own a JPA domain model where a core entity hierarchy is about to grow from three subtypes to a dozen, several of which add many mandatory columns. How would you decide which @Inheritance strategy the hierarchy should use, and what would you weigh beyond raw query speed?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Weigh read shape against constraint enforcement and change cost. SINGLE_TABLE is fastest but forfeits NOT NULL on subclass columns and grows sparse; JOINED keeps constraints and normalization at a join and INSERT per level, and every new subtype taxes existing polymorphic queries. Also ask whether the hierarchy should exist at all.

open as a page

showing 31–46 of 46