Entity Mapping & Annotations
How a POJO becomes a row: identifiers and generation strategies, columns and converters, embeddables, enums and temporals, and inheritance hierarchies. Interviewers start here because mapping choices — IDENTITY vs SEQUENCE, SINGLE_TABLE vs JOINED — lock in performance behavior for the life of the schema.
part ofHibernateoverview, primer and where to startread it →on this pageshowhide
explore
- @Entity Basics & Access Types5 questions
- Identifier Generation Strategies5 questions
- Columns & AttributeConverters6 questions
- Embeddables & Composite Keys6 questions
- Enum & Date/Time Mapping6 questions
- Inheritance Strategies5 questions
- Naming Strategies4 questions
- @Formula & Generated Values4 questions
- Entity equals/hashCode & @NaturalId5 questions
questions
page 2 of 2What does the JPA @Lob annotation do to a String or byte[] entity attribute, and what problems do teams run into with it in production?
basics
~20 s@Lob tells the provider to map the attribute to a large-object type — CLOB for String, BLOB for byte[] — instead of varchar or varbinary. In production it costs memory: the whole value is loaded with the row and rewritten on update unless you split it out.
Why does Hibernate insist that a composite identifier class implement equals() and hashCode(), and what concretely goes wrong at runtime when the implementation is missing, incomplete, or based on mutable state?
basics
~20 sHibernate keys the persistence context and the second-level cache by identifier value, comparing keys with equals/hashCode. Without a correct implementation, lookups miss: the same row loads as multiple instances, find() re-queries, merge misbehaves, and cache hits never happen. Keys must also be immutable, or the hash changes under the map.
A link entity's primary key is (order_id, product_id) and the same entity also needs @ManyToOne references to Order and Product. How do you map that in JPA/Hibernate without mapping the same two columns twice?
basics
~20 sUse derived identity: keep an @EmbeddedId holding the two id values and annotate each association with @MapsId("orderId") / @MapsId("productId"). Hibernate then owns the columns once through the associations and fills the key fields from them, so you set only the associations.
What does the Hibernate configuration property hibernate.jdbc.time_zone do, why do teams set it to UTC, and what must you be careful about when introducing it into a system that already has data?
basics
~20 sIt tells Hibernate which time zone to pass to JDBC when binding and reading timestamps, instead of the JVM default. Setting it to UTC makes stored values independent of the server's zone. Introducing it later re-interprets existing rows, so it needs a data migration, not just a config flip.
A production table stores an enum attribute as ordinal integers and the team now needs to insert a new constant in the middle of the enum and move the column to string storage. How do you carry that out safely, and how would you detect damage if someone had already reordered the constants?
basics
~20 sFreeze the current ordinal-to-name mapping from the old source, add a new string column, backfill it by mapping each stored ordinal through that frozen table in SQL, switch the mapping to EnumType.STRING, then drop the integer column. To detect prior damage, reconstruct the mapping from git history and cross-check row counts and timelines against business expectations.
A field is annotated with Hibernate's @ColumnDefault("0") and the table was created by Hibernate's schema generation, yet new rows still arrive with NULL in that column. Why, and what actually makes the database default apply?
basics
~20 s@ColumnDefault only adds DEFAULT to generated DDL. Hibernate still lists every mapped column in the INSERT, so it sends an explicit NULL and an explicit value always beats a column default. Use @DynamicInsert so null columns are omitted, plus @Generated to read the value back — or just set it in Java.
Explain why a Java equals() implementation that starts with getClass() != other.getClass(), or that reads the other object's fields directly, can misbehave when Hibernate hands you a lazily-loaded proxy — and what to write instead.
basics
~20 sA lazy reference is a generated subclass of your entity, so getClass() returns Order$HibernateProxy, not Order, and a getClass() check rejects the same row. Direct field access on a proxy reads the uninitialized subclass fields (null), bypassing the interception. Use instanceof and call getters.
You assign entity identifiers in application code — for example a UUID set in the constructor — instead of letting the database generate them. What does that change for Hibernate's handling of the entity, and what is the storage and index cost of random UUIDs?
basics
~20 sAn id that is never null removes Hibernate's usual new-versus-detached signal, so merge() must SELECT before it knows whether to insert. Random UUIDs are 16 bytes and insert at random points in the primary-key index, causing page splits and poor cache locality; time-ordered UUIDs and a native 16-byte column fix most of that.
An application maps a deep entity hierarchy with JPA's JOINED inheritance strategy and its polymorphic queries have become slow. Explain the SQL these queries generate, what makes them expensive, and the levers available to reduce the cost without abandoning the hierarchy.
basics
~20 sA base-type query under JOINED outer-joins every subclass table so Hibernate can build each concrete instance, so cost grows with hierarchy width. Levers: query the concrete subtype instead, restrict with TYPE()/TREAT(), select a DTO projection of base columns only, index the join keys, or flatten to SINGLE_TABLE.
One of your Hibernate entities maps to a table called order and has a column called user, both reserved words in the database. What options does Hibernate give you to make the generated SQL valid, and what is the risk of switching on the setting hibernate.globally_quoted_identifiers?
basics
~10 sQuote the offending names individually - backticks or escaped double quotes in @Table/@Column - or enable hibernate.auto_quote_keyword. hibernate.globally_quoted_identifiers quotes everything, which makes all identifiers case-sensitive and locks the schema to their exact spelling.
How do you make one attribute of a JPA entity use property access while the rest of the class uses field access, and when is that worth doing?
basics
~20 sDeclare the class default explicitly with @Access(AccessType.FIELD), mark the backing field @Transient so it is not mapped twice, then put @Access(AccessType.PROPERTY) plus the mapping annotations on the getter. Hibernate then calls that getter and setter for that one column.
Hibernate's sequence-backed id generators can use different optimizers — hi/lo, pooled, and pooled-lo. What does the value returned by the database sequence mean in each, and which are safe when another application inserts into the same table?
basics
~20 sWith legacy hi/lo the sequence value is a multiplier, so it bears no relation to stored ids and outside writers collide. With pooled, nextval returns the top of the reserved block; with pooled-lo it returns the bottom. Both keep sequence values inside the id space, so external callers using the same sequence stay safe.
For an association table entity such as order_line, how would you decide between a composite primary key built from the two foreign keys and a single surrogate key with a unique constraint on the pair? What drives the call, and what does each choice cost you later?
basics
~20 sComposite keys model the truth (a pair is unique) with no extra column, but every child table, cache key and query carries both columns and the key must never change. A surrogate key gives one stable, cheap handle to reference and evolve, at the cost of an extra column plus a unique constraint you must not forget.
You are setting the convention for how a multi-service system persists dates and times through JPA/Hibernate entities. How do you decide what each attribute stores — UTC instants, local wall-clock plus a zone, or plain calendar dates — and what do you standardise across the codebase?
basics
~20 sClassify each attribute first: a past moment stores an Instant in UTC; a calendar fact stores LocalDate; a future local commitment stores wall-clock plus an explicit zone id so later zone-rule changes stay correct. Then standardise UTC everywhere in storage, conversion only at edges, and one mapping style per meaning.
You are designing the identity model for a domain whose objects routinely live outside an open persistence context — cached DTO graphs, HTTP sessions, sets nested inside aggregates. How would you decide between assigning identifiers in the constructor (for example UUIDs), relying on a business/natural key, and using database-generated sequence identifiers as the basis for object equality?
basics
~20 sPick by whether an identifier exists at construction time. Assigned UUIDs always do, so equality is simple everywhere, at the cost of index locality — mitigate with time-ordered UUIDs. A true immutable business key is best when one genuinely exists. Generated sequence ids are cheapest in the database but need the constant-hashCode workaround.
You own a JPA domain model where a core entity hierarchy is about to grow from three subtypes to a dozen, several of which add many mandatory columns. How would you decide which @Inheritance strategy the hierarchy should use, and what would you weigh beyond raw query speed?
basics
~20 sWeigh read shape against constraint enforcement and change cost. SINGLE_TABLE is fastest but forfeits NOT NULL on subclass columns and grows sparse; JOINED keeps constraints and normalization at a join and INSERT per level, and every new subtype taxes existing polymorphic queries. Also ask whether the hierarchy should exist at all.
showing 31–46 of 46