What roles do Hibernate's SessionFactory and JPA's EntityManagerFactory play compared with a Session or EntityManager, and which of these objects are safe to share across threads?
answer
- factory = heavy, immutable, thread-safe, one per DB
- session = cheap, single-threaded, one unit of work
- L2 cache on the factory, L1 on the session
- always close: pool drain looks like a hang
- flush+clear in long loops
basics
~20 sThe factory is a heavyweight, immutable, thread-safe, application-scoped object holding mappings, the connection pool, caches and query plans. A Session/EntityManager is cheap, short-lived, single-threaded and owns one persistence context. Never share a session between threads.
solid answer
~50 s**SessionFactory / EntityManagerFactory** — built once at startup from the mappings and configuration. It is expensive to create (parsing annotations, building the metamodel, generating SQL templates), immutable afterwards, and **thread-safe by design**. It owns the connection pool, the second-level cache, the query-plan cache and the statistics. One per persistence unit, i.e. per database, for the life of the application. `SessionFactory extends EntityManagerFactory` since Hibernate 5.2, so `emf.unwrap(SessionFactory.class)` gets you the native view. **Session / EntityManager** — obtained from the factory, cheap to create, and **not thread-safe at all**. It holds a persistence context (the first-level cache and the queue of pending actions) and, from first need until close, a JDBC connection. Its natural scope is one unit of work: one transaction, one request, one job step. Sharing a session across threads corrupts the identity map and interleaves flushes on one JDBC connection — the resulting bugs are non-deterministic. The correct pattern is one session per thread per unit of work, always closed.
code
java · 9 lines// once per application
EntityManagerFactory emf = Persistence.createEntityManagerFactory("orders");
// per unit of work, per thread
try (EntityManager em = emf.createEntityManager()) {
em.getTransaction().begin();
em.persist(new Order("A-1"));
em.getTransaction().commit();
}go deeper
State the split: factory built once and shared, session per unit of work and never shared between threads. Know that both exist in a JPA and a Hibernate flavour.
Explain what each object holds — pool, L2 cache, metamodel versus persistence context, action queue, connection — and why that makes one thread-safe and the other not.
Bring operational detail: connection-acquisition mode, unclosed sessions draining the pool and presenting as a hang, and flush/clear or StatelessSession for long jobs.
Discuss unit-of-work boundaries across an application — how long a persistence context should live, per-datasource factory topology for multiple databases, and the memory/connection budget those choices imply.
## The factory Creating a `SessionFactory` means reading the configuration, scanning and validating every mapped class, building the JPA metamodel, precomputing SQL for load/insert/update/delete on each entity, establishing the connection pool and initialising the second-level cache regions. That work is measured in hundreds of milliseconds to seconds, which is why it happens once, at bootstrap. Afterwards the factory is effectively immutable and explicitly **thread-safe**: many threads call `openSession()`/`createEntityManager()` on it concurrently. It also holds the long-lived shared machinery: - the JDBC connection pool; - the **second-level cache** (shared across sessions, unlike the first-level cache); - the query-plan / HQL-parse cache; - `Statistics`, if enabled; - named queries and the metamodel. There should be exactly one per persistence unit. The classic anti-pattern — building a factory per request — exhausts memory and connections almost immediately, and is a question interviewers use to check whether the candidate has ever bootstrapped Hibernate themselves. ## The session `Session`/`EntityManager` is the per-unit-of-work object. Creating one allocates little more than an empty persistence context; the JDBC connection is typically acquired lazily on first database access and released at transaction end, depending on the connection-handling mode. What it owns: - the **persistence context** (first-level cache): the identity map of managed entities plus the loaded-state snapshots used for dirty checking; - the **action queue** of pending inserts, updates and deletes awaiting flush; - the current flush mode, enabled filters, and the session-level batching state. All of that is mutable, unsynchronised state, so a session is single-threaded by contract. Two threads flushing one session can interleave statements on the same JDBC connection, produce duplicate managed instances for the same row, or corrupt the action queue. Nothing in Hibernate detects this for you; symptoms show up as impossible SQL ordering, phantom `ConcurrentModificationException`s and connection-state errors. The complementary mistake is a session that lives too long. Because the persistence context never evicts anything on its own, a session used for a long job accumulates every entity it touched, holding memory and forcing dirty checking to walk an ever-growing set at each flush. The remedies are periodic `flush()` + `clear()` in batch loops, or `StatelessSession`, which has no persistence context at all. ## Lifecycle summary | | Factory | Session | |---|---|---| | Created | once, at startup | per unit of work | | Cost | expensive | cheap | | Thread-safe | yes | **no** | | Holds | pool, L2 cache, metamodel, plans | persistence context, action queue, a connection | | Closed | at shutdown | at end of the unit of work, always | ## Getting one from the other ```java EntityManagerFactory emf = Persistence.createEntityManagerFactory("orders"); SessionFactory sf = emf.unwrap(SessionFactory.class); try (EntityManager em = emf.createEntityManager()) { em.getTransaction().begin(); ... em.getTransaction().commit(); } ``` JPA 3.1's `EntityManager` implements `AutoCloseable`, so try-with-resources is the idiomatic guard against leaks. In Hibernate 6 the factory also offers `sf.inTransaction(session -> ...)` and `sf.fromTransaction(session -> ...)`, which open a session, run a transaction, commit or roll back, and close — removing the boilerplate where most leaks hide. ## What leaking looks like An unclosed session keeps its persistence context and, in the common configuration, its JDBC connection. Under load the pool drains and every thread blocks waiting for a connection, so the application appears hung rather than erroring. Diagnosing it means looking at pool metrics (active connections pinned at max) and Hibernate statistics (session-open count far exceeding session-close count) — the kind of concrete answer that separates a senior response from a textbook one.
- What actually goes wrong if two threads share one EntityManager?Its persistence context, action queue and JDBC connection are unsynchronised mutable state. Threads interleave flushes on one connection, corrupt the identity map, and can trip ConcurrentModificationException or connection-state errors. Nothing detects the misuse, so the failures are intermittent and load-dependent — which is why the rule is absolute rather than situational.
- Why is creating a SessionFactory per request catastrophic, when creating a Session per request is correct?The factory parses mappings, builds the metamodel, precomputes SQL and stands up a connection pool and cache regions — seconds of work and a lot of retained memory each time, with each instance holding its own pool. A session allocates an empty persistence context and borrows a pooled connection, so its cost is negligible.
- A long-running import job uses one session for a million rows. What goes wrong and what do you change?The persistence context retains every entity it touched, so heap grows and each flush dirty-checks an ever-larger set, slowing down quadratically. Flush and clear in fixed-size batches, or use a StatelessSession, which has no persistence context, no dirty checking and no cascades.
The factory is the workshop: built once, shared, full of tooling. A session is one worker's bench for one job — cheap to set up, private, and cleared away when the job is done. Two workers sharing a bench mid-job is how parts get mixed up.
saying these in an interview costs you the question
- Calling the factory expensive to use rather than expensive to create, and pooling factories
- Sharing a Session or EntityManager across threads because 'it is just a cache'
- Placing the second-level cache on the session or the first-level cache on the factory
- Leaving sessions unclosed and blaming the database when the pool drains
- Holding one session open for an entire batch job