An EntityManager has unsaved changes pending when raw SQL is executed through createNativeQuery. What does Hibernate do about the pending state, why does it behave differently than for a JPQL query, and how do you narrow that behaviour?
answer
- Auto-flush decided by query spaces
- Native SQL unparseable → assume all spaces
- Whole session flushed + query cache evicted
- addSynchronizedEntityClass / QuerySpace / querySpaces
- COMMIT or MANUAL flush → SQL sees pre-change DB
basics
~20 sFor JPQL, Hibernate flushes only if pending changes touch tables the query reads. It cannot parse native SQL, so it assumes the query touches everything: it flushes the whole session and invalidates all cached query results. Declaring the query's synchronized tables or entity classes narrows both.
solid answer
~60 sWith the default AUTO flush mode, Hibernate auto-flushes before a query so the query sees your pending changes. For JPQL it computes **query spaces** — the tables the query reads — and flushes only when a pending change touches one of them. Native SQL is an opaque string: Hibernate cannot know which tables it reads, so it takes the conservative route — **flush the entire persistence context** and treat the query as invalidating **all** query spaces, which evicts the query cache broadly. In a long persistence context with many dirty entities, a small native lookup can therefore trigger a large flush, and repeated native queries can keep the query cache permanently cold. The fix is to tell Hibernate what the query actually touches: ```java session.createNativeQuery(sql) .addSynchronizedEntityClass(Book.class); // or addSynchronizedQuerySpace("book") ``` and for declarations, the `querySpaces` attribute of Hibernate's `@NamedNativeQuery`. Now only those spaces drive the flush decision and the cache invalidation. The mirror-image hazard: under `FlushModeType.COMMIT` or a manual flush mode, no flush happens, so the SQL runs against a database that does not yet contain your pending inserts and updates.
go deeper
Know that a query normally triggers an automatic flush so it sees your pending changes, and that native SQL is a blunter version of that.
Explain query spaces and why native SQL forces the conservative all-spaces assumption, and name addSynchronizedEntityClass as the narrowing tool.
Diagnose it from symptoms — unexplained flush cost on a report path, query-cache hit ratio collapsing after a native query was added, a native UPDATE that appears to do nothing — and prescribe flush/execute/clear around native DML.
Treat it as a unit-of-work design question: how long persistence contexts live, whether native reads belong inside them at all, and how cache invalidation contracts are declared and reviewed across a codebase.
## The mechanism: auto-flush and query spaces A persistence context accumulates changes — new entities, modified ones, removals — and writes them at flush time. A query executed before that flush would otherwise read a database that disagrees with what the application has already done. So with the default `FlushModeType.AUTO`, Hibernate **auto-flushes before executing a query**. Flushing is not free: it walks the persistence context, dirty-checks every managed entity against its loaded-state snapshot, and issues the resulting DML. Hibernate therefore optimises the decision using **query spaces** — essentially the set of tables a query touches. Before a JPQL query it compares the query's spaces with the spaces that pending changes would write to. No overlap, no flush. Selecting books while only comments are dirty flushes nothing. The same query spaces drive **query-cache invalidation**: when a statement writes to a table, cached results of queries over that table are stale and must be dropped. ## Why native SQL is different Hibernate does not parse native SQL. It cannot: the string may use dialect syntax, CTEs, views, synonyms or functions it has no grammar for. So it has no idea which tables the query reads. Facing that, it makes the only safe assumption — **the query might touch anything** — with two consequences: 1. **The whole persistence context is flushed.** Every dirty entity is dirty-checked and written, not just those relevant to the query. In a long-lived persistence context this can be a substantial cost paid by a query that reads a single unrelated row. 2. **Every query space is treated as invalidated.** The query cache is broadly evicted. A workload that mixes native queries with query caching can end up with a cache that never survives long enough to help — and diagnosing that from a cache-hit-ratio graph without knowing this rule is genuinely hard. Note what does *not* happen: native SQL results are not made stale by this, and the second-level entity cache is not wholesale wiped by a mere read. The pain is flush cost plus query-cache churn. ## Narrowing it Hibernate lets you supply the information it could not derive: ```java List<Book> books = session.createNativeQuery("select * from book where ...", Book.class) .addSynchronizedEntityClass(Book.class) // implies the book table .getResultList(); // or by table name query.addSynchronizedQuerySpace("book"); ``` For a declared query, Hibernate's own `@NamedNativeQuery` carries a `querySpaces` attribute listing the tables. Once spaces are declared, the native query behaves like a JPQL query: it flushes only when pending changes touch those tables, and it invalidates only those regions of the query cache. A subtlety worth stating: `addSynchronizedQuerySpace("")` — an empty space set that is nonetheless *declared* — is the documented way to say "this query touches nothing", which suppresses the flush entirely. That is a sharp tool: use it only for genuinely independent reads (a lookup table, a monitoring probe), because if the query really does read a table you have pending writes for, you have just guaranteed a stale read. Declaring spaces is also a correctness statement, not just a performance one. If you declare too few and the query reads a table you have dirty entities for, Hibernate will skip a flush it needed, and the SQL sees pre-change data. ## The other direction: no flush at all The symmetric hazard is a persistence context in `FlushModeType.COMMIT` (or Hibernate's `MANUAL`). Now nothing is flushed before a query, native or not. A native `SELECT` then runs against a database that has never seen your pending inserts. Native DML is worse: an `UPDATE ... WHERE` written to affect rows you just inserted in memory silently affects nothing, and the row counts you branch on are wrong. The general rule for native DML holds regardless of flush mode: `executeUpdate` on native SQL goes straight to the database. Managed entities are not updated to match, lifecycle callbacks and cascades do not run, and second-level cache entries for the affected rows are not invalidated unless the query's synchronized spaces say so. So the safe pattern for a bulk native statement is: flush, run it, then clear the persistence context (or explicitly evict the affected regions), so nothing stale is left behind. ## How this shows up in production - A report page that runs a native query in the middle of a long unit of work becomes slow, and profiling blames a flush that has nothing to do with the report. - Query-cache hit ratio collapses after someone adds a native query to a hot path. - A native `UPDATE` appears to do nothing, because the rows it targets exist only in the persistence context. - Entities read after a native `UPDATE` show pre-update values, because the identity map still holds the old instances. ## The answer to give Name the mechanism (auto-flush driven by query spaces), state the asymmetry (Hibernate cannot parse native SQL, so it assumes everything), give the fix (`addSynchronizedEntityClass` / `addSynchronizedQuerySpace` / `querySpaces`), and mention both failure directions — over-flushing plus cache churn on one side, stale reads under `COMMIT`/manual flush mode on the other.
- What is the risk of declaring synchronized query spaces that are too narrow for what the SQL actually reads?You suppress a flush that was needed. Hibernate will decide the pending changes do not overlap the query's tables and skip writing them, so the SQL runs against pre-change data and returns stale or missing rows. You also leave query-cache entries for the untouched tables intact when the query did in fact write or read them. The declaration is a correctness assertion, so it must list every table the SQL genuinely touches.
- You run a bulk native UPDATE through executeUpdate in the middle of a unit of work. What must you do around it?Flush first so pending in-memory changes are in the database before the statement runs, then execute it, then clear or selectively evict the persistence context because the managed entities it touched still hold pre-update state and will be dirty-checked against a stale snapshot. Declare the statement's synchronized spaces so the second-level and query caches for those tables are invalidated. Lifecycle callbacks, cascades and versioning do not run for native DML, so any @Version column must be bumped by the SQL itself.
saying these in an interview costs you the question
- Believing Hibernate parses native SQL well enough to work out which tables it reads
- Assuming a small native SELECT causes only a small flush regardless of what else is pending
- Declaring synchronized spaces purely as a performance tweak without realising too-narrow spaces cause stale reads
- Expecting native DML to update managed entities, run cascades or bump a @Version column
- Thinking FlushModeType.COMMIT only delays writes and cannot change what a query sees