When you call repository.save(entity), how does Spring Data decide whether to run an INSERT or an UPDATE by default?
answer
- isNew() drives INSERT vs UPDATE
- default = @Id null (0 for primitive)
- new -> persist, existing -> merge
- works because @GeneratedValue keeps id null
- assigned ids break the null check
basics
~10 sBy default Spring Data treats an entity as new when its @Id is null (or 0 for a primitive id type). New means INSERT; otherwise it's an existing row, so UPDATE.
solid answer
~40 ssave() delegates to an EntityInformation.isNew(entity) check. The default strategy looks at the @Id property: if it's null (or numeric 0 for a primitive id), the entity is considered new and an INSERT is performed; a non-null id means it already exists, so an UPDATE runs. In Spring Data JPA this maps to EntityManager.persist for new entities and merge for existing ones. This default works cleanly with database-generated ids (@GeneratedValue), because the id is null until the row is inserted and Hibernate assigns it. The moment you assign ids yourself, the id is never null, so this default check would classify every save as an update — that's the classic pitfall that pushes you toward @Version-based detection or implementing Persistable.
code
java · 18 lines// Default strategy in action — generated id keeps the null check honest
@Entity
class Book {
@Id
@GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id; // null until first insert -> isNew() == true
private String title;
}
// SimpleJpaRepository.save (paraphrased):
// if (entityInformation.isNew(entity)) em.persist(entity); // INSERT
// else return em.merge(entity); // SELECT + UPDATE
Book b = new Book();
b.setTitle("Spring in Action");
repo.save(b); // id == null -> INSERT, id assigned by DB
b.setTitle("Spring in Action, 6ed");
repo.save(b); // id != null -> UPDATEgo deeper
Know the one-liner: new = null id -> INSERT, non-null id -> UPDATE.
Explain that isNew() drives persist vs merge, and that primitive ids use 0 not null.
Connect the default to why generated ids work and assigned ids break, and the merge SELECT cost.
Frame isNew as a pluggable EntityInformation strategy and reason about its store-specific consequences.
## The core question save() has to answer Spring Data's `CrudRepository.save(S entity)` is a single method for both creating and updating. Internally it must decide: is this a brand-new record that needs an **INSERT**, or an existing one that needs an **UPDATE**? It answers this by asking an `EntityInformation` object: `boolean isNew(entity)`. ## The default strategy: look at the @Id The default `isNew` implementation (in Spring Data Commons' `AbstractEntityInformation`) inspects the entity's identifier property (the field marked `@Id`): - If the id **type is a wrapper/object** (e.g. `Long`, `UUID`, `String`): the entity is new when `id == null`. - If the id **type is primitive** (e.g. `long`, `int`): the entity is new when the value is `0`. So `isNew` returns `true` for a fresh object whose id hasn't been set yet. ## What each store does with that verdict - **Spring Data JPA** (`SimpleJpaRepository.save`): if `isNew` is true it calls `entityManager.persist(entity)` (a pure INSERT); otherwise it calls `entityManager.merge(entity)` (which typically triggers a SELECT then an UPDATE, or an INSERT if the row is missing). - **Spring Data JDBC / R2DBC / MongoDB**: no persistence context exists, so the `isNew` verdict directly chooses an INSERT vs. UPDATE SQL statement (or Mongo insert vs. replace). ## Why the default works with generated ids With `@GeneratedValue` (JPA) or a database identity/sequence column, the id stays `null` in memory until the row is inserted; the database then assigns it. So a newly constructed object correctly reports `isNew == true`, gets inserted, and receives its id. Every subsequent `save()` of that same loaded object sees a non-null id and updates. ## The seed of the pitfall The entire default hinges on **the id being null before the first save**. If you assign the id yourself (a client-provided `UUID`, a natural key, a business code), the id is *never* null, so the default check reports `isNew == false` on a genuinely new object. In JPA that means `merge` runs when you wanted `persist` — an extra SELECT on every insert, and subtle bugs. The fixes are covered by the `@Version` strategy and by implementing `Persistable<ID>`. ## Key terms - **EntityInformation** — Spring Data's metadata object per entity type; exposes `getId`, `getIdType`, `isNew`. - **persist vs. merge** — JPA operations: `persist` is insert-only for a new managed entity; `merge` copies detached state onto a managed entity, usually preceded by a load.
- What value counts as 'new' when the @Id is a primitive long instead of Long?For a primitive id the new-check is value == 0, not null (a primitive can never be null). So an unset primitive long id (0) is treated as new; this is why people generally prefer boxed id types.
- In JPA, what's the practical difference between the persist path and the merge path save() might take?persist is a straight INSERT for a new managed entity. merge assumes the entity may already exist: Hibernate usually issues a SELECT to load current state, then an UPDATE (or an INSERT if not found), and returns a managed copy. merge therefore costs an extra query.
saying these in an interview costs you the question
- Claiming save() always runs an UPDATE if the row exists 'because it checks the database first' — the default checks the in-memory id, not the DB.
- Saying save() looks at @CreatedDate or a timestamp by default — it doesn't; the default is purely the @Id (or @Version) check.
- Thinking a primitive-id entity is new only when id is null (a primitive can't be null; the check is value 0).