What concretely goes wrong when a JPA entity class is used directly as the request and response body of an HTTP API, and what does introducing a separate DTO type at that boundary actually buy you?
answer
- Mass assignment: id, version, audit, tenant, role
- Omitted field → null → merge writes NULL
- Response shape depends on what happened to be fetched
- Schema rename becomes a breaking API change
- Load, authorise, compare version, apply named fields
basics
~20 sInbound, deserialization can set fields the caller should never control (id, version, audit, ownership) and absent fields become nulls that merge writes over real data. Outbound, the object is detached and may drag lazy graphs or leak columns. A DTO makes the writable and readable field lists explicit.
solid answer
~60 s**Inbound problems.** Deserializing a body straight onto an entity means every mapped field is potentially writable: identifiers, `@Version`, audit columns, ownership or role flags — mass assignment. Fields the client omitted arrive as `null`, and because `merge` copies whole state, those nulls are written over existing column values. Identifier presence also decides insert-versus-update, so a caller can steer which row you touch. **Outbound problems.** The object you return is detached; serializing it walks associations, so you either leak columns you never meant to expose, emit huge graphs, or hit lazy-loading errors depending on what was initialised. Bidirectional associations produce cycles that need annotations to break. **Structural problem.** Your wire contract becomes your schema: renaming a column is a breaking API change, and a mapping choice made for the database leaks to clients. A DTO makes both directions explicit: an inbound type listing exactly the fields this operation may change (plus the version token), and an outbound type listing exactly what is exposed — mapped by hand or by a mapper, inside the transaction.
code
java · 4 lines// PUT /orders { "id": 7, "status": "PAID", "tenantId": 42 }
Order saved = em.merge(bodyAsOrder); // id chooses the row,
// status/tenantId are attacker-chosen,
// omitted fields become NULLgo deeper
Say that an entity body lets callers set fields they shouldn't and that omitted fields can be written as null; a DTO names exactly what may go in and out.
Cover both directions — mass assignment and null-overwrite inbound, exposure and unstable fetch-dependent shapes outbound — and show the load-and-apply write path.
Add authorisation on the resolved entity, version-token handling, presence semantics for partial updates, and why response building belongs inside the transaction.
Treat the wire contract as an independently versioned interface: decide where mapping lives, how schema evolution stays invisible to clients, and how field-level exposure and writability are enforced consistently across services.
## Why this pattern is tempting The entity already models the domain, already has the fields, and already serializes. Skipping a DTO removes a class and a mapping step. For a read-only prototype it works. The problems appear the moment the object crosses the boundary in the **write** direction, or the moment the schema and the API need to evolve at different speeds. ## Inbound: mass assignment Deserializing JSON onto an entity means the JSON author decides the value of every mapped, settable field. That includes fields the API never intended as inputs: the primary key, the `@Version` token, `createdBy`/`createdAt`, `tenantId`, `role`, `status`, `balance`. If the object is then merged, all of them are written. This is the same class of vulnerability as mass assignment in any framework: the safe list of writable fields must be defined somewhere, and “whatever the entity happens to expose” is not a decision anyone made. A second, subtler control the caller gains: **which row you touch**. The identifier in the body selects the target of the write. If your handler merges what it is given without checking ownership against the authenticated principal, a caller can supply someone else’s identifier. ## Inbound: absent versus null JSON has no notion of “this field was not mentioned” once it lands in a Java object with a `null` field. Combined with `merge`’s whole-object copy semantics, an omitted field becomes a NULL write. A client that sends `{"id":7,"shippingAddress":"..."}` — perfectly reasonable from its point of view — nulls every other column. The usual symptoms are lost data on partial updates and confusing NOT NULL constraint violations. Related: if the mapper never sets `version`, the entity may look transient and be inserted rather than updated (see the version discussion for detached merges). The version must be an explicit part of the wire contract for updates to be safe. ## Outbound: detached graphs and leakage What you return is a detached object. Two failure modes follow. First, whatever associations happen to be initialised get serialized; whatever is not initialised either serializes as empty/null or, if a serializer touches a lazy reference after the persistence context closed, raises a lazy-initialisation error. This makes the shape of the response depend on the internal fetch decisions of the query that produced it — an unstable contract. Second, everything mapped is exposed by default: internal flags, password hashes, cost columns, soft-delete markers. Bidirectional associations serialize into cycles, which teams then patch with serializer annotations sprinkled across entity classes — persistence classes accumulating presentation concerns. ## Structural coupling With entities on the wire, the schema *is* the contract. Renaming a column, splitting a table, changing a numeric type, introducing an embeddable — each becomes a client-visible change, and each API-shaped requirement (a computed field, a differently named property, a flattened view) pushes back into the mapping. The two models have genuinely different reasons to change: the database model optimises storage, integrity and query plans; the API model optimises the consumer’s use case and stability. ## What a DTO boundary gives you 1. **An explicit writable field list per operation.** “This endpoint may change shippingAddress and note.” Everything else is unreachable by construction, not by review. 2. **Explicit read exposure.** Adding a column does not silently publish it. 3. **Presence semantics you control.** Optional wrappers, JSON Patch, or per-operation types let “absent” differ from “set to null”. 4. **Stable contracts.** Schema refactors stay behind the mapping. 5. **Predictable fetching.** Build the response inside the transaction (or with a projection query), so nothing depends on lazy state after the context closes. 6. **A natural place for the version token,** validation rules per operation, and per-role field visibility. ## The write path that pairs with it Inside the transaction: load the managed entity by identifier, authorise it against the caller, compare the incoming version token, then apply only the fields the DTO carries. Dirty checking emits an UPDATE for exactly what changed, no detached object is ever merged, and the collection-deletion hazards of merged graphs never arise. ## Reasonable exceptions Read-only internal endpoints, admin tools, and quick prototypes can survive with entities on the wire, especially with projection queries returning read-only rows. The line worth defending in an interview: **never accept an entity as an inbound write payload**; outbound is a judgement call about exposure and stability.
- Isn't writing DTOs and mapping code just boilerplate that duplicates the entity?They coincide early and diverge later, because the two models change for different reasons: storage and integrity on one side, consumer needs and contract stability on the other. The mapping code is where that divergence is expressed, and it is also where the writable-field list and exposure decisions live — decisions that otherwise exist nowhere. For genuinely read-only endpoints, a projection query straight into a DTO avoids both the entity and the hand mapping.
- How do you distinguish “field omitted” from “field explicitly set to null” in a partial update?Not with a plain entity — both land as null. Use a per-operation DTO with a presence-aware representation (Optional-typed fields, a JsonNullable-style wrapper, or an explicit patch format such as JSON Patch), or define the endpoint as a full replacement so absence is unambiguous. Then apply only the fields marked present onto the managed entity.
saying these in an interview costs you the question
- Assuming validation annotations on the entity make the inbound payload safe
- Thinking merge is a partial update that ignores null fields
- Serializing entities and patching cycles with annotations on the persistence class
- Trusting the identifier in the body without an ownership check
- Claiming DTOs are always pointless duplication