skip to content

How would you organise a service so that writes go through mapped JPA entities while reads are served by DTO projections — a lightweight read/write model split — and what does that approach cost you?

level: principalimportance: nice to knowfreq 36%

answer

  1. one DB, two models in code
  2. commands: load by id, mutate managed entity
  3. queries: JPQL/native → per-endpoint record
  4. boundary rule: entities never leave the write side
  5. costs: duplicated schema knowledge, lost entity cache, type sprawl

basics

~20 s

Keep one database and one schema. Commands load entities by id and mutate them; queries run projection queries into per-use-case DTOs and never return entities. Costs: two sets of types over one schema, duplicated knowledge of the mapping, and weaker reuse of entity behaviour and caching.

solid answer

~50 s

The split is organisational, not infrastructural: same database, same transactions, two code paths. **Write side.** Commands load managed entities by identifier, enforce invariants on the domain object, rely on dirty checking, cascades and `@Version`. Entities never leave this side. **Read side.** Queries are hand-written JPQL/HQL or native SQL projecting straight into DTOs or records shaped for one endpoint. No entity is loaded, so no persistence-context growth and no lazy navigation. Read-side code may legitimately bypass the domain model and even query views. **Boundary rule.** Entities are never returned from the write side and never serialised; DTOs are never used to persist. That single rule is what keeps the split honest. **Costs.** Schema knowledge now lives in two places, so a column rename touches mappings *and* queries; derived logic on entities may be reimplemented in SQL; the entity second-level cache stops helping reads; and you accumulate many narrow types. It pays off on read-dominated services with divergent read shapes, and is overhead on small CRUD.

code

java · 18 lines
java
// write side: managed entity, dirty checking does the UPDATE
void rename(long orderId, String name) {
    Order order = em.find(Order.class, orderId);
    order.rename(name);
}

// read side: query-shaped record, nothing managed
record OrderRow(Long id, String customer, BigDecimal total, int lines) {}

List<OrderRow> recent(int limit) {
    return em.createQuery("""
            select new com.app.query.OrderRow(o.id, c.name, o.total, size(o.lines))
            from Order o join o.customer c
            order by o.placedAt desc
            """, OrderRow.class)
        .setMaxResults(limit)
        .getResultList();
}

go deeper

for a junior

Say that writes use entities and reads use DTO queries, and that reading a DTO cannot update anything.

for a middle

Add the boundary rule and the concrete benefit — read paths stop paying persistence-context and dirty-checking costs.

for a senior

Discuss placement, enforcement, native SQL and views on the read side, and duplicated schema knowledge as the main cost.

for a principal

Weigh it as an architectural call: when read shapes justify divergence, how caching and behaviour duplication shift, how to enforce the boundary, and how it becomes a seam toward replicas or materialised reads.

## What "lightweight" means here Full CQRS usually implies separate read stores fed asynchronously, with eventual consistency and its own operational burden. The lightweight variant keeps **one database, one schema, one transaction manager** and separates only the *models in code*: a write model of mapped entities and a read model of query-shaped DTOs. Reads see committed data immediately, because they read the same tables. Nothing about it requires messaging, projections-as-materialised-views, or a second datastore. ## Shape of the code **Command side.** A command loads the aggregate by identifier, calls behaviour on it, and lets the flush write the changes. This is where mapped entities earn their cost: automatic dirty checking, cascade of state to children, `@Version` for concurrent updates, lifecycle callbacks. Result types are small — an identifier, or nothing. **Query side.** A query object owns a JPQL/HQL statement (or native SQL when the ORM's expressiveness runs out — window functions, recursive CTEs, vendor-specific JSON operators) that selects exactly the endpoint's fields into a record. It uses joins, aggregates, `limit`/`offset`, and can read database **views** that encapsulate the joining. It is allowed to know about tables rather than aggregates, and that permission is the whole point: read shapes rarely match write aggregates. **The boundary rule.** Entities do not cross into the read side, and DTOs are never handed to `persist`/`merge`. If the read side needs a value the write model computes, either duplicate the expression in SQL or move it into a shared pure function that takes primitives. ## Why the split helps - Read paths stop paying the managed-entity bill: no snapshots, no dirty checking, no persistence-context growth, no accidental per-row statements. - Read shapes can diverge freely from the aggregate: a dashboard row spanning five tables is one query, not a graph walk. - Write-side mappings can stay strict and small — fewer associations mapped purely "because a screen needs them", which in turn removes fetch-strategy landmines. - Read queries become independently optimisable: you can rewrite one to native SQL, add a covering index, or point it at a read replica without touching the domain. ## What it costs **Duplicated schema knowledge.** The mapping and the projection queries both encode column and relationship facts. A rename must update both, and only the mapping side is checked by the schema validator — string queries fail at runtime. Mitigate with an executed-query test per projection and a schema-validation check at startup. **Behaviour reimplemented in SQL.** Rules living as entity methods often reappear as `case` expressions. Where the rule is important, keep one source of truth: compute it in SQL and let the write model read it back, or compute it in Java over primitives shared by both sides. Silent divergence between the two implementations is the real hazard. **Caching.** A second-level entity cache serves entity reads; DTO queries bypass it. If reads were largely cache hits, projecting can increase database load. The read side then needs its own strategy — query cache, application cache of DTOs, or materialised views — which is more moving parts. **Type proliferation.** One record per endpoint is deliberate but adds up. Resist the pressure to merge them back into one wide half-filled DTO; that is a regression to the coupling you were escaping. **Cognitive cost.** Newcomers must learn *why* two ways exist and when to use each. Without an explicit rule, someone will return an entity from a query handler and the split quietly dissolves. ## Practical placement Give the two sides different packages and different entry types — for example `orders.domain` holding entities plus a small repository of by-id lookups, and `orders.query` holding query classes returning records. Enforce the boundary mechanically where you can: an architecture test that forbids the query package from referencing entity types, and forbids controllers from importing entity types at all, is far more durable than a convention in a wiki. ## When to adopt it Good fit: read-dominated services, endpoints whose payloads span aggregates, reporting or list-heavy screens, systems where read scale outgrows write scale. Poor fit: small CRUD where every read is a form the user immediately edits, or a team of two shipping a first version — the extra ceremony buys little and the entity read is cheap at small volumes. ## Where it stops If read load genuinely outgrows one database, this design gives a clean seam to go further: point the query side at a read replica (accepting replication lag), or materialise heavy aggregations into a table refreshed on a schedule. That step introduces staleness and is a different decision — it belongs to capacity planning, not to the in-process split described here.

  • With reads and writes hitting the same database, do you get any consistency problems from this split?
    No inherent ones — both sides read committed data from the same tables in the same transactional context, so there is no eventual consistency to reason about. Problems appear only if you later move the read side to a replica or a materialised table, which introduces replication or refresh lag and turns it into a genuine staleness decision.
  • How do you stop the boundary from eroding over time?
    Make it mechanical rather than cultural: an architecture test that forbids the query package from importing entity types and forbids web controllers from importing them at all, plus a review rule that a query handler never returns a mapped class. Conventions documented but unenforced degrade within a few sprints as people reach for the nearest available type.

saying these in an interview costs you the question

  • Equating this with full CQRS and claiming eventual consistency where a single database is used
  • Returning entities from query handlers and calling the split done
  • Persisting DTOs by merging them back into entities
  • Merging per-endpoint DTOs into one wide shared type to "reduce duplication"
  • Ignoring that DTO queries bypass the second-level entity cache and can raise database load

context