skip to content

Explain the @Aggregation annotation on a MongoRepository — how pipeline methods work and their key features.

level: seniorimportance: should knowfreq 55%

answer

  1. pipeline = String[] of stage docs
  2. $match/$group/$project/$lookup/$unwind/$sort
  3. output doc -> DTO/projection element type
  4. trailing Sort/Pageable -> $sort/$skip/$limit
  5. static pipeline; dynamic -> MongoTemplate Aggregation

basics

~20 s

@Aggregation lets a repository method run a MongoDB aggregation pipeline. You supply an array of pipeline stage strings ($match, $group, $sort, ...) with ?0 placeholders, and Spring maps each output document to the method's return type.

solid answer

~50 s

@Aggregation declares an aggregation pipeline directly on a repository method via its pipeline attribute — an array of JSON stage documents like { $match: { 'status': ?0 } }, { $group: { _id: '$dept', total: { $sum: 1 } } }. Parameters bind with ?0/?1 and SpEL. Each resulting document is mapped to the method's element type, which is often a DTO/projection rather than the entity since aggregation reshapes data. Return List, Stream, a single object, or AggregationResults<T>. A trailing Sort or Pageable adds $sort/$skip/$limit stages automatically. Collation is supported. Use it when you need grouping, faceting, computed fields, joins ($lookup), or multi-stage transforms that a plain filter can't do. Compared to MongoTemplate's typed Aggregation builder, @Aggregation is declarative and concise but static; use the template when the pipeline must be built dynamically at runtime.

code

java · 16 lines
java
public interface OrderRepository extends MongoRepository<Order, String> {

    @Aggregation(pipeline = {
        "{ $match : { 'status' : ?0 } }",
        "{ $group : { _id : '$customerId', total : { $sum : '$amount' } } }",
        "{ $sort  : { total : -1 } }",
        "{ $limit : ?1 }"
    })
    List<CustomerTotal> topCustomers(String status, int limit);
}

// Projection shaped like the pipeline output, NOT the Order entity
public interface CustomerTotal {
    @Value("#{target._id}") String getCustomerId();
    double getTotal();
}

go deeper

for a junior

Know @Aggregation runs a pipeline of stages on a repository method.

for a middle

Know the common stages, parameter binding, and DTO/projection return mapping.

for a senior

Discuss static-vs-dynamic (template) tradeoff, index-friendly early $match, and Sort/Pageable stage injection.

for a principal

Reason about allowDiskUse, $lookup cost, pipeline memory limits, and choosing aggregation vs pre-computed materialized views at scale.

**`@Aggregation`** (`org.springframework.data.mongodb.repository.Aggregation`) binds a repository method to a **MongoDB aggregation pipeline**. An aggregation pipeline is an ordered list of **stages**, each transforming the stream of documents flowing through it — the model MongoDB uses for grouping, computed fields, joins, and analytics that a single `find` filter cannot express. **Declaration:** the `pipeline` attribute is a **String array**, each element a JSON stage document: ``` @Aggregation(pipeline = { "{ $match : { 'status' : ?0 } }", "{ $group : { _id : '$department', total : { $sum : 1 } } }", "{ $sort : { total : -1 } }" }) ``` Common stages: `$match` (filter), `$group` (aggregate with accumulators like `$sum`, `$avg`, `$max`, `$push`), `$project` (reshape/compute), `$sort`, `$limit`, `$skip`, `$unwind` (flatten arrays), `$lookup` (left-join another collection), `$facet`, `$count`, `$addFields`. **Parameter binding:** positional `?0`, `?1` bind method arguments into stage documents, and SpEL `:#{...}`/`?#{...}` supports computed values, exactly like `@Query`. Same injection guidance: bind, don't concatenate untrusted input. **Return types & mapping:** each **output document** of the final stage is mapped to the method's **element type**. Because `$group`/`$project` reshape documents, the return type is frequently a **DTO / interface projection** (e.g. `DeptCount` with `department` + `total`) rather than the `@Document` entity. Supported returns include `List<T>`, `Stream<T>`, a single `T` (for pipelines yielding one doc), and `AggregationResults<T>` (which also carries raw metadata). For a scalar like a count/sum, you can map to a wrapper type or a single-value projection. **Sorting & paging:** a trailing `Sort` parameter is translated into an appended `$sort` stage; a trailing `Pageable` appends `$sort`/`$skip`/`$limit`. Note MongoDB caps some ordering/memory behavior; large `$group`/`$sort` may need `allowDiskUse` (configured via `MongoTemplate`, not the annotation) or supporting indexes on the `$match`/`$sort` fields for performance. **Collation:** the `collation` attribute enables locale-aware string comparison. **When to use `@Aggregation`:** grouping/rollups, computed/derived fields, `$lookup` joins, faceted search, de-duplication, top-N per group — anything multi-stage. Use a **derived method** or `@Query` for simple filters, and the **`MongoTemplate` typed `Aggregation` builder** (`Aggregation.newAggregation(match(...), group(...), ...)`) when the pipeline is **dynamic** (stages depend on runtime conditions) or you want type-safe stage construction. **Gotchas:** - The **element type** of a collection return is the *stage output* shape, not the entity — mismatches leave fields null. - `$match` should come **early** and hit indexes; `$match` after `$group` can't use collection indexes. - Pipelines are **static** in `@Aggregation` — you can't conditionally add stages; switch to the template. - Field references inside stages use `'$field'` syntax; forgetting the `$` compares to a literal string. - Aggregation results ignore the entity's `@Id` unless the pipeline emits `_id`.

  • Why is the return element type of an @Aggregation method usually a DTO rather than the @Document entity?
    Stages like $group and $project reshape documents into a new structure that no longer matches the entity's fields, so Spring maps each output document to a DTO/interface projection matching the pipeline's output shape.
  • When would you use MongoTemplate's Aggregation builder instead of @Aggregation?
    When the pipeline must be built dynamically — stages added or omitted based on runtime conditions — or when you want type-safe stage construction and options like allowDiskUse. @Aggregation pipelines are static strings fixed at compile time.

saying these in an interview costs you the question

  • Mapping aggregation output back to the entity type and expecting reshaped fields to populate
  • Believing you can conditionally add stages inside @Aggregation at runtime
  • Forgetting the '$field' dollar prefix so a field reference becomes a literal

context