skip to content

What performance and correctness options matter when running large MongoTemplate aggregations, and how does result mapping work?

level: principalimportance: should knowfreq 30%

answer

  1. 100 MB per stage, 16 MB per doc
  2. allowDiskUse for blocking $group/$sort spill
  3. only leading $match/$sort use indexes
  4. AggregationOptions: collation/maxTime/batchSize/comment
  5. DTO mapping: unmatched projected fields silently dropped

basics

~20 s

Filter early with an indexed $match, sort/limit for top-N, and set AggregationOptions.allowDiskUse(true) when $group/$sort exceed the 100 MB in-memory stage limit. Spring maps each result document to your output DTO via getMappedResults; unmatched fields are dropped.

solid answer

~40 s

Aggregations run server-side but have hard limits: each stage may use up to 100 MB of RAM, and a single returned document can't exceed 16 MB. For big $group/$sort you enable Aggregation.newAggregation(...).withOptions(newAggregationOptions().allowDiskUse(true).build()) so blocking stages spill to disk. Performance is driven by an early, index-eligible $match and $sort — a $sort backed by an index, or right after $match, avoids a blocking in-memory sort; $limit after $sort gives top-N cheaply. Watch for $unwind exploding document count and $lookup needing an index on the foreign field. AggregationOptions also carries collation (locale-aware string comparison), comment, batchSize, and maxTime. On mapping: aggregate(agg, In.class, Out.class) maps each output BSON doc to Out via the MappingMongoConverter; projected field names must match Out's properties or they're silently dropped. Use explain via the options/AggregationOptions or profiler to validate index usage.

code

java · 22 lines
java
import static org.springframework.data.mongodb.core.aggregation.Aggregation.*;
import org.springframework.data.mongodb.core.aggregation.AggregationOptions;
import org.springframework.data.mongodb.core.query.Collation;
import java.time.Duration;

AggregationOptions options = AggregationOptions.builder()
    .allowDiskUse(true)                         // spill big $group/$sort to disk
    .collation(Collation.of("en").strength(2)) // case-insensitive compare
    .maxTime(Duration.ofSeconds(30))            // server time budget
    .comment("nightly-revenue-report")
    .build();

Aggregation agg = newAggregation(
        match(where("createdAt").gte(since)),  // indexed, filters early
        group("customerId").sum("amount").as("total"),
        sort(Sort.by(Sort.Direction.DESC, "total")),
        limit(50))
    .withOptions(options);

List<CustomerRevenue> rows =
    mongoTemplate.aggregate(agg, Order.class, CustomerRevenue.class)
                 .getMappedResults();   // empty list, never null

go deeper

for a junior

Aware aggregations return mapped results and that filtering early is good.

for a middle

Knows getMappedResults mapping and basic match-first ordering.

for a senior

Applies allowDiskUse, collation, and index-eligible match/sort; diagnoses null-DTO-field mapping issues.

for a principal

Governs cluster impact with maxTime/comment/explain, reasons about 100MB/16MB limits, top-k sort optimization, lookup indexing, and typed-vs-untyped field mapping pitfalls at scale.

Running aggregations at scale is where correctness, resource limits, and mapping semantics converge. **Hard server limits.** - **100 MB per stage**: blocking stages (`$group`, `$sort`, `$bucket`) that need to materialize data are capped at 100 MB of RAM. Exceeding it errors unless you allow disk spill. - **16 MB per document**: any single document produced (including a big `$lookup`-joined array or a `$group` that `$push`es huge arrays) can't exceed BSON's 16 MB limit. - **`allowDiskUse`**: `AggregationOptions.builder().allowDiskUse(true).build()`, attached via `newAggregation(...).withOptions(options)`, lets blocking stages spill to temp disk files, trading speed for capacity. (On very new servers disk use may be on by default, but set it explicitly for portability.) **Performance levers.** - **Index-eligible `$match` first**: only a `$match` (and a leading `$sort`) at the *start* of the pipeline can use collection indexes. Once documents pass through a transforming stage, later `$match`es filter in memory. - **`$sort` + index**: a sort that matches an index avoids a blocking in-memory sort. `$sort` immediately followed by `$limit` lets the engine keep only top-N in memory (top-k optimization). - **`$project`/`$unset` early** to shrink documents before expensive stages reduces memory pressure. - **`$unwind` multiplies documents**; unwinding multiple arrays is combinatorial — can blow the 100 MB stage limit. - **`$lookup`** should target an indexed `foreignField`; otherwise each left doc triggers a collection scan of the foreign collection. **AggregationOptions (`org.springframework.data.mongodb.core.aggregation.AggregationOptions`).** Built via `AggregationOptions.builder()`: - `.allowDiskUse(true)` — disk spill (above). - `.collation(Collation.of("en").strength(...))` — locale-aware, case/accent-insensitive string comparison in `$match`/`$sort`/`$group`. - `.cursorBatchSize(n)` — network batch size for large result cursors. - `.maxTime(Duration)` — server-side time budget; the op aborts if exceeded. - `.comment("...")` — tag for profiler/log correlation. - `.explain(true)` (or `mongoTemplate` explain paths) — returns the query plan instead of results to verify index usage. Attach with `newAggregation(stages).withOptions(options)`. **Result mapping.** `mongoTemplate.aggregate(agg, inputType, OutputType.class)` returns `AggregationResults<OutputType>`. Each output BSON document is mapped to `OutputType` by the `MappingMongoConverter` (same converter used for entities). Key semantics: - Field names in the projected output must correspond to `OutputType` properties (respecting `@Field`); **unmatched fields are silently ignored**, a frequent "why is my DTO field null" bug. - The `_id` from `$group` maps to the property named `id`/`_id` — project it to a clean name if your DTO expects one. - `getMappedResults()` returns a `List` (empty, never null); `getUniqueMappedResult()` expects at most one row and throws otherwise. - You can map to a generic `Document`/`Map` output type when the shape is dynamic. **Typed vs untyped aggregation.** With a typed input class, Spring maps Java property names to stored field names in stages. With the untyped `newAggregation(stage...)` (no class), you must use the raw stored field names. Mixing them causes silent mismatches. **Observability.** Use `maxTime` to bound runaway pipelines, `comment` to find them in the profiler/`system.profile`, and an explain plan to confirm `$match`/`$sort` hit indexes rather than doing COLLSCAN/blocking sorts. **When to use these.** Reach for `allowDiskUse` and collation on reporting/analytics pipelines over large collections; enforce `maxTime` on user-triggered aggregations to protect the cluster; and always validate index usage on the leading `$match` and any `$lookup` foreign field before shipping.

  • A $group over a huge collection fails with an exceeded-memory error. What is the fix and its trade-off?
    Enable AggregationOptions.allowDiskUse(true) so blocking stages spill to temporary disk files instead of being capped at 100 MB RAM. The trade-off is slower execution due to disk I/O; also revisit whether an earlier indexed $match can reduce the data first.
  • Your output DTO field is always null even though the pipeline produces the value. Likely cause?
    The projected field name in the pipeline does not match the DTO property (or its @Field name). Spring's MappingMongoConverter silently ignores unmatched fields, so align the $project alias with the DTO property, and remember $group's key lives under _id.
  • Which stages in a pipeline can actually use a collection index?
    Only a $match or $sort at the very start of the pipeline (before any transforming stage), plus the foreign field of a $lookup if indexed. Once documents pass through $group/$project/$unwind, later $match/$sort operate in memory.

saying these in an interview costs you the question

  • Thinking any $match anywhere in the pipeline uses indexes
  • Believing aggregations have no size limits
  • Assuming allowDiskUse is always on / never needed
  • Not knowing unmatched projected fields are silently dropped during DTO mapping
  • Confusing collation with a filter rather than comparison rules

context