What performance and correctness options matter when running large MongoTemplate aggregations, and how does result mapping work?
answer
- 100 MB per stage, 16 MB per doc
- allowDiskUse for blocking $group/$sort spill
- only leading $match/$sort use indexes
- AggregationOptions: collation/maxTime/batchSize/comment
- DTO mapping: unmatched projected fields silently dropped
basics
~20 sFilter early with an indexed $match, sort/limit for top-N, and set AggregationOptions.allowDiskUse(true) when $group/$sort exceed the 100 MB in-memory stage limit. Spring maps each result document to your output DTO via getMappedResults; unmatched fields are dropped.
solid answer
~40 sAggregations run server-side but have hard limits: each stage may use up to 100 MB of RAM, and a single returned document can't exceed 16 MB. For big $group/$sort you enable Aggregation.newAggregation(...).withOptions(newAggregationOptions().allowDiskUse(true).build()) so blocking stages spill to disk. Performance is driven by an early, index-eligible $match and $sort — a $sort backed by an index, or right after $match, avoids a blocking in-memory sort; $limit after $sort gives top-N cheaply. Watch for $unwind exploding document count and $lookup needing an index on the foreign field. AggregationOptions also carries collation (locale-aware string comparison), comment, batchSize, and maxTime. On mapping: aggregate(agg, In.class, Out.class) maps each output BSON doc to Out via the MappingMongoConverter; projected field names must match Out's properties or they're silently dropped. Use explain via the options/AggregationOptions or profiler to validate index usage.
code
java · 22 linesimport static org.springframework.data.mongodb.core.aggregation.Aggregation.*;
import org.springframework.data.mongodb.core.aggregation.AggregationOptions;
import org.springframework.data.mongodb.core.query.Collation;
import java.time.Duration;
AggregationOptions options = AggregationOptions.builder()
.allowDiskUse(true) // spill big $group/$sort to disk
.collation(Collation.of("en").strength(2)) // case-insensitive compare
.maxTime(Duration.ofSeconds(30)) // server time budget
.comment("nightly-revenue-report")
.build();
Aggregation agg = newAggregation(
match(where("createdAt").gte(since)), // indexed, filters early
group("customerId").sum("amount").as("total"),
sort(Sort.by(Sort.Direction.DESC, "total")),
limit(50))
.withOptions(options);
List<CustomerRevenue> rows =
mongoTemplate.aggregate(agg, Order.class, CustomerRevenue.class)
.getMappedResults(); // empty list, never nullgo deeper
Aware aggregations return mapped results and that filtering early is good.
Knows getMappedResults mapping and basic match-first ordering.
Applies allowDiskUse, collation, and index-eligible match/sort; diagnoses null-DTO-field mapping issues.
Governs cluster impact with maxTime/comment/explain, reasons about 100MB/16MB limits, top-k sort optimization, lookup indexing, and typed-vs-untyped field mapping pitfalls at scale.
Running aggregations at scale is where correctness, resource limits, and mapping semantics converge. **Hard server limits.** - **100 MB per stage**: blocking stages (`$group`, `$sort`, `$bucket`) that need to materialize data are capped at 100 MB of RAM. Exceeding it errors unless you allow disk spill. - **16 MB per document**: any single document produced (including a big `$lookup`-joined array or a `$group` that `$push`es huge arrays) can't exceed BSON's 16 MB limit. - **`allowDiskUse`**: `AggregationOptions.builder().allowDiskUse(true).build()`, attached via `newAggregation(...).withOptions(options)`, lets blocking stages spill to temp disk files, trading speed for capacity. (On very new servers disk use may be on by default, but set it explicitly for portability.) **Performance levers.** - **Index-eligible `$match` first**: only a `$match` (and a leading `$sort`) at the *start* of the pipeline can use collection indexes. Once documents pass through a transforming stage, later `$match`es filter in memory. - **`$sort` + index**: a sort that matches an index avoids a blocking in-memory sort. `$sort` immediately followed by `$limit` lets the engine keep only top-N in memory (top-k optimization). - **`$project`/`$unset` early** to shrink documents before expensive stages reduces memory pressure. - **`$unwind` multiplies documents**; unwinding multiple arrays is combinatorial — can blow the 100 MB stage limit. - **`$lookup`** should target an indexed `foreignField`; otherwise each left doc triggers a collection scan of the foreign collection. **AggregationOptions (`org.springframework.data.mongodb.core.aggregation.AggregationOptions`).** Built via `AggregationOptions.builder()`: - `.allowDiskUse(true)` — disk spill (above). - `.collation(Collation.of("en").strength(...))` — locale-aware, case/accent-insensitive string comparison in `$match`/`$sort`/`$group`. - `.cursorBatchSize(n)` — network batch size for large result cursors. - `.maxTime(Duration)` — server-side time budget; the op aborts if exceeded. - `.comment("...")` — tag for profiler/log correlation. - `.explain(true)` (or `mongoTemplate` explain paths) — returns the query plan instead of results to verify index usage. Attach with `newAggregation(stages).withOptions(options)`. **Result mapping.** `mongoTemplate.aggregate(agg, inputType, OutputType.class)` returns `AggregationResults<OutputType>`. Each output BSON document is mapped to `OutputType` by the `MappingMongoConverter` (same converter used for entities). Key semantics: - Field names in the projected output must correspond to `OutputType` properties (respecting `@Field`); **unmatched fields are silently ignored**, a frequent "why is my DTO field null" bug. - The `_id` from `$group` maps to the property named `id`/`_id` — project it to a clean name if your DTO expects one. - `getMappedResults()` returns a `List` (empty, never null); `getUniqueMappedResult()` expects at most one row and throws otherwise. - You can map to a generic `Document`/`Map` output type when the shape is dynamic. **Typed vs untyped aggregation.** With a typed input class, Spring maps Java property names to stored field names in stages. With the untyped `newAggregation(stage...)` (no class), you must use the raw stored field names. Mixing them causes silent mismatches. **Observability.** Use `maxTime` to bound runaway pipelines, `comment` to find them in the profiler/`system.profile`, and an explain plan to confirm `$match`/`$sort` hit indexes rather than doing COLLSCAN/blocking sorts. **When to use these.** Reach for `allowDiskUse` and collation on reporting/analytics pipelines over large collections; enforce `maxTime` on user-triggered aggregations to protect the cluster; and always validate index usage on the leading `$match` and any `$lookup` foreign field before shipping.
- A $group over a huge collection fails with an exceeded-memory error. What is the fix and its trade-off?Enable AggregationOptions.allowDiskUse(true) so blocking stages spill to temporary disk files instead of being capped at 100 MB RAM. The trade-off is slower execution due to disk I/O; also revisit whether an earlier indexed $match can reduce the data first.
- Your output DTO field is always null even though the pipeline produces the value. Likely cause?The projected field name in the pipeline does not match the DTO property (or its @Field name). Spring's MappingMongoConverter silently ignores unmatched fields, so align the $project alias with the DTO property, and remember $group's key lives under _id.
- Which stages in a pipeline can actually use a collection index?Only a $match or $sort at the very start of the pipeline (before any transforming stage), plus the foreign field of a $lookup if indexed. Once documents pass through $group/$project/$unwind, later $match/$sort operate in memory.
saying these in an interview costs you the question
- Thinking any $match anywhere in the pipeline uses indexes
- Believing aggregations have no size limits
- Assuming allowDiskUse is always on / never needed
- Not knowing unmatched projected fields are silently dropped during DTO mapping
- Confusing collation with a filter rather than comparison rules