How do you return the top 3 highest-scoring documents per group in an aggregation pipeline?
answer
- rank within each group, not overall
- the accumulator sees documents in sort order
- pushing everything then trimming is the naive form
- newer accumulators keep only n per group
- a window function gives explicit tie semantics
basics
~10 sSort descending, then keep only the leaders per group: $sort followed by $group using $topN or $firstN (MongoDB 5.2+), or $push plus $slice on older versions. $setWindowFields with $rank and a $match also works.
solid answer
~50 sThe classic form is `$sort` then `$group`: because accumulators consume documents in the order the preceding `$sort` produced, `{ $push: "$$ROOT" }` builds a per-group array already in rank order, and a following `$set` with `{ $slice: ["$docs", 3] }` trims it. The weakness is that `$push` materialises **every** document of every group before trimming, which is fine for small groups and dangerous for large ones. Since MongoDB 5.2 the bounded accumulators fix that: `{ $topN: { output: ["$player", "$score"], sortBy: { score: -1 }, n: 3 } }` keeps only three per group and does its own ordering, so no preceding `$sort` is required; `$firstN` does the same for a stream already sorted. The third option is `$setWindowFields` with `$rank` or `$denseRank` partitioned by the group, then `$match` on rank ≤ 3 — the most SQL-like, and the only one that gives you tie semantics explicitly.
code
javascript · 7 lines// bounded: keeps only three per group (MongoDB 5.2+)
db.scores.aggregate([
{ $group: {
_id: "$team",
top: { $topN: { output: ["$player", "$score"], sortBy: { score: -1 }, n: 3 } }
} }
])go deeper
Recognise that a top-per-group result needs grouping plus ordering, and that a plain $limit caps the whole stream rather than each group. Knowing the naive sort-then-group shape is enough here.
Explain that accumulators consume documents in the preceding $sort's order, write the $push plus $slice form correctly, and know that $slice works as an expression, not only in projections.
Lead with the bounded accumulators and justify the choice on memory: per-group state proportional to n rather than to group size. Handle ties explicitly and add a tiebreaker to the sort for reproducibility.
Own the read-pattern decision: if top-N-per-key is a hot path, a maintained leaderboard document or a pre-aggregated rollup may beat recomputing it, and that trades write cost and staleness against query cost.
## The shape of the problem "Top N per group" — three best scores per player, five most recent orders per customer, the latest reading per sensor — is one of the most frequently asked aggregation questions, because it has several correct answers with genuinely different cost profiles and the choice reveals whether the candidate thinks about memory. ## Option 1: $sort then $group with $push and $slice ```javascript db.scores.aggregate([ { $sort: { team: 1, score: -1 } }, { $group: { _id: "$team", docs: { $push: "$$ROOT" } } }, { $set: { docs: { $slice: ["$docs", 3] } } } ]) ``` The accumulator receives documents in the order the previous `$sort` produced, so `docs` is already ranked; `$slice` as an **expression** — `{ $slice: [array, n] }` — takes the first three. `$$ROOT` is the system variable for the whole current document; push only the fields you need if the documents are large. The cost is the point: every document of every group is accumulated before anything is discarded. For a hundred teams of a few dozen players that is nothing. For a sensor collection with a million readings per device, that is a per-group array of a million documents, and the aggregation will strain or fail. If a candidate offers this form, the good follow-up answer is "and I would only use it when group sizes are bounded". ## Option 2: bounded accumulators ($topN, $firstN) MongoDB 5.2 added accumulators that keep only what is needed: ```javascript db.scores.aggregate([ { $group: { _id: "$team", top: { $topN: { output: ["$player", "$score"], sortBy: { score: -1 }, n: 3 } } } } ]) ``` `$topN` carries its own `sortBy`, so no preceding `$sort` stage is needed and the accumulator's state stays bounded at `n` entries per group. `$bottomN` is the ascending counterpart; `$top` and `$bottom` are the single-result forms. `$firstN` and `$lastN` are the order-dependent siblings — they take `input` and `n` and mean "the first n documents that reached this group", which is only meaningful when a `$sort` precedes them. When you have the version available, this is the answer to lead with: same result, bounded memory. ## Option 3: $setWindowFields with $rank ```javascript db.scores.aggregate([ { $setWindowFields: { partitionBy: "$team", sortBy: { score: -1 }, output: { rank: { $denseRank: {} } } } }, { $match: { rank: { $lte: 3 } } } ]) ``` This is the direct analogue of a SQL window function, available from MongoDB 5.0. Its distinguishing merit is explicit tie handling: `$rank` leaves gaps after ties (1, 1, 3), `$denseRank` does not (1, 1, 2), and `$documentNumber` ignores ties entirely and just numbers rows. If the requirement is "everyone tied for third also counts", this is the only one of the three that expresses it directly — the other two silently cut at exactly three documents, choosing arbitrarily among tied candidates. ## Ties, determinism and the sort key Whatever form you pick, a sort on a non-unique key does not define a total order, so two runs can return different members of a tie and pagination over such a result is unstable. Add a tiebreaker to the sort — `{ score: -1, _id: 1 }` — whenever the result feeds anything that must be reproducible. This is the detail that most often separates a solid answer from a strong one. ## Indexes A `$sort` at the front of a pipeline can be served by an index, which avoids a blocking sort of the whole input. A compound index matching the sort — `{ team: 1, score: -1 }` — lets the sort be satisfied by index order. Note that this helps the option-1 form, where the sort is a real leading stage; `$topN` does its own ordering inside the group. Confirm which is actually happening with `explain()` rather than assuming, and remember an index only helps a sort at a point in the pipeline where the original collection order is still available. ## The single-latest special case "Most recent document per key" is the degenerate N=1 version and is worth recognising instantly: `$sort` by key then timestamp descending, `$group` with `{ $first: "$$ROOT" }` — or, on 5.2+, `{ $top: { output: "$$ROOT", sortBy: { ts: -1 } } }`. It appears constantly in time-series and event-log work, and the `$first` form is bounded already, since only one document per group is retained. ## Choosing Bounded accumulators when the version allows and the shape is simple; window fields when ties or ranking semantics matter, or when you need the rank value itself in the output; `$push` plus `$slice` only when group sizes are known to be small or when you must run on an older deployment.
- What is wrong with $push plus $slice on a collection with very large groups?The accumulator materialises every document of every group before `$slice` discards any, so per-group state grows with group size rather than with n. On a million readings per device that is a million-element array per group and the aggregation strains or fails. `$topN`/`$firstN` bound the state at n and are the fix.
- How do the three ranking window operators differ when scores tie?`$rank` gives tied documents the same rank and then skips numbers (1, 1, 3). `$denseRank` gives them the same rank without gaps (1, 1, 2). `$documentNumber` ignores ties and numbers every document distinctly. Pick the one matching the requirement — "include everyone tied for third" needs a rank operator, not a hard $limit.
- How would you return just the most recent document per key?It is top-N with n = 1: `$sort` by key ascending and timestamp descending, then `$group` on the key with `{ $first: "$$ROOT" }`. That form is already bounded, since only one document per group is kept. On MongoDB 5.2+ the equivalent is `{ $top: { output: "$$ROOT", sortBy: { ts: -1 } } }`, which carries its own ordering.
- Why add a tiebreaker to the sort key?A sort on a non-unique field defines only a partial order, so which of several tied documents lands in the top N is arbitrary and can differ between runs, plans or shards. Appending a unique field — `{ score: -1, _id: 1 }` — makes the order total and the result reproducible, which matters for anything paginated or compared over time.
saying these in an interview costs you the question
- Applies a single $limit and thinks it caps each group
- Pushes every document per group regardless of group size
- Assumes $group returns members in sorted order without a preceding $sort
- Ignores ties and calls the arbitrary cut deterministic
- Thinks $slice is only a projection operator, not an expression