skip to content

What does cursor.batchSize() change about how a MongoDB find() returns documents?

level: middleimportance: should knowfreq 44%

answer

  1. One knob is transport, the other is result size
  2. Documents arrive in groups, not all at once
  3. The first group has a familiar default count
  4. Message size caps every later group
  5. limit() is what actually bounds the total

basics

~20 s

batchSize() sets how many documents the server puts in each network batch, not how many the query returns in total. It tunes round trips against per-batch memory; limit() is what caps the total number of documents.

solid answer

~40 s

`find()` does not send all matching documents at once — it returns a cursor. The first reply carries an initial batch (by default up to 101 documents, or less if 16 MB is reached first), and the driver fetches later batches with `getMore` as you iterate, each batch capped by the 16 MB maximum BSON message size. `batchSize(n)` changes documents per batch: smaller batches mean more round trips but less memory held per fetch and a shorter window before the cursor is touched again; larger batches mean fewer round trips but more memory on both sides. It does **not** limit the result set — `find().limit(5).batchSize(10)` yields five documents. Most applications should leave it alone and only tune it when documents are unusually large or per-document processing is slow.

code

javascript · 7 lines
javascript
// batchSize tunes round trips; limit bounds the result set
db.events.find({ type: "click" }).batchSize(500).forEach(doc => process(doc));

db.events.find({ type: "click" }).limit(5).batchSize(10);  // returns 5 docs

// aggregate has its own cursor batchSize option
db.events.aggregate([{ $match: { type: "click" } }], { cursor: { batchSize: 200 } });

go deeper

for a junior

Know that find() gives you a cursor you iterate, not a finished list, and that limit() — not batchSize() — is how you ask for fewer documents.

for a middle

Explain the mechanics: an initial batch of up to 101 documents, getMore for the rest, each batch bounded by the 16 MB message size, and batchSize as a round-trips-versus-memory knob.

for a senior

Give the concrete reasons you would ever touch it — very large documents, or slow per-document processing that leaves a cursor idle — and say plainly that the usual answer is to leave the default and bound the query instead.

for a principal

Own the pattern choice: long-lived cursors couple a job's runtime to server-side state. Decide when a workload should stream through a cursor at all versus resume through re-issued bounded queries.

## find() returns a cursor, not a result set When you run `db.c.find(filter)`, the server does not materialize and ship every matching document. It creates a **cursor** — server-side state identifying the query and its progress — and returns a first batch of documents plus a cursor id. The driver then issues `getMore` commands against that id to pull further batches, transparently, as your loop advances. When the cursor is exhausted the server returns a cursor id of zero and releases the state. This streaming design is what lets a query over a billion documents run in constant client memory. It is also why cursors have lifecycle concerns — batches, timeouts, and explicit closing — that a "give me the rows" API does not. ## Default batching For a `find()`, the first batch contains up to 101 documents, or fewer if the batch would exceed the 16 MB maximum BSON message size. Subsequent `getMore` batches are not capped at 101; they are bounded by that 16 MB message size. So with small documents you get a modest first batch and then much larger ones; with 1 MB documents you get roughly a dozen or so per batch regardless of counts. These defaults are deliberately chosen so that the common case — a query returning a screenful of results — completes in one round trip, while a large scan streams efficiently. ## What batchSize actually controls `cursor.batchSize(n)` asks the server to put at most `n` documents in each batch, including the first. It is purely a **transport tuning knob**: - **Smaller batches**: more `getMore` round trips, therefore more latency overhead on a slow or distant link, but less memory materialized per fetch and a smaller amount of work discarded if the client stops early. - **Larger batches**: fewer round trips and better throughput on a fast link, at the cost of holding more documents in memory on both server and client per fetch, and a longer pause between touching the cursor. The 16 MB message ceiling still applies: a `batchSize` larger than what fits is silently satisfied with fewer documents. ## batchSize is not limit This is the confusion the question exists to catch. `limit(n)` bounds the **total** number of documents the query will ever return; `batchSize(n)` bounds documents **per batch**. `db.c.find().limit(5).batchSize(10)` returns five documents in total. `db.c.find().limit(1000).batchSize(10)` returns a thousand documents across roughly a hundred round trips. When a `limit` is smaller than the batch size, the server sends only up to the limit and can close the cursor immediately, which is why a small `limit` is the right way to say "I only want a few" — never a small `batchSize`. ## When tuning is justified Leave the default alone unless you have a specific reason. Real ones: - **Large documents.** If documents are hundreds of kilobytes, a default batch can be a big allocation on both sides. A smaller `batchSize` smooths memory use. - **Slow per-document processing.** If each document triggers an external call taking seconds, a large batch means a long gap between `getMore` calls, and the cursor sits idle on the server — which interacts badly with the idle-cursor timeout. Smaller batches touch the cursor more often, though the more robust fix is to stop holding a long-lived cursor at all and page with a range query instead. - **Low-latency first result.** A UI that wants to render the first rows fast may prefer a small first batch, though a `limit` usually expresses that better. - **Aggregation cursors.** The aggregate command has its own `cursor: { batchSize: n }` option with the same meaning; note that `{ batchSize: 0 }` there means "return the cursor without any documents in the first reply", which is used to start a pipeline and stream everything through `getMore`. ## Iteration in practice In `mongosh`, iterating with `hasNext()`/`next()` or `forEach()` walks the cursor and fetches batches automatically; `toArray()` drains it fully into client memory, which defeats the point of streaming for a large result. Drivers expose the same duality — an async iterator that streams, and a `toArray`-style helper that buffers. Choosing `toArray()` on an unbounded query is a far more common production incident than any `batchSize` mistake. Cursors should be closed when you stop early. Driver iterators close on exhaustion or when their scope ends; if you break out of a loop manually, close explicitly so the server can drop the state rather than waiting for the idle timeout. ## What to say in an interview Say that `find()` streams through a cursor, that the first batch defaults to 101 documents with later batches bounded by the 16 MB message size, that `batchSize` moves documents per round trip while `limit` caps the total, and that the honest default advice is not to tune it without a measured reason.

  • How many documents does the first batch of a find() contain by default?
    Up to 101 documents, or fewer if that many would exceed the 16 MB maximum BSON message size. Subsequent `getMore` batches are not capped at 101 — they are bounded by that message size — so with small documents later batches are considerably larger than the first.
  • What does db.c.find().limit(5).batchSize(10) return?
    Five documents. `limit` bounds the total result set; `batchSize` only controls how many documents travel per network round trip. Since the limit is smaller than the batch size, everything arrives in one reply and the server can close the cursor immediately.
  • When is calling toArray() on a cursor the wrong choice?
    Whenever the result set is unbounded or large. `toArray()` drains the whole cursor into client memory, which discards the streaming property that makes cursors safe and is a routine cause of out-of-memory incidents. Iterate the cursor instead, or bound the query with a `limit` or a range predicate.

batchSize is the size of each trayload a waiter carries from the kitchen; limit is how many dishes you actually ordered. Bigger trays mean fewer trips, not more food.

saying these in an interview costs you the question

  • Says batchSize limits how many documents the query returns
  • Thinks find() ships the entire result set at once
  • Believes every batch is capped at 101 documents
  • Sets a huge batchSize as a general performance tweak
  • Calls toArray() on an unbounded query

context