A batch job iterating a MongoDB cursor fails with CursorNotFound — what causes that and how do you avoid it?
answer
- The server reclaims state that looks abandoned
- Processing time counts as idle time
- A session ending takes its cursors with it
- Failover destroys cursor position too
- Hold no cursor across slow work
basics
~20 sThe server closed the cursor before the job asked for the next batch — usually the 10-minute idle-cursor timeout, an expired logical session, or a primary stepdown. Fix it by not holding a long-lived cursor: page with a resumable range query instead.
solid answer
~50 sA cursor is server-side state, and the server reclaims it when it looks abandoned. By default an idle cursor is closed after about 10 minutes (controlled by the `cursorTimeoutMillis` server parameter), so a job that spends minutes processing each batch before calling `getMore` finds its cursor gone. Cursors also die when the logical session that owns them expires — drivers open sessions implicitly, and an idle session times out after 30 minutes by default — and when the node steps down, restarts, or the connection drops. The `noCursorTimeout` option suppresses the idle timeout but not session expiry or failover, and it leaks server state if you forget to close the cursor. The robust fix is to stop holding a cursor across slow work: read a bounded page with a range query on `_id`, process it, then issue the next query from the last key.
code
javascript · 3 lines// fragile: cursor sits idle while slowWork runs
const c = db.jobs.find({ status: "pending" });
while (c.hasNext()) slowWork(c.next());go deeper
Know that a cursor is state on the server, not a local list, and that leaving it untouched for too long lets the server close it — so process results promptly.
Explain the mechanics: the roughly 10-minute idle-cursor timeout governed by cursorTimeoutMillis, the logical session that owns the cursor, and the getMore call that fails once the cursor id is gone.
Diagnose it properly — measure the gap between batches, check for a stepdown at that moment — and present the resumable range-query loop rather than reaching for noCursorTimeout or a cluster-wide setting change.
Own the job design standard: long batch jobs should be restartable from a persisted anchor rather than tied to server-side cursor state, so that failover and deploys are non-events rather than incidents.
## Why cursors expire at all When `find()` returns, the server keeps state for the cursor: the query plan's position, its snapshot of the storage engine's view, and resources tied to it. If every client that walked away left a cursor behind, that state would accumulate without bound. So MongoDB reclaims cursors that look abandoned, and the client discovers this as a **CursorNotFound** error on the next `getMore`. ## The three ways a cursor dies under you **1. Idle timeout.** A cursor that is not touched for roughly 10 minutes is closed by a background reaper. The window is set by the `cursorTimeoutMillis` server parameter (600000 ms by default). "Touched" means a `getMore`; time spent processing documents already delivered in the current batch counts as idle. This is the classic failure: a job pulls a batch, does a slow external call per document, and takes longer than the timeout to come back for more. **2. Session expiry.** Modern drivers associate every operation with a logical session, and cursors belong to their session. When a session ends — explicitly, or by expiring after its idle timeout (30 minutes by default, from `localLogicalSessionTimeoutMinutes`) — its cursors are killed with it. This is why `noCursorTimeout` is not an unconditional guarantee: it exempts the cursor from the idle-cursor reaper, not from its session ending. **3. Topology and connection events.** A primary stepdown or election, a `mongod` restart, or a dropped connection all destroy cursor state. Ordinary retryable-write and retryable-read machinery does not resurrect a partially consumed cursor: there is no way to replay the position. On a sharded cluster the router's cursor and the per-shard cursors can be lost the same way. An administrator can also end cursors explicitly with `killCursors` or by killing the session, which is how an operator stops a runaway scan. ## Why noCursorTimeout is the wrong first answer `db.c.find(...).noCursorTimeout()` (the `noCursorTimeout` find option) tells the server not to reap this cursor for idleness. It is real and occasionally correct, but it has three problems: - If the client crashes or forgets to close, the cursor lives until its session ends — server state held for no one. - It does nothing about session expiry or failover, so it converts a reliable 10-minute failure into an intermittent one, which is worse to debug. - It encourages the underlying design mistake: coupling a long-running job's correctness to a single piece of server-side state. If you do use it, close the cursor explicitly in a finally-style block, and keep the session alive for the duration. ## The resumable pattern The pattern that survives timeouts, stepdowns and restarts is to hold no long-lived cursor at all. Read a bounded page anchored on a unique, indexed key, process it fully, then issue the next query from the last key you saw: ```javascript let lastId = null; while (true) { const q = lastId ? { _id: { $gt: lastId } } : {}; const page = db.jobs.find(q).sort({ _id: 1 }).limit(500).toArray(); if (page.length === 0) break; for (const doc of page) slowWork(doc); lastId = page[page.length - 1]._id; } ``` Each iteration is a fresh, short-lived query. Nothing is held open during `slowWork`. If the process dies, persist `lastId` and the job restarts where it stopped — the same property that makes range paging the right answer for HTTP APIs makes it the right answer for batch jobs. Note this is an incremental scan, not a consistent snapshot: documents inserted with a higher `_id` while the job runs will be picked up, and documents already passed will not be revisited. ## Reducing exposure without changing the pattern If you must keep a single cursor, shrink the idle window: smaller batches mean more frequent `getMore` calls, so the cursor is touched more often. Better still, separate fetching from processing — drain the cursor quickly into a bounded queue and do the slow work off it — so the cursor's lifetime is bounded by read speed rather than by processing speed. Do not raise `cursorTimeoutMillis` cluster-wide to paper over one job. It is a global safety valve against leaked cursors; loosening it for everyone to fix one badly shaped job is a poor trade. ## Diagnosing it CursorNotFound (error code 43) surfaces on the `getMore`, so the stack trace points at the iteration, not the original `find` — people often look for a bug in the wrong place. Check how long the job spends between batches, whether a stepdown or restart happened at that time, and whether the batch size makes the gap between `getMore` calls longer than the idle timeout. ## What to say in an interview Name all three causes rather than only the timeout, explain why `noCursorTimeout` is a partial and leaky fix, and present the resumable range-query loop as the design that removes the failure mode instead of postponing it.
- Does the noCursorTimeout option guarantee the cursor stays alive?No. It exempts the cursor from the idle-cursor reaper only. The cursor still dies when its logical session ends — sessions expire after 30 minutes of inactivity by default — and when the node steps down, restarts, or the connection drops. It also leaks server-side state if the client forgets to close the cursor, so use it sparingly and close explicitly.
- Why does the CursorNotFound error appear on getMore rather than on the original find?The `find` succeeded and created the cursor; the failure happens when the client asks for the next batch and the server no longer has that cursor id. The stack trace therefore points at the iteration site, which misleads people into looking for a bug in the loop instead of measuring the gap between batches or checking for a failover.
- Would raising cursorTimeoutMillis be a reasonable fix for one slow job?Rarely. It is a cluster-wide safety valve against leaked cursors, so relaxing it for everyone to accommodate one badly shaped job trades a global protection for a local convenience. Reshape the job instead — process off a bounded queue, or page with a resumable range query — and leave the parameter alone.
- How does the resumable paging loop differ from a cursor in what it guarantees?A cursor gives you one consistent traversal of a single query execution; the paging loop is an incremental scan made of many independent queries. Documents inserted ahead of the anchor while the job runs will be picked up, and passed documents are not revisited. In exchange it survives timeouts, stepdowns and process restarts, since the anchor key is all the state you need.
saying these in an interview costs you the question
- Reaches for noCursorTimeout as the complete fix
- Blames the network without checking processing time per batch
- Thinks a cursor survives a primary stepdown
- Raises the cluster-wide timeout to fix one job
- Looks for the bug at the find() rather than the getMore