What happens to a MongoDB transaction that runs longer than transactionLifetimeLimitSeconds?
answer
- Sixty seconds by default
- A background task does the aborting
- Think about what an open transaction is holding
- Locks and a pinned snapshot cost cache
- Raising the knob is rarely the fix
basics
~10 sThe server aborts it. transactionLifetimeLimitSeconds defaults to 60, and a periodic cleanup aborts any transaction older than that; subsequent operations and the commit then fail. The fix is smaller transactions, not a bigger limit.
solid answer
~50 sMongoDB deliberately caps transaction lifetime: `transactionLifetimeLimitSeconds` defaults to 60 seconds, and a background job aborts transactions that exceed it, after which further operations on that transaction and its commit fail. The reason is resource pressure — an open transaction holds a storage-engine snapshot and its locks, which pins older document versions in cache and blocks conflicting writers. A related default keeps things moving: a transaction waits only briefly to acquire a lock (`maxTransactionLockRequestTimeoutMillis` defaults to 5 ms) and then fails fast with a write conflict labelled as transient, so a retry can start clean instead of queueing. Raising the lifetime limit trades that pressure for tolerance and is rarely the right answer; the documented guidance is to keep a transaction to a modest number of documents — around 1,000 or fewer — and to break bulk work into resumable batches instead.
code
javascript · 10 lines// Server parameter, default 60
db.adminCommand({ getParameter: 1, transactionLifetimeLimitSeconds: 1 });
// Find transactions currently open and how long they have been running
db.getSiblingDB("admin").aggregate([
{ $currentOp: {} },
{ $match: { "transaction": { $exists: true } } },
{ $project: { opid: 1, "transaction.parameters.txnNumber": 1,
"transaction.timeOpenMicros": 1 } }
]);go deeper
Know that MongoDB transactions are not allowed to run indefinitely: the default lifetime is 60 seconds and the server aborts anything older, so transactions are meant to be short.
Explain why the cap exists — an open transaction pins a storage-engine snapshot and holds locks — and that conflicting transactions fail fast with a write conflict rather than queueing.
Diagnose it in production: find the long-open transaction with currentOp, get non-database work out of the block, split bulk changes into idempotent resumable batches, and explain why a huge commit spikes replication lag.
Set the guardrails — what a transaction is allowed to be used for, an agreed size ceiling, and a review rule that treats raising the lifetime limit as an escalation requiring justification rather than a routine tuning step.
## The limit and what it does `transactionLifetimeLimitSeconds` is a server parameter with a default of 60 seconds. A periodic server task looks for transactions whose lifetime has exceeded it and aborts them. From the application's point of view, the next operation you send on that transaction — or the commit — fails, and the work is gone. Nothing is partially applied: abort is abort. The driver side has its own clock too. `maxCommitTimeMS` bounds how long the commit may take, and `withTransaction()` gives up retrying after roughly two minutes. A transaction can therefore end because your callback was slow, because the commit was slow, or because the server's lifetime limit expired first. ## Why MongoDB caps it at all An open transaction is not free while it sits there: - It holds a **storage-engine snapshot**. Every document version that existed when the snapshot was taken must remain reachable for as long as the transaction lives, so cache and history accumulate instead of being reclaimed. - It holds **locks** on everything it has written. Other transactions that touch the same documents cannot proceed. - Its writes are **staged** and become visible only at commit, so the longer it runs the larger the eventual burst. Without a cap, one forgotten transaction — a client that opened one and then blocked on a slow HTTP call — would degrade the whole node. The 60-second default is the server refusing to let a single misbehaving client do that. ## Fast failure on conflict, by design Related and often conflated: `maxTransactionLockRequestTimeoutMillis` defaults to 5 milliseconds. A transaction that finds a document locked by another transaction does not queue behind it for long; it fails quickly with a write conflict, which the driver sees carrying the transient-transaction label, and `withTransaction()` re-runs the whole thing. MongoDB prefers a cheap retry over a long wait. The operational signature of a hot document under transactional load is therefore not slow queries but a rising retry rate — and eventually the retry limit being hit. ## The oplog story Older MongoDB versions required an entire transaction to fit in a single oplog entry, which imposed an effective 16 MB ceiling on the transaction's total change set. That constraint is gone: since MongoDB 4.2 a large transaction is written as as many oplog entries as it needs. The 16 MB BSON limit still applies to each individual document, but it no longer caps the transaction. What replaces it is a softer, more operational limit. All of a large transaction's oplog entries become visible together at commit, so committing a very large transaction produces a burst of replication traffic. Secondaries apply it as a unit, lag spikes, and the majority-commit point — which every `w: "majority"` write in the system waits on — moves later. A single 200,000-document transaction can therefore be felt by unrelated traffic. MongoDB's own guidance is to keep a transaction to about 1,000 documents or fewer. ## What to do instead **Shrink the transaction.** Most oversized transactions are doing bulk work that has no cross-document invariant at all — a backfill, a status sweep. That work does not need a transaction; it needs an idempotent, resumable batch loop with a checkpoint (the last processed `_id`), so a failure resumes rather than restarts. **Move work out of the transaction.** Anything that is not a database operation — reading a file, calling a service, computing a large payload — should happen before the transaction opens. A transaction should be a tight burst of database operations, not the scope of a business workflow. **Model the invariant into one document** where you can, so the atomic write is free and there is no lifetime to exceed. **Only then consider the knob.** Raising `transactionLifetimeLimitSeconds` is legitimate for a specific, understood workload — a controlled migration on a maintenance window, for instance — but as a response to "our transactions time out" it converts a fast, visible failure into slow cache pressure and long lock holds that are much harder to diagnose. ## Diagnosing it The error surfaces as the transaction having been aborted or no longer existing. Look at how long the callback actually takes end to end, whether anything non-database sits inside it, how many documents it touches, and whether the abort rate correlates with contention on a small set of hot documents. `currentOp` shows transactions in progress with their start time, which is the quickest way to catch a client holding one open. ## What an interviewer is testing Whether you treat the limit as an obstacle or as a signal. The good answer states the default, explains the snapshot-and-lock cost that motivates it, mentions that the old 16 MB oplog ceiling is gone but that big commits still hurt replication, and lands on restructuring the work rather than raising the parameter.
- Why does MongoDB fail a conflicting transaction after a few milliseconds rather than making it wait?Because waiting compounds the cost. A transaction that queues for a lock keeps its own snapshot and locks alive while it waits, so contention cascades. The default lock request timeout of 5 ms makes the loser fail fast with a write conflict carrying the transient-transaction label, and the driver re-runs the whole transaction from a fresh snapshot. Under contention you see retries rather than a growing queue.
- Is a MongoDB transaction still limited to 16 MB by the oplog?No. That was true before MongoDB 4.2, when a transaction had to fit in a single oplog entry. Since 4.2 a large transaction is written as as many oplog entries as it needs, so the practical limits are time, locks and resource pressure rather than 16 MB. The 16 MB BSON limit still applies to each individual document. Very large commits remain a bad idea for replication reasons, not because of a hard cap.
- What is the replication cost of committing a transaction that modified 200,000 documents?Its oplog entries become visible together at commit, so secondaries receive and apply a large batch as a unit. Replication lag spikes, and the majority-commit point advances late — which delays every `w: "majority"` write in the system, including ones from unrelated workloads. That is why the guidance is roughly 1,000 documents per transaction and why bulk work belongs in resumable batches instead.
saying these in an interview costs you the question
- Says the transaction just keeps running until the client commits
- Believes a transaction is still capped at 16 MB of oplog
- Reaches straight for raising transactionLifetimeLimitSeconds
- Puts an external API call inside the transaction and blames MongoDB
- Thinks a conflicting transaction waits indefinitely for the lock