How does a MongoDB TTL index with expireAfterSeconds actually delete expired documents?
answer
- A background thread, not a per-document timer
- Runs on a cycle measured in tens of seconds
- The removals are ordinary deletes
- Only one member of a replica set does the work
- The field must really hold a date
basics
~20 sA background thread wakes roughly every 60 seconds, finds documents whose indexed date is older than expireAfterSeconds, and deletes them with ordinary delete operations. Expiry is therefore approximate, not instant, and only the primary performs the deletes.
solid answer
~50 sYou create one with `db.sessions.createIndex({ createdAt: 1 }, { expireAfterSeconds: 3600 })`. A background TTL monitor runs on roughly a 60-second cycle, scans the index for entries whose date plus `expireAfterSeconds` is in the past, and issues **normal delete operations** for them. Consequences worth stating: removal is eventual, so a document can linger a minute or more past its deadline and your queries must still filter on the date if correctness depends on it; the deletes consume write throughput, generate oplog entries, and can compete with your workload on a big backlog. On a replica set only the **primary** expires documents — secondaries just apply the replicated deletes. The index must be single-field on a BSON date (or an array of dates, where the earliest wins); a missing or non-date value never expires. `collMod` changes `expireAfterSeconds` without rebuilding.
code
javascript · 2 lines// Collection-wide lifetime: one hour after createdAt
db.sessions.createIndex({ createdAt: 1 }, { expireAfterSeconds: 3600 })go deeper
Recall the createIndex call with expireAfterSeconds, that the field must hold a date, and that MongoDB removes the documents for you some time after they expire.
Explain the background monitor and its roughly one-minute cycle, why expiry is approximate, the expireAfterSeconds: 0 per-document pattern, and the single-field restriction.
Show that TTL removals are ordinary writes competing with your workload, that only the primary performs them, and how you would shorten a TTL on a huge collection without causing a delete storm.
Own retention as a policy question: what the data lifecycle should be, whether expiry or archival to cheaper storage is right, and how the delete volume shapes oplog sizing and cluster capacity.
## The mechanism A TTL index is an ordinary single-field index plus one option: ```javascript db.sessions.createIndex({ createdAt: 1 }, { expireAfterSeconds: 3600 }) ``` Nothing about the *storage* of the document changes. What changes is that a background thread in `mongod` — the TTL monitor — now has a job to do. It wakes on a fixed cycle, by default about every 60 seconds, walks the TTL index looking for entries whose indexed date plus `expireAfterSeconds` lies in the past, and deletes those documents. The deletes are completely ordinary: they take the same locks, produce the same oplog entries, and fire the same change-stream events as a delete your application issued. That single fact explains almost every TTL question an interviewer will ask. ## Expiry is approximate Between the moment a document becomes eligible and the moment it disappears there is at minimum the remainder of the current monitor cycle, plus however long the deletion pass takes. Under load, or when a large backlog has accumulated (say you shortened the TTL, or the server was down for a day), the lag can be much larger than a minute. The engineering consequence is that **you cannot treat the collection as if expired documents are gone**. If a session must stop being valid at exactly the one-hour mark, the read path still has to check the timestamp; the TTL index is a space-reclamation mechanism, not an access-control mechanism. Candidates who answer "the document is removed the instant it expires" are usually the ones whose sessions stay usable for an extra minute in production. ## The deletes cost real write throughput Because they are normal deletes, TTL removals consume the same resources as application writes: index maintenance for every index on the collection, oplog space, replication bandwidth, and cache churn. Expiring millions of documents at once is a genuine load event. Two practical rules follow. First, be careful about *shortening* a TTL on a large collection — you have just queued a very large delete storm. Second, watch for a workload where the TTL monitor never catches up because the insert rate exceeds the delete rate; the collection grows without bound even though the index is configured correctly. A related surprise: freed space returns to the storage engine but does not necessarily shrink files on disk, so a TTL collection can hold steady in document count while occupying much more disk than the live data suggests. ## What actually expires Rules that get asked directly: - The index must be **single-field**. A compound index cannot carry `expireAfterSeconds`. - The indexed field must hold a **BSON date**. A string like `"2026-01-01"`, a number, or a missing field means the document is never expired. This is the single most common bug: the app stored an ISO string and the collection never shrinks. - If the field holds an **array of dates**, the earliest date in the array determines expiry. - TTL indexes are not supported on capped collections, since MongoDB cannot remove arbitrary documents from one. - The `_id` index cannot be made a TTL index. ## The expireAfterSeconds: 0 pattern Setting `expireAfterSeconds: 0` means "expire at the time stored in the field" rather than "expire immediately". You write an explicit `expiresAt` date into each document and the monitor removes it once the clock passes that value. This is how you get per-document lifetimes — a two-hour password-reset token and a thirty-day export in the same collection — instead of one collection-wide duration. ```javascript db.tokens.createIndex({ expiresAt: 1 }, { expireAfterSeconds: 0 }) db.tokens.insertOne({ userId: 7, expiresAt: new Date(Date.now() + 2 * 3600 * 1000) }) ``` ## Replica sets Only the primary runs TTL deletions. Secondaries keep the TTL index but do not act on it independently; they receive the deletes through the oplog like any other write. This preserves the ordering guarantee that makes replication work — if each member expired documents on its own clock, members would diverge. It also means that during a period where a member lags, expired documents remain readable on that secondary until it catches up, which matters if you serve reads from secondaries. ## Changing and monitoring You can change the duration in place without rebuilding the index: ```javascript db.runCommand({ collMod: "sessions", index: { keyPattern: { createdAt: 1 }, expireAfterSeconds: 86400 } }) ``` Server status exposes counters for how many documents the TTL monitor has deleted and how many passes it has made; those are the metrics to alert on when you suspect the monitor is falling behind. And because a TTL index is a real index, it still serves ordinary range queries on that date field — you do not need a second index for "documents created in the last hour".
- A collection has a TTL index on createdAt but never shrinks. What do you check first?Whether the field actually holds BSON dates. If the application wrote an ISO-8601 string or an epoch number, the TTL monitor skips those documents entirely and the collection grows forever. After that, check that the index really carries expireAfterSeconds, that you are looking at the primary, and that the insert rate is not simply outrunning the deletion pass.
- How do you give each document its own lifetime instead of one collection-wide duration?Create the index with expireAfterSeconds: 0 on a field such as expiresAt, and write the absolute expiry timestamp into each document. Zero means expire when the stored date passes, not expire immediately, so a two-hour token and a thirty-day export can live in the same collection with different deadlines.
- Why should you be cautious about shortening expireAfterSeconds on a very large collection?Because it makes a huge number of documents eligible at once, and the TTL monitor will delete them as ordinary writes — taking locks, maintaining every index on the collection, and filling the oplog. That can starve application traffic and blow out replication lag. Step the value down gradually, or delete in controlled batches yourself.
It is a janitor who walks the building about once a minute and empties the bins that are past their pickup time — not a self-destruct timer wired into each bin.
saying these in an interview costs you the question
- Says documents vanish at the exact expiry instant
- Thinks TTL deletion bypasses the oplog or replication
- Assumes each replica set member expires its own copy
- Believes expireAfterSeconds: 0 deletes documents immediately
- Tries to put expireAfterSeconds on a compound index