How does time-to-live expiry work in a log-structured store, and why does expired data keep using disk space and read time until compaction?
answer
- expiry is a stamp, not a timer
- checked when read or merged
- hidden now, removed later
- some stores still return it
- whole files expire together
basics
~20 sA time-to-live stamps each cell or column group with an expiry; nothing is deleted at that moment. Expired cells are skipped when read or merged, and their space returns only when compaction rewrites or drops the files holding them.
solid answer
~50 sA **time-to-live (TTL)** is set per write, per row, or as a default for a table or column family, and records when data expires. The store does **not** run a timer that deletes it. Instead, expiry is a property checked later: reads compare each cell's expiry with the current time, and compaction drops expired cells when it rewrites their files. So expiry is **deferred deletion**. Until compaction reaches the files, expired cells still occupy disk and reads may still have to step over them. Stores differ in visibility: many hide expired cells from reads immediately, while some garbage-collection schemes may keep returning data due for removal until background collection has run, so reads should filter it. In replicated stores with peer repair, an expired cell may also be kept as a delete marker for the grace period. Whole files whose data has all expired can be dropped outright, which is why TTL pairs well with time-grouped files.
go deeper
Know that a TTL makes data expire automatically and that disk space is reclaimed later, not at the expiry moment.
Explain lazy expiry at read and compaction time, the visibility difference between stores, and why expired data still costs reads.
Design tables so expiry drops whole files, with one TTL per table and time in the key, and size disk for the lag.
Be ready to map legal retention and erasure deadlines onto TTL, compaction lag and backup retention with margins.
## What a TTL is A **time-to-live** attaches an expiry to data: - per write (every cell written by this mutation expires in N seconds), - per row, or - as a **default** for a whole table or column family. Some stores express the same idea as a **garbage-collection rule** on a column family, such as "keep versions younger than 30 days". ## Expiry is checked, not scheduled A log-structured store does not set a timer per cell. Expiry is evaluated **lazily**: 1. **At read time**: when a read merges entries, it compares each cell's expiry to the current time and treats expired cells as absent — in stores that filter on read. 2. **At compaction time**: when files are rewritten, expired cells are dropped from the output (or, in some replicated stores, turned into delete markers kept for a grace period before being purged). So TTL is best described as **deferred deletion**: the data becomes logically invisible, then physically disappears whenever compaction gets to it. ## Visibility differs between stores | behaviour | what a read sees after expiry | |---|---| | filter on read | nothing; the cell is skipped immediately | | collect in background | the data may still be returned until collection runs; reads should filter by age themselves | Knowing which applies matters when retention is a correctness or compliance requirement. ## Why expired data still costs - **Disk**: the bytes stay in their files until those files are compacted. - **Reads**: a read over a range full of expired cells must still read and skip them. - **Compaction**: rewriting files just to drop expired cells is itself I/O. - **Markers**: where expired cells become delete markers, those markers add their own read cost until purged. ## Making expiry cheap 1. **Group data by time**: if a file holds only data written in one period with one TTL, the whole file expires together and can be **dropped without rewriting** — the basis of time-window style strategies, and of stores that delete files containing only expired rows. 2. **Use one TTL per table**: mixed TTLs keep files partly alive forever. 3. **Prefer TTL over explicit deletes** for bulk ageing-out: the expiry is written once with the data instead of a later delete per row. 4. **Keep a time component in the key**, so reads of recent data do not scan expired ranges. ## Interview angle The key sentence is "TTL is deferred deletion": invisible at read or collection time, reclaimed at compaction. Add the visibility difference between stores and how time-grouped files make expiry nearly free.
- Why is a table default TTL often better than setting TTL on each write?It guarantees every row expires, even from writers that forget, and keeps expiry uniform, so time-grouped files expire together. Per-write TTLs invite mixed lifetimes that keep files partly alive.
- A compliance rule says data must be gone 30 days after collection. Is a 30-day TTL enough?Not by itself. Expired data can remain on disk, in replicas and in backups until compaction and backup rotation remove it. The TTL must be shorter than the deadline by the worst-case compaction and backup lag.
saying these in an interview costs you the question
- Believing expired data is deleted from disk at the exact expiry moment
- Assuming every store hides expired data from reads immediately
- Mixing very different TTLs in one table and expecting files to expire cleanly
- Treating a TTL as satisfying a hard erasure deadline on its own