Why does time-window compaction suit time-series data with a fixed expiry in a log-structured store, and what kinds of writes break it?
answer
- one file per window
- closed windows stay closed
- drop whole expired files
- old data in new windows
- mixed expiry, late writes
basics
~20 sIt merges each time window's data into one file that is never compacted again, so a fully expired window is dropped as a whole file. Back-dated writes, updates to old data and mixed expiries mix windows and defeat it.
solid answer
~50 sTime-series data is mostly **written once, in time order, and expires after a fixed period**. Time-window compaction exploits that: files are grouped into **windows** (for example a day); inside the current window files are merged normally; once the window closes, its files are merged into **one file that is never compacted again**. When every cell in that file has expired, the store can **drop the whole file** instead of rewriting data and purging tombstones one by one. That keeps write amplification very low and makes expiry cheap. It breaks when files mix old and new data: **back-dated or out-of-order writes** (backfills, client-supplied timestamps), **updates or deletes** to old windows, **different expiry periods** in one table, and background repair that copies old data into the current window. A file can also be dropped only if it does not hide older data elsewhere.
go deeper
Know that time-window compaction groups data by time and drops whole files once all their data has expired.
Explain the one-file-per-closed-window mechanism, how to size windows, and why expiry becomes cheap.
Identify writes that mix windows, such as backfills, updates, mixed expiry and repair, and redesign tables to avoid them.
Be ready to set retention and schema conventions for time-series data so that expiry stays cheap as the platform grows.
## The workload it is built for Many time-series tables share a shape: - rows are **appended in roughly time order** (readings, events, logs); - they are **rarely updated or deleted** individually; - they **expire** after a fixed retention period; - reads ask for recent time ranges. General-purpose strategies handle this poorly. They repeatedly rewrite data that will never change, and expiry produces tombstones that must be merged away row by row. ## How time-window compaction works 1. Each file is assigned to a **window** by the timestamps of its data (an hour, a day…). 2. Inside the **current** window, newly flushed files are merged with a normal (typically size-tiered) policy. 3. When a window **closes**, its files are merged into a **single file**, and that file is **not compacted again**. 4. When **all** data in a window's file has expired, the file is **dropped whole**, provided it cannot be hiding older data in other files. Some range-served stores provide a closely related, date-aware strategy that tiers files by age for the same purpose. ## Why it is efficient | concern | general strategy | time-window | |---|---|---| | rewriting immutable old data | repeated as files merge | once, when the window closes | | expiry | per-row markers merged away | delete whole files | | reads of recent data | may touch old files | recent windows only, if files carry time bounds | Choosing window size: aim for a manageable number of windows across the retention period — for example, with a 90-day expiry, windows of a few days rather than hours. ## What breaks it The strategy assumes a file holds data from **one** time window. Anything that puts old and new data in the same file undermines it: - **Back-dated or out-of-order writes**: a backfill, or clients supplying their own timestamps, write old data into the current memory buffer, which flushes into a current-window file. - **Updates and deletes to old data**: they land in today's files but refer to old windows, so an old file cannot be dropped while it still holds data a newer marker must hide, and the marker must stay while old data exists. - **Mixed expiry periods**: if some rows live 7 days and others 365, a window's file is never fully expired, so it is never dropped. - **Repair copying old data**: in replicated stores, a repair that writes old data through the normal path mixes it into current files. - **Changing the window size** later: existing files are not re-split. ## Design rules that keep it working 1. Use it only for **append-mostly** data with a **single retention** period per table. 2. Write with **server timestamps**, or at least timestamps close to real time. 3. Route **backfills** and corrections to a separate table, or accept they will linger. 4. Put a **time bucket in the key** as well, so partitions stay bounded and reads target recent windows. ## Interview angle Explain the one-file-per-closed-window mechanism, why dropping whole files makes expiry cheap, and name the specific writes that mix windows.
- Why can a fully expired file sometimes not be dropped?If it contains tombstones that hide older data in other files, dropping it would let that older data reappear. The store must check that nothing older is shadowed before dropping it, or wait until the older files are gone too.
- How should a historical backfill be loaded into a time-window compacted table?Ideally into a separate table or with a process that writes each window's data so it flushes into its own files. Streaming years of old rows through the normal path mixes them into current-window files that will not expire cleanly.
saying these in an interview costs you the question
- Using time-window compaction for data that is frequently updated
- Mixing rows with different expiry periods in one time-window table
- Letting clients write arbitrary back-dated timestamps into it
- Believing expired rows are always removed exactly at their expiry time