What do Prometheus's two retention limits bound, and how do you size a server's disk from them?
answer
- Two independent bounds on the same data
- Whichever fires first does the deleting
- Deletion happens a whole unit at a time
- Deleting a series only marks it first
- Series count divided by interval, times bytes
basics
~20 sTime-based retention bounds how old a sample may be; size-based retention bounds how much disk the database occupies. Whichever is reached first deletes the oldest blocks. Size a disk as samples per second times retention seconds times one to two bytes.
solid answer
~40 s`--storage.tsdb.retention.time` bounds age and defaults to fifteen days; `--storage.tsdb.retention.size` bounds bytes and is off unless you set it. Both may be set, and whichever triggers first removes data. Removal happens at **block granularity** — a block is only dropped once its entire time range has fallen outside the window — so the oldest sample on disk can be older than the nominal retention, and the size bound is best-effort rather than a hard cap. Sizing is arithmetic: samples per second is active series divided by scrape interval, and each stored sample costs roughly one to two bytes once compressed. Multiply by the retention window in seconds, then leave real headroom: compaction writes the merged block before deleting its sources, and the WAL and memory-mapped head chunks share the same volume.
code
bash · 6 lines# time bound plus a size bound as insurance against a cardinality jump
prometheus \
--config.file=/etc/prometheus/prometheus.yml \
--storage.tsdb.path=/var/lib/prometheus \
--storage.tsdb.retention.time=15d \
--storage.tsdb.retention.size=120GBgo deeper
Know that Prometheus keeps data for a limited window on local disk, that the default window is fifteen days, and that it is set by a command-line flag rather than in the configuration file.
Explain both bounds and that whichever is hit first triggers deletion, that deletion happens at block granularity rather than per sample, and be able to do the samples-per-second times seconds times bytes-per-sample arithmetic out loud.
Show the operational judgement: set a size bound as insurance, leave compaction headroom, know that an admin-API delete only writes tombstones, and treat active series count as the leading indicator of a disk problem days before the graph turns.
Own the policy. Decide how long raw data is worth keeping on the servers themselves versus in a long-term tier, what the estate spends on that, and what standard you hold teams to when their series count moves everyone's storage bill.
Prometheus does not tier data, expire individual samples, or run a vacuum. It has two retention bounds and one deletion mechanism, and understanding both is what stops a monitoring server filling its disk at three in the morning. ## The two bounds | Bound | Flag | What it limits | Default | |---|---|---|---| | Time | `--storage.tsdb.retention.time` | how old the oldest sample may be | fifteen days | | Size | `--storage.tsdb.retention.size` | how much disk the database may occupy | unset, so disabled | They are independent and may both be configured. Prometheus enforces each by deleting the oldest data until the bound is satisfied, so in practice **whichever is reached first wins**. Setting only the time bound is the common failure: a cardinality jump doubles the bytes per day and the disk fills long before fifteen days have passed. Setting a size bound as a floor under the time bound is the cheap insurance. ## Why deletion is coarse, and why deleting a series returns nothing Retention is enforced against **whole blocks**. A block is deleted only when its entire time range lies outside the window, so immediately after a compaction has produced a wide block, the oldest sample on disk may be noticeably older than your configured retention. That is expected behaviour, not a bug, and it is another reason to leave headroom. Deleting specific series is a different mechanism with the same coarseness. The admin API endpoint `/api/v1/admin/tsdb/delete_series` — which only exists when the server was started with `--web.enable-admin-api` — does not erase anything. It records **tombstones**: ranges marked deleted, honoured by queries so the data disappears from results immediately. The bytes come back only when the affected block is next rewritten, either by ordinary compaction or by explicitly calling `/api/v1/admin/tsdb/clean_tombstones`. Engineers who delete a runaway metric and then watch the disk graph for relief are usually the ones who learn this in production. The same applies to a metric you simply stop scraping. Its series stay in every block that already contains them until those blocks age out; dropping a noisy job frees space at the pace of retention, not immediately. ## Sizing a volume The arithmetic is short and interviewers do ask for it: 1. **Samples per second** = active series ÷ scrape interval in seconds. 2. **Total samples** = samples per second × retention window in seconds. 3. **Bytes on disk** ≈ total samples × one to two bytes per sample, the range Prometheus's compression typically lands in. Worked through on a real estate: a cheese-ageing inventory platform runs 47 hosts with 12 containers each — 564 scrape targets — and each target exposes about 1,180 series, giving roughly 665,520 active series. At a 15-second scrape interval that is 44,368 samples per second. Over the default fifteen-day window (1,296,000 seconds) that is about 57.5 billion samples, and at 1.7 bytes per sample roughly 98 GB, or about 91 GiB. That number is the **steady-state block data only**. What it does not include: - The write-ahead log and the memory-mapped head chunks, which live on the same filesystem and are sized by the last couple of hours of ingest. - Compaction working space, because the compactor writes the merged block before it deletes the sources. A volume with no slack cannot compact, and a database that cannot compact cannot enforce retention. - Growth. Series counts on a container platform ratchet upward with every deployment that adds a label value. So the honest provisioning answer for the estate above is not 91 GiB but something closer to 150 GiB, with a size-based retention bound set below the volume size so Prometheus sheds data rather than dying. ## What a full disk actually does Prometheus does not degrade gracefully when its volume fills. It cannot write the WAL, cannot cut a block, and cannot compact — and because it cannot compact, it cannot delete. That is why the size bound matters: it is the mechanism that keeps the server inside its own volume without a human. Alert on free space on the data volume as a first-class monitoring-the-monitoring signal, and alert well before the point where compaction can no longer find room, not at ninety-five percent. Two habits separate people who have run this from people who have read about it. First, they set both bounds. Second, they treat a jump in the count of active series as the leading indicator of a disk problem, because that number moves days before the disk graph does.
- Retention is set to fifteen days but the oldest sample on disk is seventeen days old. Is that a bug?No. Retention is enforced by deleting whole blocks, and a block is removed only once its entire time range has fallen outside the window. After compaction has merged several two-hour blocks into a wide one, that block survives until its newest sample ages out, so overshooting the nominal retention by up to a block width is normal.
- How much headroom do you leave above the computed size, and why?Enough that the largest compaction you expect can write its output before the sources are deleted, plus room for the write-ahead log and memory-mapped head chunks on the same volume, plus growth in series count. A rule of thumb of half again the computed steady state is defensible; the real point is that a full volume stops compaction, and stopped compaction stops retention.
saying these in an interview costs you the question
- Thinks size-based retention is a hard per-sample cap
- Expects an admin-API delete to shrink the disk at once
- Sizes the disk from series count while ignoring scrape interval
- Provisions exactly the computed size with no compaction headroom
- Believes retention removes individual samples rather than blocks
- Sets only the time bound and never watches free space