skip to content

Storage and Operations

Running Prometheus at size: TSDB blocks and the WAL, retention and compaction, then the moment local storage is not enough and you reach for remote_write into Thanos, Cortex or Mimir. Interviewers ask how you get long retention and a global view out of a system designed to be single-node.

on this pageshow

questions

4

In Prometheus's local TSDB, how does a scraped sample travel from the WAL to a compacted block, and what does a restart replay?

level: middleimportance: must knowfreq 58%

answer

  1. Two destinations for every scraped sample
  2. One of them is only for recovery
  3. The recent window lives in memory
  4. Every two hours something is cut to disk
  5. Start-up has to rebuild the head

basics

~20 s

Every scraped sample is appended to both the write-ahead log and the in-memory head. Roughly every two hours the head's oldest range is cut into an immutable block, which compaction later merges. A restart replays the log to rebuild the head.

solid answer

~40 s

Prometheus writes each sample twice: into a write-ahead log segment under `wal/`, purely for crash recovery, and into the in-memory **head**, which is what recent queries read. Completed head chunks are written out and memory-mapped from `chunks_head/`, so the process does not hold every raw sample in RAM. About every two hours the oldest completed range of the head is cut into an immutable block directory named by a ULID, holding `chunks/`, an `index`, `meta.json` and `tombstones`; the WAL is then checkpointed and superseded segments are deleted. A background compactor merges adjacent blocks into wider ones with a single index. On start-up Prometheus replays the WAL and reloads the memory-mapped head chunks to rebuild the head, which on a large server is minutes before it will answer a query.

code

text · 14 lines
text
data/
├── 01HQ8Z3K9VQ7X4M2B6WN5TCFD0/   # an immutable block, named by ULID
│   ├── chunks/
│   │   └── 000001
│   ├── index
│   ├── meta.json
│   └── tombstones
├── chunks_head/                   # memory-mapped head chunks
│   └── 000001
├── wal/
│   ├── checkpoint.00000270/
│   └── 00000271
├── lock
└── queries.active

go deeper

for a junior

Be ready to say that Prometheus keeps its own data on local disk in a directory it owns, that recent data is in memory and older data is in files, and that a write-ahead log exists so a crash does not lose the recent data.

for a middle

An interviewer expects you to explain the mechanics: sample goes to both the log and the in-memory head, the head is cut into an immutable block roughly every two hours, the log is checkpointed and truncated afterwards, and compaction merges blocks into wider ones.

for a senior

Show that you have operated this. Talk about restart cost scaling with head size, staggering restarts across replicas, leaving disk headroom because compaction writes before it deletes, and watching head series count as the leading indicator of memory trouble.

for a principal

Own the consequence of the single-node design: block immutability is what makes shipping blocks elsewhere possible, and head size is what caps how big one server may get. Be able to say where you would draw the line between a bigger server and a different architecture.

Prometheus stores metrics in its own embedded time-series database rooted at the directory given by `--storage.tsdb.path` (`data/` by default). There is no external database and no clustering: one Prometheus process owns that directory exclusively, and almost everything about operating Prometheus at size follows from that fact. ## Two destinations for every sample When a scrape completes, each sample — a metric name plus its label set, one timestamp and one value — goes two places at once. 1. **The write-ahead log.** The sample is appended to a segment file under `wal/`. The WAL is append-only and sequential, so writing to it is cheap. It exists only so that an unclean shutdown does not lose acknowledged data; **no query ever reads the WAL**. 2. **The head.** The same sample is appended to the head, the mutable, most-recent part of the database that lives in memory. Every query over recent time is served from the head. Inside the head, samples for one series accumulate into a chunk. A chunk is closed when it reaches its sample limit or when its time range elapses; closed head chunks are written out and memory-mapped from `chunks_head/`, so the process keeps chunk metadata resident but not every raw value. This is the reason head memory tracks the count of **active series** far more strongly than it tracks the scrape interval: each series carries fixed per-series overhead whether it is sampled every fifteen seconds or every minute. ## Cutting a block, then compacting it The head only covers the most recent window. Roughly every two hours Prometheus cuts the oldest completed range of the head into a **persistent block** — a directory named with a ULID containing: | Entry | What it holds | |---|---| | `chunks/` | the compressed sample data, in numbered segment files | | `index` | the inverted index from label pairs to series, and from series to chunks | | `meta.json` | the block's time range, series and sample counts, and its compaction level | | `tombstones` | ranges marked for deletion but not yet physically removed | A block is **immutable**. Nothing is written into it after it is cut. That is what keeps the design simple: readers never take a lock, and a block can safely be copied or uploaded elsewhere while Prometheus is running. Once a block is durable on disk, the WAL no longer needs the records that produced it. Prometheus writes a **checkpoint** beside the segments — a compacted directory holding the series records and any samples still required — and deletes the segments it supersedes. That truncation is the only thing that stops the WAL growing without bound, and it is why WAL size is roughly a function of ingest rate over the last couple of hours rather than of retention. A background **compactor** then merges adjacent blocks into wider ones. Compaction is not tidying for its own sake; it pays for itself three ways: - One index over a wide range is far cheaper to search than many small indexes covering the same range. - Series that were deleted, or that simply stopped being scraped, drop out of the merged index instead of being carried forward in every small block. - Retention can then be enforced by deleting whole blocks, which is a directory removal rather than a rewrite. The compactor writes the merged block first and removes the sources afterwards, so a volume sized to exactly the steady-state data will eventually wedge with no room to compact. ## What a restart actually costs On start-up Prometheus reloads the memory-mapped head chunks and replays the remaining WAL segments to rebuild the head exactly as it was. Only when that finishes does the server begin scraping and answering queries. Replay time is driven by how much data the head holds, which in turn is driven by active series count — so the servers that most need to be restarted quickly are the ones slowest to come back. Practical consequences engineers are expected to draw from this: - **Restarts, not writes, are the expensive operation.** A config reload is cheap; a process restart on a multi-million-series server is minutes of blindness. - **Run replicas, and never restart them together.** Two identically configured servers scraping the same targets let you take one down without a gap. - **Give the volume headroom.** WAL, memory-mapped head chunks, blocks and compaction output all share one filesystem. - **`--storage.tsdb.wal-compression` trades a little CPU for smaller WAL segments**, which means fewer bytes to read back on replay. - **The head is where cardinality hurts first.** A sudden flood of new series inflates head memory immediately, and inflates replay time on the next restart. Because a block is a self-contained, immutable directory, it is also the unit that every long-term storage design for Prometheus builds on: uploading blocks to object storage is possible precisely because nothing will ever rewrite them in place.

  • Why does Prometheus memory grow with the number of active series rather than with how often you scrape them?
    Each active series costs fixed overhead in the head: its label set, its position in the in-memory index, and an open chunk. Halving the scrape interval doubles the samples inside existing chunks, which compress well and are memory-mapped out once closed. Doubling the number of series doubles the per-series structures that must stay resident, so cardinality moves memory far more than sample rate does.
  • A Prometheus server takes eleven minutes after a restart before it answers queries. What is it doing, and what actually shortens that?
    It is reloading memory-mapped head chunks and replaying the WAL to rebuild the head, and it will not scrape or serve until that completes. Replay time scales with head size, so the real lever is fewer active series; compressing WAL records reduces the bytes read back. Operationally the fix is a second replica scraping the same targets, restarted at a different time, so the gap is never user-visible.
  • Why can a block be uploaded to object storage while Prometheus is still running?
    Because a block is immutable once cut. Its `chunks/`, `index` and `meta.json` are written, fsynced and never modified again — only merged into a new block or deleted wholesale. That makes the directory safe to copy at any time, which is the property every long-term storage design for Prometheus is built on.

The write-ahead log is the running journal a bookkeeper scribbles as each transaction happens; the block is the bound ledger volume closed at the end of the period. The journal is only ever read again to rebuild the current page after a fire.

saying these in an interview costs you the question

  • Thinks queries are served from the write-ahead log
  • Believes blocks are appended to after they are cut
  • Says samples are written straight to a block on scrape
  • Expects restart to be instant regardless of head size
  • Assumes memory is driven by scrape interval, not series count
  • Treats compaction as optional housekeeping with no query benefit
open as a page

How does Prometheus remote_write behave when the receiver is slow or down, and how does federation differ?

level: seniorimportance: must knowfreq 62%

basics

~20 s

remote_write pushes samples out of the local write-ahead log, so a slow receiver grows the send queue and memory while scraping continues untouched; samples are lost only once the log is truncated past them. Federation instead pulls the latest value of selected series.

open as a page

What do Prometheus's two retention limits bound, and how do you size a server's disk from them?

level: middleimportance: should knowfreq 47%

basics

~20 s

Time-based retention bounds how old a sample may be; size-based retention bounds how much disk the database occupies. Whichever is reached first deletes the oldest blocks. Size a disk as samples per second times retention seconds times one to two bytes.

open as a page

Your Prometheus pairs keep six weeks locally, but a regulator wants six months and one global view. How would you choose between a sidecar-plus-object-store design and a receive-and-ingest design, and how does each deduplicate the replica pair?

level: principalimportance: should knowfreq 40%

basics

~20 s

A sidecar design leaves each Prometheus untouched and uploads its immutable blocks to object storage, deduplicating the replicas at query time by a replica label. A receive design has every Prometheus remote-write into a distributed ingestion tier that drops one replica's samples on arrival.

open as a page