Your Prometheus pairs keep six weeks locally, but a regulator wants six months and one global view. How would you choose between a sidecar-plus-object-store design and a receive-and-ingest design, and how does each deduplicate the replica pair?
answer
- Both end with blocks in a bucket
- One uploads, the other is written to
- Where does the state live
- Dedupe on read versus dedupe on write
- Labels identifying cluster and replica
basics
~20 sA sidecar design leaves each Prometheus untouched and uploads its immutable blocks to object storage, deduplicating the replicas at query time by a replica label. A receive design has every Prometheus remote-write into a distributed ingestion tier that drops one replica's samples on arrival.
solid answer
~50 sBoth designs end with blocks in object storage and a querier that fans out; they differ in where the state and the risk sit. In a **sidecar** design a process beside each Prometheus uploads its cut blocks to a bucket, a store gateway serves them back, and a querier merges recent data from the sidecars with historic data from the bucket — ingestion stays exactly as it is today, and the long-term tier can be down without touching monitoring. In a **receive-and-ingest** design each Prometheus remote-writes into a distributed tier whose ingesters build and flush blocks, which buys multi-tenancy and an inbound-only network path at the price of running a stateful cluster on the ingest path. Deduplication follows the same split: query-time merge on a replica label for the sidecar design, versus dropping the non-elected replica at ingestion for the receive design.
code
yaml · 6 lines# on replica A; replica B differs only in the last value
global:
scrape_interval: 15s
external_labels:
cluster: cellar-eu-1
replica: ago deeper
Know that a single Prometheus keeps a limited window of data on its own disk, and that keeping months of history or querying many servers at once needs an additional system built around it rather than a setting.
Be able to describe the two shapes: blocks uploaded from beside each server into object storage, versus samples remote-written into a shared ingestion cluster, with a querier in front of either. Know that external labels are what identify each source.
Show that you have run one. Explain why local compaction is disabled under an uploader, how replicas are deduplicated in each design, what downsampling buys a wide query, and what still works when the long-term tier is unavailable.
This is the judgement call. Weigh operating a stateful write path against a read-path-only failure mode, decide on tenancy, network direction and blast radius, and be able to defend the cost and the retention policy to someone who will audit it.
Prometheus is single-node by design, and both long retention and a global view are things it deliberately does not do. Two architectures fill that gap, and choosing between them is a judgement call about where you hold state. The precondition for either is that every Prometheus carries a distinct `external_labels` set in its `global` configuration: those labels identify which server produced a block or a sample, and every deduplication downstream is built on them. ## What each design moves where | | Sidecar plus object store | Receive and ingest | |---|---|---| | How data leaves Prometheus | a process beside it uploads cut blocks | Prometheus remote-writes samples | | Network direction | outbound to a bucket, plus inbound queries | outbound writes into a shared cluster | | Where recent data is queried from | still the Prometheus itself | the ingestion tier's memory | | Where historic data is queried from | a gateway reading the bucket | a gateway reading the bucket | | State you now operate | a bucket, a gateway, a querier, a compactor | all of that plus a replicated write path | | If the long-term tier is down | monitoring and alerting are unaffected | senders queue, then lose, and recent data is gone | | Multi-tenancy | bolted on, per-bucket at best | first-class, because writes are addressed | In the sidecar shape, the sidecar uploads each two-hour block as Prometheus cuts it, a **store gateway** makes the bucket's blocks queryable, a **querier** fans out to the live servers and the gateway and merges the results, and a **compactor** runs against the bucket, doing the merging and downsampling the local servers no longer do. Local compaction must be disabled on any Prometheus whose blocks are uploaded, so the uploader always sees stable, uniform blocks rather than ones rewritten underneath it. In the receive shape, a distributor accepts remote writes, hashes each series onto a set of **ingesters**, and each ingester builds blocks in memory and flushes them to the same object storage. You gain a genuine write API, tenant isolation, and the ability to accept data from networks you cannot reach inbound — and you acquire a replicated, stateful cluster in the path of every sample, with a replication factor to choose, a hash ring to operate, and rollouts that must not lose unflushed data. ## Deduplicating a highly-available pair Two Prometheus servers scraping the same targets produce two near-identical copies with slightly different timestamps. Neither design tolerates that by accident. - **Sidecar design — dedupe on read.** Each server's `external_labels` include a replica identifier as well as the labels that say which cluster it belongs to. Both replicas upload; the querier is told which label distinguishes replicas, and at query time it merges series that are identical apart from that label, preferring one replica's samples and filling from the other where there are gaps. The upside is that a replica outage is invisible — the other replica's blocks cover the hole. The cost is storing two copies, and keeping the bucket compactor from merging two replicas' overlapping blocks in ways you did not intend. - **Receive design — dedupe on write.** Both replicas remote-write, and the ingestion tier runs an availability tracker that elects one replica per cluster as leader and discards the other's samples while the leader keeps sending; when the leader goes quiet it fails over. Only one copy is stored, so the storage bill is halved, but a failover leaves a small seam, and the mechanism depends on both servers carrying agreed cluster and replica labels. ## How to actually decide 1. **Can you operate a stateful distributed system on the ingest path?** If the answer is honestly no, the sidecar design is the one that fails safely: a bucket outage costs you history, not alerting. 2. **How many tenants?** One platform team with a handful of clusters does not need addressed writes; dozens of teams needing isolation and per-tenant limits is what the receive design was built for. 3. **Which way does the network go?** Clusters you cannot reach inbound — a customer site, a restricted environment — can push, but cannot be queried by a central querier reaching into them. That constraint often decides it. 4. **What blast radius will you accept?** Sidecar keeps failure in the read path; receive puts it in the write path, where an outage means data that was never stored anywhere. 5. **What does it cost?** Object storage is cheap; ingester memory and the replication factor multiplying it are not. ## What six months of evidence actually demands Long retention for a regulator is not the same problem as long retention for a dashboard. Beyond keeping the blocks, you owe three things. **Downsampling**, because nobody queries six months at raw resolution — the compactor writes reduced-resolution copies alongside the raw blocks and the querier picks the right one for the query's step. **A retention policy expressed per resolution**, so raw data can age out while the coarse copies survive the full window. And **a story about the bucket itself**: versioning or object-lock so the evidence cannot be quietly deleted, and a restore you have actually rehearsed rather than assumed. The failure mode to name out loud is the tempting one: bigger local disks. That gives no global view, does not survive the server, and moves restart and compaction cost in exactly the wrong direction.
- Why must local compaction be disabled on a Prometheus whose blocks are being uploaded to object storage?Because the uploader treats a cut block as final. If the local compactor merges two-hour blocks into wider ones underneath it, the bucket ends up holding both the small blocks already uploaded and a wider block covering the same range, which leaves overlapping data for the querier and the bucket compactor to reconcile. Uploading uniform, never-rewritten blocks and letting a compactor in the bucket do the merging keeps one authority for each range.
- Six months of raw samples makes a year-long dashboard panel unusable. What does the long-term tier do about that?It downsamples. A compactor running against the bucket writes reduced-resolution copies of each block alongside the raw data, typically at a five-minute and a one-hour step, and the querier selects a resolution appropriate to the query's step so a wide range reads far fewer points. Retention is then set per resolution, letting raw data expire while the coarse series survive the full compliance window.
- What breaks in each design when the object store itself is unavailable for an hour?In the sidecar design, historic queries fail and uploads stall, but every Prometheus keeps scraping, storing and alerting on local data, so nothing is lost. In the receive design, ingesters cannot flush; they hold data in memory and the outage becomes a capacity problem in the write path, with real data loss if an ingester is restarted before it can flush.
saying these in an interview costs you the question
- Proposes a bigger local disk as the answer to long retention
- Thinks two replicas can simply both write into one store
- Believes a remote-write ingestion tier is stateless
- Leaves local compaction on while uploading blocks to a bucket
- Treats object storage as free and skips downsampling entirely
- Cannot say what still works when the long-term tier is down