When would you reach for metric_relabel_configs in Prometheus rather than relabel_configs, and what can it not save you from?
answer
- two stages, two different inputs
- one sees targets, the other sees samples
- matching on the metric name
- the target still pays the cost
- already-stored series stay until retention
basics
~20 smetric_relabel_configs runs on every sample after a scrape, so it can drop a whole metric family or a label the target insists on emitting. It cannot stop the target producing them, and it never removes series already stored.
solid answer
~50 sThe two stages have different inputs. `relabel_configs` runs on the discovered **target** before any request is made; `metric_relabel_configs` runs on each **sample** returned by the scrape, just before ingestion, and sees the sample's own labels including `__name__`. That is what the second stage can do and the first cannot: drop a metric family by name, strip a label off every series, or keep an allowlist — typically when a third-party process you cannot modify emits something expensive. What it cannot do is equally important. The target still computes and serialises everything, and the payload still crosses the network. Series already written stay until retention expires. Rules and dashboards referring to what you dropped go quiet rather than failing. And stripping a label that distinguished two series makes them collide, so the scrape fails with duplicate samples.
code
yaml · 7 linesmetric_relabel_configs:
- source_labels: [__name__]
action: drop
regex: freight_leg_shipment_.*
- source_labels: [__name__]
action: keep
regex: (freight|fleet)_.*go deeper
Recall that a scrape job has two relabeling stages and that one runs before the request while the other runs on the samples that come back. Knowing which is which is enough at this level.
Explain what the later stage can match on, why matching a metric name is only possible there, and why discovery metadata is unavailable to it.
Show the limits from experience: the target still pays the cost, stored series survive until retention runs out, stripped labels can collide and break an entire scrape, and silent rules make dependent alerts stop without failing.
Own the policy. Decide when ingestion-side filtering is legitimate containment versus a habit that hides upstream problems, who may add such a rule, and how each one is recorded so it can be removed rather than inherited.
Every Prometheus scrape job can carry two relabeling lists, and although they share a rule syntax they are separate stages with separate inputs. Mixing them up is a fast way to write a configuration that quietly does nothing. ## Two stages, two inputs | | `relabel_configs` | `metric_relabel_configs` | |---|---|---| | Runs | Before the scrape | After the scrape, before storage | | Operates on | One discovered target's label set | Each returned sample's label set | | Can see | `__meta_*` discovery metadata | `__name__` and the sample's own labels | | Can drop | A target, so it is never polled | A sample, so it is never stored | | Can change | Where and how the scrape is made | What the stored series is called | The asymmetry is worth stating explicitly, because it is the interview's real question. The second stage **cannot** see discovery metadata — those labels were discarded at the end of target relabeling — and it cannot decide whether a target is scraped at all. The first stage **cannot** see anything about the metrics, because it runs before the request that would return them. ## What only the second stage can do Matching on `__name__` gives you a per-metric filter. In practice this is nearly always containment of something you do not control: - A process you cannot modify exposes a metric family with an identifier baked into a label. - A library ships a debug family that is worthless in production and enormous in volume. - A migration means two names for the same thing during the overlap, and you want only one stored. - A label carries the same value on every series and simply wastes space. A worked case. A touring-logistics platform runs an incident nobody can explain from the existing dashboards: queries that used to answer instantly now time out, and the graphs everyone relies on are blank at exactly the moment they are needed. The cause turns out to be a freight-tracking process that began attaching a per-shipment identifier to a duration metric after a library upgrade, adding roughly 61,400 series a day. Nobody on the team owns that code, and the fix upstream is a release away. The containment is a `drop` on `__name__` for that family in `metric_relabel_configs`, applied on the next configuration reload. Volume stops growing immediately. ## What it will not save you from 1. **The target still does the work.** It computes, allocates and serialises every one of those series on every scrape, and the whole payload crosses the network and is parsed. You have moved the cost, not removed it. 2. **Existing series do not vanish.** The rule only stops new samples. The old series remain queryable and keep occupying the store until retention expires — and in that estate a compliance floor of 96 hours of retention means four more days of slow queries before the graphs recover on their own. Explicit deletion is possible through the administrative API when the server was started with `--web.enable-admin-api`, but that is a deliberate operator action, not a side effect of your rule. 3. **Dropping a label can make series collide.** If two series differed only by the label you removed, they become the same series within one scrape, and the scrape fails with duplicate samples for the same timestamp. You then lose everything from that target, not just the metric you were trying to trim. When the label is genuinely the distinguishing dimension, drop the family instead. 4. **Things that referenced the metric go quiet.** Recording and alerting rules written against a dropped name do not error; they return no data. An alert that never fires looks exactly like a healthy system. 5. **The list itself costs.** Every rule runs against every sample of every scrape on that job. A long allowlist on a high-volume job is not free. ## Where the fix really belongs Treat the second stage as a containment tool with a short expected life, not as the design. The order of preference is: stop emitting it at the source; if you do not own the source, aggregate or filter it as close to the source as you can; and only then filter at ingestion in the server. Whichever you choose, write down why the rule exists — a bare `drop` on a metric name, discovered by someone two years later during an incident, is indistinguishable from a mistake, and removing it is how a fleet quietly rediscovers the problem it was hiding.
- You dropped a label with this stage and the whole scrape started failing. What happened?The label was the only thing distinguishing several series from that target. Removing it made them identical, so the scrape produced repeated samples for the same series at the same timestamp and was rejected outright. The result is worse than the original problem: you lose every metric from that target. When a label really is the distinguishing dimension, drop the metric family, or aggregate at the source.
- The rule is deployed but queries are still slow days later. Why?Because the rule only prevents new samples. Every series written before it still exists and is still selected and scanned by queries that match it, until retention expires. Explicit deletion via the administrative API is available when the server is started with the flag that enables it, but otherwise recovery is on the retention clock, not on the deploy clock.
- Why can this stage not filter on the discovery metadata that selected the target?Because those labels are gone. Every label whose name still begins with a double underscore is discarded at the end of target relabeling, before any request is made. By the time samples exist, the only labels available are the sample's own — its name and its dimensions — plus the target labels that survived the first stage.
saying these in an interview costs you the question
- Thinks the rule stops the target producing the metric
- Expects dropped series to disappear from stored data immediately
- Strips a label without checking whether series then collide
- Believes this stage can filter which targets are scraped
- Tries to match on discovery metadata that no longer exists
- Leaves an unexplained drop rule with no record of why