In Prometheus, what happens to a time series when its target stops exposing it, and why does it vanish from an instant query instead of reading zero?
answer
- A missing series is not a zero
- One value, one timestamp, one series
- The server appends something when a line disappears
- Instant queries look back a bounded window
basics
~20 sA Prometheus sample is one float64 value at one millisecond timestamp on one series. When a target stops exposing a series, the server appends a stale marker, so queries after it return nothing — absence is not zero.
solid answer
~40 sA sample is just a `float64` and an `int64` millisecond timestamp; the labels belong to the series, whose identity is its complete label set including `__name__`. When a line a target used to expose is missing from a scrape, Prometheus appends a **stale marker** on that series at that scrape's timestamp. An instant query resolves each series by looking backwards from the query time for its most recent sample, within a bounded window — five minutes by default, set by `--query.lookback-delta`. A stale marker inside that window suppresses the older samples, so the series is absent from the result rather than reporting its last value. It is never zero: a missing counter means nobody reported, not that nothing happened, and a missing queue-depth gauge is no claim that the queue is empty.
code
text · 6 linesscrape @ 10:00:00 seedorder_dispatch_queue_depth{region="eu-west-3"} 43
scrape @ 10:00:15 seedorder_dispatch_queue_depth{region="eu-west-3"} 41
scrape @ 10:00:30 (the line is absent from the target's output)
-> stale marker appended on that series at 10:00:30
instant query @ 10:00:45 -> no result for the series (not 0, not 41)go deeper
Recall that a Prometheus sample is a value plus a timestamp, and that a series with no recent sample is missing from a query result rather than reading zero. Absence and zero are different answers.
Explain the stale marker written when a target stops exposing a line, and the bounded lookback an instant query performs, so you can say precisely why the old value is not carried forward.
Show what this means operationally: expressions over an absent series yield nothing to compare, so anything depending on a value has to handle absence explicitly rather than assuming a number will arrive.
Own the gap this leaves: a service that was never instrumented produces no series to observe at all, so coverage has to be checked against an inventory of what should be reporting, not against the telemetry itself.
## What a sample is, and what identifies the series it belongs to A Prometheus **sample** is exactly two things: a `float64` value and an `int64` timestamp in milliseconds since the Unix epoch. It carries no labels of its own. The labels belong to the **series**, and a series' identity is its complete label set including the reserved `__name__` label that holds the metric name. `seedorder_dispatch_queue_depth{region="eu-west-3"}` is one series; change any label value and you have a different series with its own independent history. This is why there is no such thing as "updating" a series or "renaming" one. A target that changes a label value stops writing to the old series and starts writing to a new one, and the old series keeps whatever samples it already had until retention removes them. ## What happens when a line stops appearing Suppose a target that has been exposing a line simply stops exposing it — the metric was removed, the code path is gone, or the whole target disappeared from service discovery. Prometheus does not leave the series dangling with its last value. At the scrape where the line is missing, the server appends a **stale marker**: a special value written on that series at that timestamp, which query evaluation understands as "this series has no value from here on". That marker is what turns a disappearance into a clean end rather than a value that lingers. ## Instant queries and the lookback window An instant query asks for the value of every matching series *at one timestamp*. Samples do not usually land exactly on that timestamp, so evaluation looks **backwards** from it for the most recent sample on each series, up to a bounded window — five minutes by default, set by `--query.lookback-delta`. Three outcomes: - there is a real sample inside the window, and its value answers the query; - the most recent sample is older than the window, and the series is simply not in the result; - a stale marker lies inside the window, and the series is not in the result either — the marker suppresses the older real samples behind it. Without stale markers, the second rule alone would carry a dead series' last value forward for the whole lookback window before it faded out. The marker collapses that to the next scrape. | Situation | What an instant query returns | |---|---| | a real sample within the lookback window | that sample's value | | the newest sample is older than the window | the series is absent from the result | | the target stopped exposing the line | absent from the stale marker's timestamp onward | | the target was removed from discovery | its series go stale, and so does the synthetic `up` series for it | | the series never existed | absent | ## Why absent rather than zero Because Prometheus stores no default value and knows nothing about what a series *should* contain. Filling in a zero would be inventing data, and for most metrics it would be a lie in the most dangerous direction: - A missing cumulative counter is not "nothing happened". It is "nobody told us", and those two look identical on a graph if you substitute zero. - A missing gauge such as a queue depth is not an empty queue. Zero would be a positive claim of emptiness that no measurement supports. - Substituting zero would make a broken exporter, a crashed process, a firewall rule and a genuinely idle service indistinguishable — which is precisely the distinction an operator most needs. The consequence you feel is that expressions over a vanished series produce **no series at all**, not a series reading zero. A comparison filters nothing because there is nothing to filter, and anything downstream that expected a value has to be written to treat absence explicitly rather than assuming a number will always arrive. ## The case where this bites hardest The seed-catalogue ordering service is rolled out to a new region and the pods come up without their metrics port exposed. Nothing in the metrics stack complains loudly: no series carrying `region="eu-west-3"` ever exists, so every dashboard that groups by region simply shows one fewer line and every threshold expression silently matches nothing. The only telemetry actually arriving from the region is an 18 GB-per-day log stream that nobody is watching, and the gap survives until somebody notices the missing line on a panel. That failure mode — *never instrumented* rather than *instrumented and broken* — is invisible to anything built on the series themselves, because you cannot alert on the absence of a series you have never seen by looking at that series. Catching it needs a check written against what you expect to exist: a comparison against the inventory of what should be reporting, or a rule that asserts a series is present rather than one that asserts its value is high. ## The short version to say out loud A sample is a value and a millisecond timestamp on a series identified by its full label set. When a target stops exposing a series, Prometheus writes a stale marker and evaluation stops resolving it. Absence is a first-class state, distinct from zero, and every query and alert you write has to decide what it means.
- How far back does a Prometheus instant query look for a series' latest sample?Up to the configured lookback window, five minutes by default, controlled by `--query.lookback-delta`. Within that window the most recent real sample answers the query. If the newest sample is older than the window, or a stale marker lies inside it, the series is absent from the result entirely rather than reporting an old value.
- Why does the stale marker exist at all, rather than just letting the series age out?Without it, a dead series would keep answering with its last value for the whole lookback window before fading. That is five minutes of a stale number presented as current, which is long enough for a threshold to be evaluated against a value nobody is measuring any more. The marker collapses that to the very next scrape.
- If a whole target disappears from service discovery, is anything left behind?The series it was exposing get stale markers at the following evaluation, and the synthetic `up` series for that target stops as well. The samples already written remain in storage until retention removes them, so historical queries over the period the target existed still return data — only timestamps after the marker resolve to nothing.
An unplugged thermometer does not read zero degrees; it reads nothing at all. Prometheus draws exactly that distinction, and the stale marker is how it records that the thermometer was unplugged.
saying these in an interview costs you the question
- Expects a missing counter to read zero
- Thinks the last value is carried forward indefinitely
- Treats no data as a healthy zero
- Believes removing a target deletes its stored samples
- Assumes query results interpolate across a gap