In Prometheus, what is stored when a scrape times out, and why must scrape_timeout not exceed scrape_interval?
answer
- Nothing partial survives
- The report series are written anyway
- Panels go blank rather than flatline
- A request may not outlive its own cycle
- Watch duration before up flips
basics
~20 sA timed-out Prometheus scrape stores none of the target's samples; the partial body is discarded. It still writes an up value of 0 plus the scrape report series, and marks the target's earlier series stale. Prometheus refuses to load a timeout larger than the interval.
solid answer
~50 sA scrape is all or nothing. If it times out, is refused, returns a non-200, fails to parse, or trips `sample_limit`, Prometheus keeps none of that target's samples — the part of the body that did arrive is thrown away with the rest, because a half-ingested scrape would be a wrong picture rather than a smaller one. What is still written is the report: `up` at 0, `scrape_duration_seconds` at roughly the timeout, and `scrape_samples_scraped` at 0. Prometheus also emits staleness markers for the series that target had been producing, so panels go blank immediately instead of holding the last value for the lookback window. A `scrape_timeout` greater than the applicable `scrape_interval` is refused at configuration load: a request that can still be running when the next tick arrives would overlap itself and destroy the even spacing the whole model assumes.
code
text · 4 linesup{job="turbine-gateway",instance="gw-11:9100"} 0
scrape_duration_seconds{job="turbine-gateway",instance="gw-11:9100"} 10.001
scrape_samples_scraped{job="turbine-gateway",instance="gw-11:9100"} 0
scrape_series_added{job="turbine-gateway",instance="gw-11:9100"} 0go deeper
Know that a scrape either works or does not, that up records which, and that a target failing to answer is visible in Prometheus rather than being silent. Recognising a timeout in the targets page error text is enough here.
Explain the mechanics: no partial ingest, the report series written on failure, and the constraint that a timeout must fit inside its interval. Be able to give the overlap argument for why that constraint exists.
Show that you diagnose from scrape duration before up flips, that you can separate a timeout from a refusal or a parse failure, and that you know staleness markers are why panels empty instead of holding stale values.
Own the standard: which jobs get which cadence, what a team must do when a target outgrows the scrape path, and how you keep an estate from drifting into timeouts sized to hide slow endpoints rather than fix them.
## A scrape is all or nothing When a Prometheus scrape fails, the server keeps none of the target's samples. That is true whether the failure was a refused connection, a TLS error, a non-200 response, a body that stopped arriving before `scrape_timeout` expired, a body that failed to parse partway through, or a scrape that breached a configured `sample_limit`. The half of the response that did arrive is discarded with the rest; there is no partial ingest, and no attempt is made to keep the prefix that parsed cleanly. That design is deliberate. Half a scrape is not a smaller truth, it is a different one: some series present, others silently missing, all of them stamped as if they were a complete picture of the target at that instant. Anything built on top — a ratio between two series, a sum across a job — would quietly produce a wrong number rather than an obvious gap. ## What is still recorded The scrape's own report series are written regardless, so the failure itself is data: | Series after a failed scrape | Value | |---|---| | `up` | 0 | | `scrape_duration_seconds` | the elapsed time, which for a timeout sits at roughly the configured limit | | `scrape_samples_scraped` | 0 | | `scrape_series_added` | 0 | Alongside that, the target's page in the Prometheus web interface carries the last error text and the last scrape duration for each target, which is normally the fastest way to separate "refused" from "timed out" from "invalid response". The most useful of these in day-to-day operation is `scrape_duration_seconds` on the scrapes that still *succeed*. A target creeping from 0.4 s to 8.6 s against a ten-second timeout is a target that will start failing this week, and the duration series says so days before `up` flips. ## Why the graph ends instead of flatlining Beyond the report series, a failed scrape triggers **staleness markers** for the series that target was previously producing. A staleness marker is a special entry meaning "this series has no current value", and its effect is that an instant query stops returning the series straight away instead of continuing to return the last sample for the length of the query lookback window (five minutes by default). This is why a dashboard panel for a dead target goes blank rather than flatlining at its last value, and it is worth being explicit about in an interview: without staleness markers, a target that died at 09:00 would still be reporting a healthy-looking number at 09:04. ## Why the timeout can never exceed the interval Prometheus rejects a scrape configuration in which `scrape_timeout` is larger than the `scrape_interval` that applies to the same job — the configuration fails to load, and on a reload the running server keeps the old one. It is not a tuning choice you might get away with, and the reasoning is worth being able to give: - A request permitted to run longer than the gap between requests can still be in flight when the next tick arrives, so scrapes of one target would overlap or queue. - Samples would then arrive at a cadence unrelated to the configured one, breaking the assumption that a target's series are evenly spaced. - The load imposed on the target would no longer be bounded by the interval — a slow target would accumulate concurrent readers exactly when it is least able to serve them. When `scrape_timeout` is not specified at all, it is inherited rather than invented, and it never ends up longer than the interval it lives under. The fix for a target that genuinely needs longer than the interval is never a bigger timeout: it is a longer interval for that job, or an exporter that stops doing expensive work on the scrape path. ## Diagnosing it on a real estate A wind-farm maintenance planner scrapes 47 site gateways every fifteen seconds with a ten-second timeout, and its own API carries a 320 ms p99 budget. One morning three gateways start showing zeros in `up` and their panels go blank. The order of work: 1. Read `scrape_duration_seconds` for those targets over the previous week. A slow climb points at the target or at a query behind it getting more expensive; a step change points at a deployment or a network path. 2. Check whether the failing targets share anything — a site, a subnet, a release — since three at once is rarely three independent faults. 3. Read the last error text on the targets page. "Context deadline exceeded" is a timeout, and a very different problem from a connection refused or a certificate that just expired. 4. Confirm whether the target is slow for everyone or only for the scraper, remembering that each Prometheus replica is an independent reader and the target is answering all of them. The wrong first move — raising the timeout — is unavailable anyway once it reaches the interval, and that constraint is doing you a favour: it forces the conversation back to why the response got expensive rather than letting the scrape quietly consume the whole cycle.
- A target's scrape duration has climbed from 0.4 s to 8.6 s but up is still 1. What do you do?Treat it as an outage that has not happened yet. The duration series is the leading indicator: at a ten-second timeout that target has weeks of headroom left at best, and it will fail first under exactly the load that makes you want the data. Find what got expensive — usually a growing enumeration behind the endpoint or a target now exposing far more series — rather than waiting for the zeros.
- Why not simply raise the timeout to give a slow target more room?Because the ceiling is the interval, and the configuration is refused above it. The real remedies are a longer interval for that job, splitting one enormous endpoint into several targets, or moving expensive work off the scrape path into a background refresh. Raising the timeout only trades a fast honest failure for a scrape that consumes most of its own cycle and still leaves you nothing when it eventually breaches.
- What is the difference between a staleness marker and a series simply having no new samples?A staleness marker is a positive statement that the series has ended, so an instant query returns nothing for it from that moment. Absence of new samples alone is indistinguishable from a slow writer, so a query would keep returning the last value until the lookback window expires. That distinction is why a dead target's panel empties immediately rather than holding a comforting last reading for minutes.
saying these in an interview costs you the question
- Believes the samples parsed before the timeout are kept
- Says a failed scrape writes nothing at all
- Thinks a dead target's series flatline at their last value
- Treats a timeout larger than the interval as an aggressive tuning choice
- Reads up as a health check of the application rather than of the scrape
- Ignores scrape duration until up has already started flipping