skip to content

Scraping and Exporters

Prometheus pulls rather than receives: targets expose a plain-text /metrics endpoint and get scraped every interval. Expect questions on scrape_interval vs scrape_timeout, when to reach for node_exporter, cAdvisor or blackbox, and how to instrument something that cannot be scraped.

on this pageshow

questions

4

In Prometheus, what happens on each scrape of one target, and what do scrape_interval and scrape_timeout bound?

level: juniorimportance: must knowfreq 84%

answer

  1. One poll of one target, on a ticker
  2. Two durations bound different things
  3. The scraper writes series about itself
  4. up is written even on failure
  5. Every sample shares the scrape's timestamp

basics

~20 s

Each cycle Prometheus sends one HTTP GET to the target's metrics path, parses the body, and stores every sample under that scrape's start timestamp. scrape_interval sets how often it repeats; scrape_timeout bounds one request. Prometheus also records an up sample per scrape.

solid answer

~40 s

Prometheus polls each target on its own schedule. On every tick it opens a single HTTP GET to the target's `scheme`, address and `metrics_path` (default `/metrics`), sending an `X-Prometheus-Scrape-Timeout-Seconds` header so the target can bound its own work. It reads and parses the body, then appends every sample with one timestamp — the moment the scrape began — so all series from that target line up on the same grid. `scrape_interval` sets how often the cycle repeats and therefore your resolution floor; `scrape_timeout` bounds a single request and may not exceed the interval. Whatever the outcome, Prometheus writes its own report series for that scrape: `up` (1 or 0), `scrape_duration_seconds`, `scrape_samples_scraped`, `scrape_samples_post_metric_relabeling` and `scrape_series_added`. `up` is written even when the target never answered, which is what makes "this target is down" queryable at all.

code

yaml · 16 lines
yaml
global:
  scrape_interval: 15s
  scrape_timeout: 10s

scrape_configs:
  - job_name: turbine-gateway
    metrics_path: /metrics
    scheme: http
    static_configs:
      - targets: ['gw-11.windfarm.internal:9100']

  - job_name: site-historian
    scrape_interval: 60s
    scrape_timeout: 45s
    static_configs:
      - targets: ['historian.windfarm.internal:9104']

go deeper

for a junior

Be ready to say that Prometheus polls each target over HTTP on a schedule, that the conventional path is /metrics, and that a series called up records whether that poll worked. Naming the two duration settings is expected.

for a middle

Explain the mechanics: one request per target per interval, every sample stamped with the scrape's start time, and the report series the server writes about its own scrape. Be able to say what each of those counts measures.

for a senior

An interviewer expects you to weigh interval against resolution, storage volume and the cost imposed on targets, and to know that a zero up sample and a missing up series are different failures with different causes.

for a principal

Own the estate-wide policy: which jobs earn fine resolution, what interval you standardise on by default, and how you stop teams paying for sub-minute cadence on signals nobody inspects faster than a shift change.

## What one scrape cycle actually does A Prometheus **target** is one address plus a path that the Prometheus server polls; a **scrape** is one poll of one target. Each target runs on its own ticker, offset by a hash of its label set, so a job with hundreds of targets spreads its requests across the interval instead of stampeding on the same second. On each tick the server: 1. Builds a URL from the target's `scheme` (default `http`), its address, and its `metrics_path` (default `/metrics`). 2. Issues a single HTTP GET, advertising the exposition formats it accepts and sending an `X-Prometheus-Scrape-Timeout-Seconds` header so a well-written target can trim its own work to fit. 3. Reads the whole body, bounded by `scrape_timeout` and, if configured, `body_size_limit`. 4. Parses the body into samples and applies any per-sample relabelling and limits such as `sample_limit`. 5. Appends every surviving sample **with a single timestamp** — the moment the scrape began — so all series from that target land on the same grid point. 6. Appends its own report series describing the scrape. 7. Compares the series seen now with those seen last time and writes a **staleness marker** for any that have disappeared, so queries stop returning them immediately rather than at the end of the lookback window. Step 5 is the one candidates miss. The timestamp is not when a counter was incremented inside the application; it is when Prometheus asked. If the exposition itself carries explicit timestamps and `honor_timestamps` is left at its default of true, those are used instead — which is how a translating exporter or a federated server can preserve original times. ## What each duration bounds | Setting | Bounds | Failure mode when misjudged | |---|---|---| | `scrape_interval` | how often one target is polled | too long and short-lived events fall between samples; too short and both storage and target CPU grow for resolution nobody reads | | `scrape_timeout` | how long one HTTP request and parse may take | too short and a merely slow target is recorded as down; larger than the interval and the configuration is refused | | `sample_limit` | how many samples one scrape may yield | guards the server against a cardinality explosion, at the price of losing the entire scrape when it trips | Prometheus's own defaults are a one-minute interval and a ten-second timeout. Both may be set globally and overridden per job, which is the normal shape in practice: a short global interval, with the one job full of expensive targets moved out to something slower. ## The series Prometheus writes about itself Every scrape produces these, carrying the target's `job` and `instance` labels: | Series | Meaning | |---|---| | `up` | 1 when the scrape succeeded, 0 when it failed for any reason | | `scrape_duration_seconds` | wall-clock time of the request and the parse | | `scrape_samples_scraped` | how many samples the target exposed | | `scrape_samples_post_metric_relabeling` | how many of them survived per-sample relabelling | | `scrape_series_added` | series new to this target since the previous scrape | Two properties make them valuable. First, they are written by the scraper *about the scrape*, so they exist even when the target is a smoking hole — which is precisely why "this target is down" is an expressible fact rather than an absence of data. Second, `scrape_series_added` is the early warning for cardinality: a target that quietly begins stamping a request identifier into a label shows up here long before storage complains. Note the difference between an `up` sample of 0 and no `up` sample at all. Zero means Prometheus tried and failed. Nothing at all means the target is not in the target list any more — nobody is even trying. Watching only for the first and never noticing the second is a common hole in an alerting setup. ## What the cycle deliberately does not do - It does not ask for history. A scrape returns the target's current values, so anything that happened between two polls is only recoverable if the target exposes it cumulatively — which is why a counter that only ever goes up survives a missed scrape and a gauge does not. - It does not care what produced the numbers. The server sees an HTTP body; whether behind it sits an application, a translating exporter process, or a shell script writing a file is invisible to it. - It does not coordinate with other scrapers. Every Prometheus server, every high-availability replica, and every engineer running an ad-hoc request against the metrics path is an independent reader with its own clock. ## Choosing the interval A maintenance planner for a wind farm polls 47 site gateways. At a fifteen-second interval its 2,310 series each produce 5,760 samples a day; moving that job to sixty seconds cuts the volume fourfold, which is the first lever anyone reaches for when a telemetry budget is halved. What it costs is the ability to see anything shorter than a minute. Totals survive, because counters are cumulative and the next scrape still carries the full sum; a gauge does not, so an inverter temperature that spikes for eight seconds and settles simply never happened as far as the time-series database is concerned. That is a per-job decision about what each signal is for, not a global knob to be turned once and forgotten.

  • Two Prometheus servers scrape the same target. Does the target see one request or two?
    Two. Scraping is driven entirely by the reader, so every server, every high-availability replica and every ad-hoc request against the metrics path is an independent HTTP call with its own timing. The target has no idea how many are watching. That is why an endpoint that does expensive work per request scales with the number of scrapers rather than with user traffic, and why an exporter's cost has to be reasoned about per reader.
  • Why do all samples from one scrape share a timestamp rather than the time each value was read?
    Because Prometheus stamps the scrape, not the reading. One timestamp per target per cycle puts every series from that target on the same grid point, so two metrics gathered together can be combined without interpolation. The cost is that the timestamp records when Prometheus asked, not when the underlying event happened, and the ordering of things that occurred between two scrapes is simply not recorded.
  • What does it mean if a target's up series stops existing altogether?
    That the target is no longer in Prometheus's target list — discovery dropped it, or the job was removed — rather than that it is down. A zero means a target that answered badly; a missing series means one nobody is polling. An alert written only against a zero value goes quiet in exactly the case where an entire job silently disappears from the configuration.

A night watchman walking a fixed round: he notes what each door looked like at 02:15, and his own logbook records that he made the round at all — even on the night he found the gate impassable.

saying these in an interview costs you the question

  • Thinks the target pushes its metrics to Prometheus each interval
  • Believes a scrape returns everything that happened since the last poll
  • Says scrape_timeout bounds how long the whole job takes
  • Assumes the sample timestamp is when the value changed inside the application
  • Cannot say what the up series is or who writes it
  • Thinks a failed scrape simply leaves no trace at all
open as a page

In a Prometheus estate, what do node_exporter, cAdvisor and blackbox_exporter each measure, and when does only a probe answer the question?

level: middleimportance: should knowfreq 64%

basics

~20 s

node_exporter reports one host's OS and hardware metrics, cAdvisor reports per-container resource use from control-group accounting, and blackbox_exporter probes an endpoint from outside over HTTP, TCP, ICMP or DNS. Only the probe sees the path a client actually travels.

open as a page

In Prometheus, what is stored when a scrape times out, and why must scrape_timeout not exceed scrape_interval?

level: seniorimportance: should knowfreq 56%

basics

~20 s

A timed-out Prometheus scrape stores none of the target's samples; the partial body is discarded. It still writes an up value of 0 plus the scrape report series, and marks the target's earlier series stale. Prometheus refuses to load a timeout larger than the interval.

open as a page

When writing a Prometheus exporter for a system you cannot instrument, what work belongs on the scrape path?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Fetch on the scrape path while it stays cheap, so every scrape returns current state. If the fetch is expensive, refresh it on the exporter's own background timer, serve the cached snapshot, and expose whether the last refresh succeeded and when.

open as a page