How does Prometheus find its scrape targets, and what does a service-discovery mechanism hand it besides a list of addresses?
answer
- the config describes how to find
- discovery re-runs, it is not startup-only
- targets are label sets, not strings
- __address__ plus read-only __meta_ labels
- files on disk versus a platform API
basics
~20 sPrometheus builds its target list from the discovery blocks in each scrape job: a static list, files on disk, or a platform API such as Kubernetes or EC2. Each target arrives carrying read-only metadata labels describing its origin.
solid answer
~40 sEvery entry under `scrape_configs` names one or more discovery mechanisms: `static_configs` for a fixed list, `file_sd_configs` for JSON or YAML files on disk, and API-backed ones such as `kubernetes_sd_configs`, `ec2_sd_configs` or `consul_sd_configs`. A mechanism does not return strings; it returns **label sets**. Each target carries an `__address__` label holding host and port, and a pile of read-only `__meta_*` labels describing its origin: namespace, pod annotations, EC2 tags. Those meta labels are the raw material `relabel_configs` filters and rewrites, and any label still starting with `__` is discarded once target relabeling finishes, so metadata never reaches storage unless you copy it into a normal label. File-based discovery is re-read from disk and needs no platform credentials; API-backed discovery reflects churn within seconds but depends on that API being reachable and on read permissions.
code
yaml · 6 linesscrape_configs:
- job_name: tour-freight-nodes
file_sd_configs:
- files:
- /etc/prometheus/targets/*.json
refresh_interval: 30sgo deeper
Be ready to say why a scrape configuration describes how to find targets instead of naming them, and to name three mechanisms: a static list, files on disk, and a platform API such as Kubernetes.
Explain that a mechanism returns label sets rather than addresses, name the reserved labels that steer the scrape, and say what happens to the discovered metadata once relabeling is done.
Show you have operated this: which page tells you what discovery really returned, what needs a config reload and what does not, how you generate target files safely, and what a stalled discovery source looks like from the outside.
Own the choice of where the fleet's truth lives. Argue when a reviewable generated file beats a live platform query, what each choice costs in credentials and API load, and how teams get their services scraped without editing a central file.
Prometheus scrapes exactly what its configuration tells it to scrape. In an estate where processes are replaced hourly, nobody can maintain that list by hand, so the configuration describes **how to find** targets rather than naming them. Service discovery is the part of the server that turns that description into a live set of things to poll, and it re-runs continuously rather than once at startup. ## What a scrape job actually contains Each entry in `scrape_configs` has a `job_name`, options that govern the HTTP scrape itself, one or more **discovery mechanisms**, and usually a `relabel_configs` list. The discovery mechanisms are the plural blocks: - `static_configs` — a literal list of addresses written in the config file. - `file_sd_configs` — a set of file globs; the files hold target lists and are re-read when they change and on a `refresh_interval`. - `kubernetes_sd_configs`, `ec2_sd_configs`, `gce_sd_configs`, `consul_sd_configs`, `dns_sd_configs` and others — mechanisms that ask a platform what currently exists. A single job may list several mechanisms, and their results are combined into one target set for that job. ## A target is a label set, not an address This is the point most candidates miss. Discovery hands Prometheus a **label set per target**, and the address is simply one of the labels. A handful of reserved names control the scrape: | Label | What it controls | |---|---| | `__address__` | The host and port the scrape request goes to | | `__scheme__` | `http` or `https` | | `__metrics_path__` | The path requested, `/metrics` by default | | `__param_<name>` | A URL query parameter added to the request | | `__meta_*` | Read-only metadata supplied by the discovery mechanism | The `__meta_*` labels are where the mechanism describes what it found: `__meta_kubernetes_namespace` and `__meta_kubernetes_pod_name` from a cluster, `__meta_ec2_instance_id` and `__meta_ec2_tag_<key>` from a cloud API, `__meta_filepath` from a file-based mechanism. They exist so that relabeling has something to decide on. Two consequences follow. First, every label whose name still begins with `__` is removed once target relabeling has finished, so metadata is invisible on the stored series unless a rule copies it into an ordinary label. Second, if you never set `instance` yourself, it defaults to the value of `__address__`, and `job` defaults to the job's `job_name`. That is why two targets discovered at the same host and port collapse into the same identity unless you distinguish them. ## File-based discovery versus an API-backed mechanism Both produce the same kind of output; they differ entirely in operations. | | File-based | API-backed | |---|---|---| | Where truth lives | Files a generator writes | The platform's own API | | How change is noticed | File change plus a refresh interval | A watch or a poll against the API | | Credentials needed | None beyond filesystem access | Read permissions on the platform | | Failure mode | Stale file, silently old target list | Discovery stalls when the API is unreachable | | Who can change it | Whatever writes the files | Anyone who can create objects in the platform | | Reviewability | The file can be committed and diffed | Emergent, and only observable after the fact | File-based discovery is not a beginner's fallback. It is the standard adaptor for anything the server cannot query directly: hardware, appliances, a partner's endpoints, or an inventory that already exists in a configuration-management system. You point a job at a glob, have a generator write the files, and the two systems stay decoupled. ## Operating it 1. **Look at what discovery produced, not at what you expected.** The server's `/service-discovery` page lists discovered targets with their labels before and after relabeling, including the ones that were dropped; `/targets` shows only what survived. 2. **Know what needs a reload.** Editing the configuration file needs a `SIGHUP` or a POST to `/-/reload` when the server was started with `--web.enable-lifecycle`. Changing the *contents* of a file-based target file needs nothing — that is the whole point of the mechanism. 3. **Write target files atomically.** Write to a temporary name and rename it into place, so a half-written file is never read mid-update. 4. **Count your watches.** Every job with its own API-backed mechanism is another consumer of that API, and a large configuration can put real load on it. ## Where the boundary sits Keep three concerns separate when you answer. Discovery answers *what exists right now*. Relabeling answers *which of those do we scrape, and what do we call them*. The scrape itself — how often, how long it may take, what the endpoint returns — is a third concern that is configured on the same job but is not part of discovery at all.
- A target appears on the server's targets page with none of the metadata you expected on its series. Where did it go?Nowhere yet — it was discarded. Every label whose name starts with `__`, including all `__meta_*` labels, is removed once target relabeling finishes. To keep one you must copy it into an ordinary label with a `replace` or `labelmap` rule. The `/service-discovery` page shows the discovered label set before that discard, which is the fastest way to confirm the metadata was really there.
- If you never configure an instance label, what does Prometheus put there?It defaults to the value of `__address__` as it stands after relabeling, and `job` defaults to the job's `job_name`. That default is convenient but it means identity is tied to host and port: rewrite `__address__` in a relabel rule and you have silently renamed every series from that target.
- Why keep file-based discovery around in an estate that is mostly Kubernetes?Because plenty of things are not in the cluster and cannot be queried: network hardware, managed services, appliances, a partner's endpoint, anything already inventoried by configuration management. File-based discovery needs no credentials on those systems, the files can be generated from whatever inventory is authoritative, and they can be reviewed and diffed before they take effect.
Discovery is closer to a passenger manifest than a phone list: every row carries the address to call plus everything the platform already knows about who is there.
saying these in an interview costs you the question
- Thinks every target must be listed by hand in the configuration
- Believes metadata labels are stored on the resulting time series
- Cannot name any discovery mechanism beyond a static list
- Thinks changing a file-based target file requires restarting the server
- Assumes discovery runs once at startup rather than continuously