Which roles can Prometheus's Kubernetes service discovery use, and why is the discovered set almost never what you want to scrape?
answer
- six roles, not one
- the role decides what becomes a target
- discovery is deliberately unfiltered
- one target per declared container port
- the annotation filters only via relabeling
basics
~20 sPrometheus's Kubernetes discovery runs in one of six roles: node, service, pod, endpoints, endpointslice or ingress. Each returns every matching object in the cluster, unfiltered by design, so relabeling rules decide which of them actually become scrape targets.
solid answer
~50 sA `kubernetes_sd_configs` block takes a `role`, and the role decides what an object turns into. `node` gives one target per cluster node; `pod` gives a target per declared container port; `endpoints` and `endpointslice` give a target per backing address behind a Service; `service` gives one per Service port addressed by the Service's DNS name; `ingress` gives one per ingress path and is meant for probing from outside. Each target carries `__meta_kubernetes_*` labels — namespace, pod name, node name, and the object's own labels and annotations. Discovery is deliberately unfiltered: it returns the whole cluster, including system namespaces, sidecars and ports that serve no metrics. `relabel_configs` does the filtering, usually a `keep` on a namespace or on an annotation your platform has agreed on. You can also narrow at the API level with `namespaces` or `selectors`, which is cheaper on a large cluster.
code
yaml · 11 linesscrape_configs:
- job_name: touring-logistics-pods
kubernetes_sd_configs:
- role: pod
namespaces:
names:
- touring-logistics
relabel_configs:
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
action: keep
regex: "true"go deeper
Recall that this mechanism watches the Kubernetes API and that the role setting decides what kind of object becomes a target. Know that pods and nodes are among the choices.
Name the roles and say what each yields per object, explain how annotations and labels arrive as metadata, and explain why a filter is mandatory rather than optional.
Demonstrate judgement about where to filter: coarse boundaries pushed to the API versus per-workload opt-in in relabeling, plus what a role returning nothing usually means about permissions.
Own the convention itself. Decide how teams declare that a workload should be scraped, who may change it, and how you keep discovery's cost on the cluster API from growing with every new scrape job.
`kubernetes_sd_configs` makes Prometheus a client of the Kubernetes API. It watches objects of one kind and turns each into targets, continuously, so pods that appear and vanish are picked up and retired without a config change. The single most important thing to understand is that the mechanism is intentionally **indiscriminate**: it reports what the cluster contains, not what is worth scraping. ## The roles and what each one yields | Role | One target per | Typical use | |---|---|---| | `node` | Cluster node, addressed at the Kubelet's port | Node and Kubelet-level metrics | | `pod` | Declared container port on each pod | Application metrics, per instance | | `endpoints` | Address behind a Service, per port | Application metrics via the Service's backing set | | `endpointslice` | Address in an EndpointSlice, per port | The same, scaling better on large Services | | `service` | Port of each Service, addressed by its DNS name | Black-box probing of a Service as a whole | | `ingress` | Path of each ingress | Black-box probing from the outside in | Two of these deserve a warning. The `service` role's address is the Service's **DNS name**, so a scrape is load balanced to some replica — perfectly correct for a probe that asks "is this Service answering", and quite wrong for per-instance metrics, where each scrape would hit an arbitrary pod and the series would be nonsense. The `pod` role creates a target for each **declared container port**, so a pod with an application port, an admin port and a sidecar's port is discovered three times; the endpoint that serves metrics is one of them. ## What comes attached Every target carries metadata under the `__meta_kubernetes_` prefix. The ones you reach for constantly are `__meta_kubernetes_namespace`, `__meta_kubernetes_pod_name`, `__meta_kubernetes_node_name`, `__meta_kubernetes_service_name`, and the generated families `__meta_kubernetes_pod_label_<name>` and `__meta_kubernetes_pod_annotation_<name>`, where characters that are not valid in a label name become underscores. That is why the annotation `prometheus.io/scrape` shows up as `__meta_kubernetes_pod_annotation_prometheus_io_scrape`. Note what that annotation is and is not. Prometheus attaches **no built-in meaning** to it. It is a widely copied convention that only works because somebody wrote a `keep` rule that tests it. A candidate who says the server "skips pods without the annotation" has misunderstood the whole mechanism. The mechanism also needs permission: the server must be able to list and watch the object kinds its roles cover. A role that returns nothing at all, on a healthy cluster, is usually an authorisation problem rather than a filtering one. ## Why the discovered set is never the set you want Take a touring-logistics platform whose cluster runs 1,847 pods across 34 namespaces. The `pod` role there discovers something close to 2,900 targets, and perhaps 260 of them serve metrics. The rest are: - Pods in system namespaces that nobody asked you to monitor. - Non-metrics ports: the application's own traffic port, an admin port, a health port. - Sidecars, each with its own declared ports. - Jobs and short-lived pods that will be gone before the next scrape. - Objects whose metrics are already collected by a different job, so scraping them again duplicates every series under a second job label. Scraping all of them is not merely wasteful. Every target that answers with something that is not an exposition payload becomes a failing target, and a targets page with thousands of red rows is a page nobody reads — which is exactly how an incident ends up unexplainable from the dashboards you already had. ## Narrowing it, and where There are two places to cut, and they are not equivalent. 1. **At the API, before discovery returns.** `namespaces` restricts the mechanism to named namespaces, and `selectors` pushes label and field selectors down to the API server. Less data crosses the wire and less is held in memory. Use this for coarse, stable boundaries. 2. **In `relabel_configs`, after discovery.** A `keep` on an annotation, a label or a port name. This is where per-workload opt-in belongs, because it is a policy your teams can act on without touching the Prometheus configuration. A healthy configuration usually does both: bound the blast radius at the API, then express the opt-in convention in relabeling. And keep the number of distinct discovery blocks modest — every job with its own `kubernetes_sd_configs` is another watch against the same API server.
- The endpoints role returns twelve targets for a single Service. Is that a misconfiguration?No — that is the role working correctly. It creates a target per backing address per port, so a Service with twelve ready replicas yields twelve targets, one per instance. That is what you want for application metrics, since each replica's counters are its own. You only want a single target when you are probing the Service as a unit, which is the `service` role's job.
- How would you reduce the load discovery itself puts on the API server in a large cluster?Push the filtering down: restrict the mechanism with `namespaces` and with `selectors`, so the API server returns fewer objects rather than the server discarding them afterwards. Then consolidate — fewer scrape jobs each with its own discovery block means fewer watches on the same objects. Choosing `endpointslice` over `endpoints` also helps on Services with very many backing addresses.
- Does Prometheus itself do anything with a prometheus.io/scrape annotation?Nothing. It surfaces every annotation as a `__meta_kubernetes_pod_annotation_*` label and stops there. The convention works only because a `keep` rule in the scrape job tests that label. Any other annotation name would work identically, and a cluster where nobody wrote the rule will scrape the annotated pods and the unannotated ones alike.
saying these in an interview costs you the question
- Uses the service role for per-instance metrics and scrapes a random replica
- Thinks setting the role alone filters out irrelevant workloads
- Believes prometheus.io/scrape is a built-in server feature
- Cannot say what the endpoints role returns for one Service
- Assumes one target per pod regardless of declared container ports
- Blames filtering when an empty role is really missing API permissions