skip to content

In a Prometheus estate, what do node_exporter, cAdvisor and blackbox_exporter each measure, and when does only a probe answer the question?

level: middleimportance: should knowfreq 64%

answer

  1. Three different vantage points
  2. Host, container, and from outside
  3. One of them owns no internal state
  4. Certificates are invisible from inside a process
  5. probe_success is a verdict, not a cause

basics

~20 s

node_exporter reports one host's OS and hardware metrics, cAdvisor reports per-container resource use from control-group accounting, and blackbox_exporter probes an endpoint from outside over HTTP, TCP, ICMP or DNS. Only the probe sees the path a client actually travels.

solid answer

~50 s

All three are separate processes that Prometheus scrapes, and they differ by vantage point. `node_exporter` runs one per machine and reads the kernel's own interfaces: CPU, memory, load, filesystem fill, disk and network I/O. cAdvisor reads control-group accounting to attribute CPU, memory and I/O to individual containers — on Kubernetes the kubelet already exposes it, so deploying a second copy usually just duplicates series. `blackbox_exporter` reads nothing internal at all: it makes a request to an endpoint and reports what came back, over HTTP, TCP, ICMP, DNS or gRPC, as `probe_success`, `probe_duration_seconds` and module-specific series. The probe earns its place where instrumentation cannot reach: DNS, a load balancer or proxy you do not own, an expiring TLS certificate, a vendor appliance with no metrics, and the case where the target is entirely dead and therefore cannot report anything about itself.

code

bash · 4 lines
bash
curl -s 'http://blackbox:9115/probe?module=http_2xx&target=https://planner.windfarm.example/status'
# probe_success 1
# probe_duration_seconds 0.184
# probe_http_status_code 200

go deeper

for a junior

Be able to say that these are separate processes Prometheus scrapes, and match each to its scope: one machine, one container, one endpoint checked from outside. Knowing the metric prefixes helps you read an unfamiliar dashboard.

for a middle

Explain what each one reads and why the readings differ: kernel interfaces, control-group accounting, and an actual request. Be ready to say why the probe is configured with the target as a parameter rather than one exporter per checked endpoint.

for a senior

Show judgement about coverage and cost: which failures only an outside-in check catches, where a probe's vantage point makes it misleading, and how you avoid duplicating container metrics an existing component already exposes.

for a principal

Own the fleet policy: what every host and container gets by default, which external dependencies are worth probing and from where, and how you keep the probe layer from becoming a single process whose death silences every availability signal at once.

## Three different vantage points In Prometheus terminology an **exporter** is a separate process that Prometheus scrapes, which translates some system's state into metrics — not a component inside your application shipping telemetry outwards. The three most common ones differ mainly in where they stand relative to the thing being measured. | Exporter | Runs | Reads | Answers | |---|---|---|---| | `node_exporter` | one per Unix-like host | the kernel's own interfaces, such as `/proc` and `/sys` | is this machine healthy: CPU, memory, load, filesystem fill, disk and network I/O | | cAdvisor | one per container host | control-group accounting for each container | which container is consuming what: CPU seconds, working-set memory, container filesystem and network | | `blackbox_exporter` | anywhere you choose | nothing internal at all — it makes a request | can this endpoint be reached, and how well, from where I am standing | Their series carry distinguishable prefixes — `node_` for host metrics, `container_` for cAdvisor's per-container series, `probe_` for probe results — which is a useful clue when you meet an unfamiliar dashboard. On Kubernetes, cAdvisor is normally not deployed separately at all: the kubelet already exposes container metrics gathered this way, so an extra deployment usually means duplicated series under a second `instance` label. The blackbox exporter also works differently in shape. One process serves many probed endpoints, because the address to probe and the module to use are handed to it as query parameters on its probe path (`/probe?module=http_2xx&target=…`), and the modules themselves — an HTTP check expecting a 2xx, a TCP connect, an ICMP ping, a DNS lookup — are defined in the exporter's own configuration file. It reports results as `probe_success`, `probe_duration_seconds`, and module-specific series such as `probe_http_status_code` or `probe_ssl_earliest_cert_expiry`. ## What only a probe can tell you Instrumenting a service tells you what the service thinks happened. A probe tells you what a client experiences, and the gap between the two contains most of the interesting failures: - **The parts of the path you do not own.** DNS resolution, a load balancer, a reverse proxy, a service mesh, a corporate firewall: none of them appear in the service's own metrics, and any of them can be the outage. - **The certificate.** An expiring TLS certificate is invisible from inside the process — the application is perfectly healthy right up to the moment nobody can talk to it. A probe reports the presented certificate's expiry, which is one of the highest-value alerts in a small monitoring setup. - **Total absence.** A process that has crashed, a host that has lost power, a network that has partitioned: none of them can report anything. A scrape failure tells you Prometheus could not reach it; a probe from a chosen vantage point tells you whether a client on that path can, which is a different question. - **Things with no instrumentation at all** — a third-party API, a vendor appliance, a printer, a site VPN concentrator. If you cannot put code in it, an outside-in check is the only measurement available. - **The user's route rather than the internal one.** Probing the public address exercises everything a customer traverses, which internal-only monitoring systematically skips. ## Where a probe misleads - It measures **one vantage point**. A probe running on the same host as the service it checks proves almost nothing about the network, and a probe from a single region says nothing about another. - `probe_success` is a verdict, not an explanation. Reading the per-phase durations and the status code series is what turns "the check failed" into "TLS handshake never completed". - A probe adds real traffic. Frequent probing of an expensive endpoint is load you are choosing to add, which matters most when the endpoint is already struggling. - One exporter process probing hundreds of endpoints is a concentration of blast radius: when it dies, every probe result stops at once, so the exporter's own scrape health deserves an alert of its own. ## Choosing in practice The wind-farm maintenance planner runs its services in containers across 6 hosts and serves an API with a 320 ms p99 budget, and its telemetry budget has just been halved. A defensible split: `node_exporter` on each host, because a full disk on the box is the failure nobody's application metrics predict; container metrics from the kubelet's existing exposure rather than a second deployment, since the duplicate is pure cost; and a small, deliberately chosen set of probes — the public planner API, the certificate on it, and each of the 11 site VPN endpoints that the planner cannot instrument because they are vendor appliances. What is cut first when budget is halved is usually the per-container series nobody queries, not the probes. The probes are a handful of series each and cover exactly the failures that would otherwise be discovered by a phone call from a site engineer.

  • Where should a blackbox probe run relative to what it probes?
    Somewhere that represents the client you care about. A probe running on the same host as the service tests the application and essentially nothing else — not DNS, not the load balancer, not the network. Put probes where your users are, or at least on the other side of the infrastructure you are trying to validate, and accept that each probe is a claim about one path only.
  • Should node_exporter run inside each container, one per container?
    No. It reports the host's numbers, so a copy per container gives you N identical readings of the same machine wearing N different instance labels: the CPU of one host counted many times, and cardinality multiplied for nothing. Run one per host as a daemon, with the host's kernel interfaces available to it, and use container-level accounting for per-container attribution.
  • A probe reports success but users report errors. What does that tell you?
    That the probe and the users are not doing the same thing. Common causes: the probe checks a lightweight status path rather than a real operation, it runs inside the network while users come from outside, it does not present the authentication or headers real clients send, or the failure is a slow response within the module's timeout. Probes are only as honest as the request they make.

saying these in an interview costs you the question

  • Says node_exporter reports the application's own request metrics
  • Deploys cAdvisor separately when the kubelet already exposes it
  • Believes a blackbox probe removes the need to instrument the service
  • Runs the probe on the probed host and calls it end-to-end
  • Treats probe_success as an explanation rather than a signal
  • Thinks these exporters run inside the application process