In Linkerd, what are the "golden metrics" that `linkerd viz stat` reports for a meshed workload, where do those numbers come from, and why might a meshed service show no success rate at all?
answer
- success rate, RPS, latency percentiles
- measured by the proxy, not your code
- viz is a separate install from the core
- 5xx means failure; 4xx does not
- no L7 protocol means no request metrics
basics
~20 slinkerd viz stat reports success rate, requests per second and latency percentiles, computed from metrics that the linkerd2-proxy sidecars expose and the viz extension's Prometheus scrapes. A workload shows none of them when its traffic is not HTTP or gRPC.
solid answer
~40 sThe golden metrics are success rate, request rate (RPS) and latency percentiles — typically p50, p95 and p99. They are not instrumented by your application: each `linkerd2-proxy` sidecar counts and times the requests passing through it and exposes them on its admin endpoint, and the `linkerd-viz` extension's Prometheus scrapes those proxies. `linkerd viz stat deploy -n myapp` then queries that data, and `--from`/`--to` narrow it to one caller-callee edge. Success rate means non-5xx for HTTP, or a non-failure `grpc-status` for gRPC. A workload shows an empty success rate when the proxy never saw requests it could classify: the traffic is not HTTP or gRPC, the port is marked opaque, the calling side is unmeshed so no outbound metrics exist, or nothing has been called in the window Prometheus retains.
go deeper
Recall the three golden metrics by name and that Linkerd's sidecar produces them for free. Be able to run linkerd viz stat deploy -n <namespace> and say what each column means.
Explain the pipeline end to end — proxy counters, viz Prometheus scrape, metrics API, CLI — and how success is classified for HTTP versus gRPC, including why 4xx does not count as failure.
Diagnose an empty stat output: unmeshed caller, opaque or non-HTTP port, missing viz extension, expired retention. Be ready to say how you would move these metrics into durable storage for alerting.
Own the observability strategy: what a mesh gives you for free, where it stops (it sees hops, not business outcomes or traces without further work), and how mesh metrics fit alongside application instrumentation without duplicating cost.
## What the golden metrics are "Golden metrics" in Linkerd means the three signals you want first when a service misbehaves: - **Success rate** — the proportion of requests the proxy classified as successful. - **RPS** — requests per second through that workload. - **Latency percentiles** — usually p50, p95 and p99. The selling point is that you get them for every meshed workload without touching application code, because the measurement is made by the proxy in the request path rather than by a library you had to add. ## Where the numbers come from Every injected pod runs the `linkerd2-proxy` container. It counts requests and records latency histograms, labelled with the workload, namespace and — crucially — the peer on the other end of the call. It exposes those counters on its admin port for scraping. The `linkerd-viz` extension supplies the rest: a Prometheus that scrapes every proxy, a metrics API in front of it, the `linkerd viz` CLI subcommands, and the dashboard. Installing the core control plane alone gives you mTLS and traffic management but no `stat` output — viz is a separate install step, and forgetting it is a common reason the command returns nothing. The everyday commands: ```bash linkerd viz stat deploy -n emojivoto linkerd viz stat deploy/web --from deploy/vote-bot -n emojivoto linkerd viz edges deploy -n emojivoto linkerd viz tap deploy/web -n emojivoto ``` `stat` aggregates over a window; `--from` and `--to` restrict it to a single caller-callee edge, which is how you tell "my service is failing" apart from "one specific caller is getting failures from me". `edges` shows who actually talks to whom, including whether the connection is mTLS'd. `tap` is different in kind: it streams a live sample of individual requests with method, path, response code and duration, which is what you reach for when the aggregate says 98% and you need to see the failing 2%. ## How success is classified For HTTP, the proxy counts a response as a failure when its status is 5xx. Anything else — including 4xx — counts as success, which surprises people: a service returning 100% 404s shows a 100% success rate, because from the mesh's point of view the server answered correctly and the client asked for the wrong thing. For gRPC, classification uses the `grpc-status` trailer rather than the HTTP status, since gRPC returns 200 at the HTTP layer even for application errors. ## Why a workload can show no success rate This is the diagnostic half of the question, and it has a small set of causes: 1. **The traffic is not HTTP or gRPC.** Linkerd's per-request metrics require a protocol it can parse. A database or a custom binary protocol yields connection-level TCP metrics only, so there is no success rate or RPS to report. 2. **The port is declared opaque.** `config.linkerd.io/opaque-ports` tells the proxy to skip protocol detection deliberately — the correct setting for non-HTTP traffic, and it removes L7 metrics by design. 3. **The pod is not meshed.** No sidecar, no metrics. `linkerd viz stat` will show the workload with empty columns; check that the namespace or pod carries `linkerd.io/inject: enabled` and that the pod really has the extra container. 4. **The caller is unmeshed.** Metrics exist per hop. If an unmeshed client calls a meshed server, the server's inbound metrics exist but there is no outbound record from the caller, so an edge-scoped `--from` query comes back empty. 5. **Nothing happened in the window.** The bundled Prometheus keeps a short retention window suited to live debugging, not to historical analysis. If the traffic was hours ago, or the viz Prometheus was restarted, the data is simply gone. ## The operational caveat worth raising The on-cluster Prometheus that ships with `linkerd-viz` is intended for interactive debugging. For dashboards, alerting and post-incident review you federate or scrape the proxies into your own long-term Prometheus instead of relying on the bundled one. Saying that unprompted signals that you have run this in production rather than only through the tutorial.
- A meshed service returns nothing but 404s, yet `linkerd viz stat` shows a 100% success rate. Is that a bug?No. Linkerd classifies HTTP responses as failures only on 5xx, so a well-formed 404 counts as a successful exchange — the server answered. If 4xx rates matter to you, look at per-route data or `linkerd viz tap`, which shows individual response codes, and alert on those separately.
- When would you use `linkerd viz tap` instead of `linkerd viz stat`?When the aggregate is not enough. `stat` gives rates over a window; `tap` streams a live sample of individual requests with path, method, status and duration, so you can see which specific requests are failing or slow. It is a sampling debug tool, not a metrics source, and it costs proxy work while running.
- Would you point production alerting at the Prometheus that ships with linkerd-viz?No. That instance is sized for interactive debugging with a short retention window and no durability guarantees — a restart loses history. Federate it or scrape the proxy metrics endpoints into your own long-term Prometheus, and build dashboards and alerts there.
saying these in an interview costs you the question
- Thinks the application must be instrumented to get the metrics
- Assumes 4xx responses lower the reported success rate
- Expects stat to work without installing the viz extension
- Treats the bundled Prometheus as a long-term metrics store
- Believes every protocol produces per-request metrics