skip to content

Explain KIP-714 client telemetry push: what problem it solves, how the protocol works, and how it relates to the MetricsReporter SPI.

level: seniorimportance: should knowfreq 35%

answer

  1. broker-side collection, clients PUSH metrics
  2. GetTelemetrySubscriptions + PushTelemetry RPCs
  3. OTLP/OpenTelemetry protobuf encoding
  4. kafka-client-metrics.sh subscriptions, client-instance-id
  5. broker plugin: ClientTelemetryReceiver / ClientTelemetry; GA 3.7

basics

~20 s

KIP-714 lets brokers collect standardized client metrics by having clients push them to the broker over the Kafka protocol, instead of operators scraping each client's JMX. The broker subscribes clients to metrics, clients push them as OpenTelemetry-encoded payloads on an interval.

solid answer

~50 s

KIP-714 (Client metrics and observability, GA in Kafka 3.7) solves the problem that client-side JMX is often unreachable to the people running the cluster — clients live in many apps/teams. Instead, the broker becomes the collection point. An admin defines client metrics subscriptions (`kafka-client-metrics.sh`, stored as ClientMetrics configs) selecting clients by client-id/matching and a push interval. Clients fetch their subscription via the new `GetTelemetrySubscriptions` API, then periodically send a `PushTelemetry` request carrying metrics encoded in OpenTelemetry (OTLP) protobuf format. The broker forwards these to a server-side plugin implementing the `ClientTelemetryReceiver` interface (a `MetricsReporter` that also implements `ClientTelemetry`), which exports them onward (e.g. to an OTel collector). It complements, not replaces, the local MetricsReporter SPI: clients still expose JMX, but now the broker can centrally subscribe to and receive a standardized OTLP metric set without per-client scraping.

go deeper

for a junior

Know it exists: brokers can centrally collect client metrics that clients push, instead of scraping each client's JMX.

for a middle

Explain the push direction, the subscription concept, and that payloads are OpenTelemetry-encoded.

for a senior

Walk the GetTelemetrySubscriptions/PushTelemetry RPCs, OTLP encoding, broker ClientTelemetryReceiver plugin, and version requirements.

for a principal

Design a fleet-wide observability strategy: subscription scoping, interval/cost control, receiver-to-collector pipeline, and how it coexists with existing JMX/Prometheus.

**The problem.** Historically, the only way to see Kafka *client* metrics was JMX on each client JVM (or a per-client MetricsReporter the app team installed). The cluster operators — the people who actually need to diagnose 'why is this client slow / mis-configured' — usually have no access to those client JVMs. KIP-714 ('Client metrics and observability') inverts the flow: clients *push* a standardized metric set to the broker, so the operator gets visibility from the cluster side. It went GA in Apache Kafka 3.7. **The protocol (two new RPCs):** 1. `GetTelemetrySubscriptions` — on connect, a supporting client asks the broker what metrics it should report. The broker replies with: a unique `client-instance-id` (a UUID identifying this client instance, also exposed via `KafkaProducer/Consumer.clientInstanceId()`), the list of subscribed metric name prefixes, the accepted compression types, and the push interval (`PushIntervalMs`). 2. `PushTelemetry` — the client periodically sends its current metric values, serialized as **OpenTelemetry Metrics Protocol (OTLP) protobuf**, optionally compressed. The broker acknowledges; back-off and a terminating flag on close are part of the flow. **Subscriptions / admin surface.** An operator creates *client metric subscriptions* using `kafka-client-metrics.sh` (or the Admin API / `ConfigResource` of type CLIENT_METRICS). A subscription has a name, a `metrics` prefix list (empty = all), an `interval.ms`, and `match` rules (e.g. `client_id`, `client_software_name`, `client_software_version`, transactional id) to target subsets of clients. This is dynamic config — no client restart needed; the next `GetTelemetrySubscriptions` picks it up. **Server-side plugin.** The broker doesn't store metrics; it hands each received payload to a plugin. You implement `org.apache.kafka.server.telemetry.ClientTelemetryReceiver` (exposed through a `MetricsReporter` that also implements `org.apache.kafka.common.telemetry.ClientTelemetry`), registered via the broker's `metric.reporters`. The receiver gets the decoded payload + context (client instance id, etc.) and forwards it to your observability backend (commonly an OpenTelemetry collector). Confluent and others ship such receivers. **Standardized metric set.** KIP-714 defines canonical client metric names (e.g. `org.apache.kafka.producer.*`, connection, request latency, throttling, errors) with OTel-compatible semantics, so all client languages report comparably — important because non-Java clients (librdkafka-based) also implement the push protocol. **Relationship to MetricsReporter SPI:** - The *client-side* `MetricsReporter`/`metric.reporters` SPI (covered separately) is local export (JMX or your own sink) and is unchanged. - KIP-714 reuses that machinery: the client's telemetry sender is itself wired through the metrics subsystem, and the *broker-side* receiver is a `MetricsReporter` + `ClientTelemetry`. So KIP-714 is a broker-driven, OTLP-standardized push *on top of* the existing SPI, not a replacement. **Edge cases / controls:** - Clients can disable participation with `enable.metrics.push=false` (default true on supporting clients). - If no subscription matches, clients push nothing — zero overhead. - The broker controls the interval, so operators can throttle volume centrally; clients honor the returned `PushIntervalMs` and back-off. - `clientInstanceId()` lets you correlate a broker-collected series back to a specific client instance. - It requires both broker and client support (3.7+ for full GA); mixed-version fleets simply fall back to no push for old clients.

  • Where do the metrics actually go once the broker receives a PushTelemetry request — does Kafka store them?
    No. The broker decodes the OTLP payload and hands it to a registered ClientTelemetryReceiver plugin (a MetricsReporter implementing ClientTelemetry), which forwards them to an external backend such as an OpenTelemetry collector. Kafka itself doesn't persist client telemetry.
  • How does an operator target only a subset of clients for telemetry, and is a restart needed?
    Create a client metrics subscription (kafka-client-metrics.sh / CLIENT_METRICS dynamic config) with match rules on client_id, client_software_name/version, etc., plus a metrics prefix list and interval.ms. It's dynamic config — clients pick it up on their next GetTelemetrySubscriptions, no restart.
  • Why use OTLP/OpenTelemetry encoding rather than Kafka's own metric format?
    OTLP is a vendor-neutral standard, so any observability backend can ingest it and non-Java clients report comparable semantics. It makes the cross-language client metric set uniform and pipeable straight into an OTel collector.

saying these in an interview costs you the question

  • Saying KIP-714 replaces JMX/MetricsReporter — it complements them; clients still expose local metrics.
  • Claiming the broker stores/queries telemetry like a TSDB — it forwards to a plugin/collector and persists nothing.
  • Saying clients decide what to report — the broker's subscription drives the metric set and interval.
  • Describing it as a pull/scrape model — it's a client-initiated push over the Kafka protocol.
  • Forgetting the OTLP/OpenTelemetry encoding and that it requires 3.7+ broker and client support.

context