skip to content

Which JMX metrics expose failed authentications and SSL/SASL handshake failures on a Kafka broker, and how would you alert on them?

level: middleimportance: must knowfreq 60%

answer

  1. MBean: socket-server-metrics
  2. failed-authentication-total / -rate
  3. per listener + networkProcessor
  4. SSL handshake + SASL both increment it
  5. re-auth: KIP-368 sensors

basics

~10 s

The broker exposes per-listener authentication counters under kafka.server:type=socket-server-metrics, including failed-authentication-total / failed-authentication-rate and successful-authentication-total. Scrape them via JMX (e.g. JMX exporter to Prometheus) and alert when the failed rate spikes.

solid answer

~40 s

Kafka's network/security layer emits authentication metrics under the JMX MBean kafka.server:type=socket-server-metrics, tagged by listener and networkProcessor. The key sensors are failed-authentication-total and failed-authentication-rate, successful-authentication-total/-rate, and on newer versions successful-reauthentication and failed-reauthentication for SASL re-auth (KIP-368). SSL handshake failures and SASL credential failures both increment failed-authentication. You expose these with a JMX-to-Prometheus exporter (jmx_exporter agent) or any JMX scraper, then alert on a sustained nonzero failed-authentication-rate or a sudden jump in failed-authentication-total per listener — that pattern indicates bad credentials, expired/untrusted certs, or a brute-force attempt. Pair it with the authorizer log (ACL denials) and expired-connection metrics for a fuller picture, and break alerts down per listener so an internal-listener cert problem is distinguishable from external client noise.

go deeper

for a junior

Know failed-authentication-total exists under socket-server-metrics and means logins are failing.

for a middle

Map SSL handshake and SASL failures to the same metric, scrape per listener via JMX exporter, and alert sensibly.

for a senior

Separate auth-metric signals from authorizer-log signals, reason about re-auth (KIP-368) and baseline-relative alerting.

for a principal

Design the org's broker observability: exporter rollout, per-listener SLOs, correlation of auth metrics with audit logs in the SIEM.

## Authentication vs authorization metrics Kafka distinguishes two failure domains: - **Authentication** — proving *who* you are (TLS mutual auth / SASL login). Failures here are exposed as **metrics**. - **Authorization** — what you're *allowed* to do (ACLs). Denials here are exposed via the **authorizer log**, not these metrics. This question is about the authentication metrics. ## The MBean and sensors The broker's socket server publishes security sensors under: ``` kafka.server:type=socket-server-metrics,listener=<NAME>,networkProcessor=<N> ``` Key attributes: - `successful-authentication-total` / `successful-authentication-rate` - `failed-authentication-total` / `failed-authentication-rate` - `successful-reauthentication-total`, `failed-reauthentication-total`, `reauthentication-latency` (SASL re-authentication, KIP-368) - `successful-authentication-no-reauth-total` A **failed-authentication** increment can come from: an SSL/TLS handshake failure (untrusted or expired client cert when `ssl.client.auth=required`), a SASL mechanism failure (wrong password, unknown SCRAM user, bad Kerberos ticket, invalid OAuth token), or a protocol mismatch. Note that the metrics are **per listener** (e.g. `INTERNAL`, `EXTERNAL`, `CONTROLLER`) and **per network processor thread**, so to get a cluster view you sum across processors and keep the listener dimension. ## How to collect and alert 1. Run a **JMX exporter** as a Java agent on each broker (the Prometheus `jmx_exporter` is standard), mapping `kafka.server:type=socket-server-metrics` attributes to time series. 2. In Prometheus/Grafana, build an alert such as: `sum by(listener)(rate(kafka_server_socketservermetrics_failed_authentication_total[5m])) > 0` sustained for several minutes, or a relative spike vs baseline. 3. Keep the **listener** label so you can tell apart: external client misconfig (high but maybe expected churn) vs internal/controller auth failures (almost always real and urgent). ## Edge cases / gotchas - A *small constant* failed rate is normal in the wild (clients with stale creds retrying). Alert on **deviation from baseline**, not absolute zero, to avoid noise. - SSL handshake failures may also surface in broker logs and as connection-close events; the metric is the cheap, aggregatable signal. - Re-authentication failures (KIP-368) matter when `connections.max.reauth.ms` is set — an expired-but-still-open connection that fails to re-auth gets closed; watch `failed-reauthentication-total`. - These are *broker* metrics; clients also expose their own connection metrics, but for audit you trust the broker side. - Resetting/restart zeroes the `-total` counters, so prefer `rate()`/`increase()` over raw totals in alerting.

  • If a client's TLS cert expires, which metric moves — an authentication metric or an authorization metric?
    An authentication metric: failed-authentication-total on the relevant listener increments, because the TLS handshake / mutual-auth step fails before any ACL check. ACL/authorization denials are unrelated and appear in the authorizer log.
  • Why alert on rate or increase rather than the raw failed-authentication-total?
    Totals are monotonic counters that reset to zero on broker restart, so absolute values are misleading. rate()/increase() over a window captures the actual surge and survives restarts, and lets you alert on deviation from baseline.

saying these in an interview costs you the question

  • Saying ACL denials show up in failed-authentication metrics (those are authorization, in the authorizer log).
  • Quoting a single cluster-wide counter and ignoring the per-listener/per-processor dimensions.
  • Alerting on raw -total counters that reset on restart instead of rate/increase.

context