What does Kafka's expired-connections-killed-count metric measure, and which configuration drives it?
answer
- KIP-368 re-authentication
- connections.max.reauth.ms
- expired-connections-killed-count
- OAuth/Kerberos short-lived creds
- kill on lapsed credential, not initial login
basics
~10 sexpired-connections-killed-count counts connections the broker closed because a SASL credential (e.g. an OAuth token or Kerberos ticket) expired and the client did not re-authenticate in time. It is driven by connections.max.reauth.ms (KIP-368).
solid answer
~40 sKIP-368 lets brokers enforce that long-lived connections re-authenticate before their SASL credential expires. You enable it per listener with listener.name.<name>.<mechanism>.connections.max.reauth.ms (or the listener-wide connections.max.reauth.ms). When set, the broker tells the client the credential lifetime; the client must re-authenticate within that window. If a connection's credential expires and it neither re-authenticated nor was closed, the broker forcibly kills it and increments the expired-connections-killed-count sensor under kafka.server:type=socket-server-metrics. Monitoring this metric matters for security audit because a rising count means clients are running with expired credentials (e.g. stale OAuth tokens) and being cut off — a sign of token-refresh problems or attempts to keep using lapsed credentials. You'd graph it per listener alongside failed-reauthentication-total and alert on sustained nonzero values.
go deeper
Recognize the metric name relates to connections being closed for expired credentials.
Tie it to connections.max.reauth.ms and short-lived OAuth/Kerberos credentials.
Explain the KIP-368 re-auth handshake, default-off behavior, client-compatibility constraints, and how the metric signals credential-rotation problems.
Set policy for credential lifetimes and reauth windows fleet-wide, plan client-compatibility rollout, and integrate the signal into security dashboards.
## The problem KIP-368 solves Before KIP-368, a Kafka connection authenticated **once** at connect time and then lived indefinitely. With short-lived credentials — OAuth bearer tokens, Kerberos tickets, delegation tokens — a connection could keep operating long after its credential had *expired*, because the broker never re-checked. That is a security gap: a revoked/expired identity stays effective for the life of the TCP connection. ## What the config does `connections.max.reauth.ms` (settable per listener+mechanism via `listener.name.<listener>.<saslMechanism>.connections.max.reauth.ms`, or as a listener-wide default) tells the broker to **require re-authentication** on the connection within a bounded window. Mechanics: 1. On authentication, the broker computes a session lifetime = min(`connections.max.reauth.ms`, remaining credential lifetime) and communicates it to a KIP-368-aware client. 2. The client must complete a SASL **re-authentication** handshake before that window closes. 3. If the credential expires and the client has **not** re-authenticated, the broker **closes (kills) the connection**. ## The metric When the broker kills such a connection, it increments: ``` kafka.server:type=socket-server-metrics,listener=<NAME> -> expired-connections-killed-count ``` Related sensors on the same MBean: `successful-reauthentication-total`, `failed-reauthentication-total`, `reauthentication-latency-{avg,max}`. ## Why it matters for audit/monitoring - A **nonzero, growing** `expired-connections-killed-count` means clients are operating with expired credentials and getting forcibly disconnected. Causes: broken OAuth token-refresh logic, clock skew making tokens look expired, old client libraries that don't support re-auth, or someone trying to ride a lapsed credential. - It is a direct, quantifiable signal that your credential-rotation hygiene is working *or* failing. ## Edge cases / gotchas - **Old clients** that predate KIP-368 cannot re-authenticate; if you enforce a reauth window they will simply be disconnected and must reconnect — so rolling this out requires compatible clients. - Setting `connections.max.reauth.ms=0` (the default) disables enforcement — no killing, metric stays flat. You must opt in. - The kill is at the connection layer; the client typically reconnects and re-authenticates, so application impact is usually a brief blip, but a storm of kills indicates a real refresh problem. - This is distinct from `failed-authentication-total` (initial login failures) — expired-connections-killed-count is specifically about *already-established* sessions whose credential lapsed.
- What is the default value of connections.max.reauth.ms and what does it imply?The default is 0, which disables KIP-368 re-authentication enforcement entirely. Connections live indefinitely on their initial credential, and expired-connections-killed-count stays at zero, so you must explicitly opt in to get the security benefit.
- How does expired-connections-killed-count differ from failed-authentication-total?failed-authentication-total counts initial login attempts that fail. expired-connections-killed-count counts already-authenticated, long-lived connections that the broker forcibly closes because their credential expired and they did not re-authenticate in time.
saying these in an interview costs you the question
- Confusing it with initial-login failures (failed-authentication-total).
- Claiming it is on by default (default connections.max.reauth.ms is 0 / disabled).
- Saying it works without client support (pre-KIP-368 clients can't re-auth and just get disconnected).