skip to content

What does Kafka's expired-connections-killed-count metric measure, and which configuration drives it?

level: seniorimportance: should knowfreq 35%

answer

  1. KIP-368 re-authentication
  2. connections.max.reauth.ms
  3. expired-connections-killed-count
  4. OAuth/Kerberos short-lived creds
  5. kill on lapsed credential, not initial login

basics

~10 s

expired-connections-killed-count counts connections the broker closed because a SASL credential (e.g. an OAuth token or Kerberos ticket) expired and the client did not re-authenticate in time. It is driven by connections.max.reauth.ms (KIP-368).

solid answer

~40 s

KIP-368 lets brokers enforce that long-lived connections re-authenticate before their SASL credential expires. You enable it per listener with listener.name.<name>.<mechanism>.connections.max.reauth.ms (or the listener-wide connections.max.reauth.ms). When set, the broker tells the client the credential lifetime; the client must re-authenticate within that window. If a connection's credential expires and it neither re-authenticated nor was closed, the broker forcibly kills it and increments the expired-connections-killed-count sensor under kafka.server:type=socket-server-metrics. Monitoring this metric matters for security audit because a rising count means clients are running with expired credentials (e.g. stale OAuth tokens) and being cut off — a sign of token-refresh problems or attempts to keep using lapsed credentials. You'd graph it per listener alongside failed-reauthentication-total and alert on sustained nonzero values.

go deeper

for a junior

Recognize the metric name relates to connections being closed for expired credentials.

for a middle

Tie it to connections.max.reauth.ms and short-lived OAuth/Kerberos credentials.

for a senior

Explain the KIP-368 re-auth handshake, default-off behavior, client-compatibility constraints, and how the metric signals credential-rotation problems.

for a principal

Set policy for credential lifetimes and reauth windows fleet-wide, plan client-compatibility rollout, and integrate the signal into security dashboards.

## The problem KIP-368 solves Before KIP-368, a Kafka connection authenticated **once** at connect time and then lived indefinitely. With short-lived credentials — OAuth bearer tokens, Kerberos tickets, delegation tokens — a connection could keep operating long after its credential had *expired*, because the broker never re-checked. That is a security gap: a revoked/expired identity stays effective for the life of the TCP connection. ## What the config does `connections.max.reauth.ms` (settable per listener+mechanism via `listener.name.<listener>.<saslMechanism>.connections.max.reauth.ms`, or as a listener-wide default) tells the broker to **require re-authentication** on the connection within a bounded window. Mechanics: 1. On authentication, the broker computes a session lifetime = min(`connections.max.reauth.ms`, remaining credential lifetime) and communicates it to a KIP-368-aware client. 2. The client must complete a SASL **re-authentication** handshake before that window closes. 3. If the credential expires and the client has **not** re-authenticated, the broker **closes (kills) the connection**. ## The metric When the broker kills such a connection, it increments: ``` kafka.server:type=socket-server-metrics,listener=<NAME> -> expired-connections-killed-count ``` Related sensors on the same MBean: `successful-reauthentication-total`, `failed-reauthentication-total`, `reauthentication-latency-{avg,max}`. ## Why it matters for audit/monitoring - A **nonzero, growing** `expired-connections-killed-count` means clients are operating with expired credentials and getting forcibly disconnected. Causes: broken OAuth token-refresh logic, clock skew making tokens look expired, old client libraries that don't support re-auth, or someone trying to ride a lapsed credential. - It is a direct, quantifiable signal that your credential-rotation hygiene is working *or* failing. ## Edge cases / gotchas - **Old clients** that predate KIP-368 cannot re-authenticate; if you enforce a reauth window they will simply be disconnected and must reconnect — so rolling this out requires compatible clients. - Setting `connections.max.reauth.ms=0` (the default) disables enforcement — no killing, metric stays flat. You must opt in. - The kill is at the connection layer; the client typically reconnects and re-authenticates, so application impact is usually a brief blip, but a storm of kills indicates a real refresh problem. - This is distinct from `failed-authentication-total` (initial login failures) — expired-connections-killed-count is specifically about *already-established* sessions whose credential lapsed.

  • What is the default value of connections.max.reauth.ms and what does it imply?
    The default is 0, which disables KIP-368 re-authentication enforcement entirely. Connections live indefinitely on their initial credential, and expired-connections-killed-count stays at zero, so you must explicitly opt in to get the security benefit.
  • How does expired-connections-killed-count differ from failed-authentication-total?
    failed-authentication-total counts initial login attempts that fail. expired-connections-killed-count counts already-authenticated, long-lived connections that the broker forcibly closes because their credential expired and they did not re-authenticate in time.

saying these in an interview costs you the question

  • Confusing it with initial-login failures (failed-authentication-total).
  • Claiming it is on by default (default connections.max.reauth.ms is 0 / disabled).
  • Saying it works without client support (pre-KIP-368 clients can't re-auth and just get disconnected).

context