skip to content

How does Kafka handle Kerberos ticket renewal, and what does sasl.kerberos.ticket.renew.window.factor control?

level: seniorimportance: should knowfreq 35%

answer

  1. renew.window.factor default 0.8 = renew at 80%
  2. jitter 0.05 avoids herd
  3. min.time.before.relogin 60s floor
  4. renew vs full keytab relogin
  5. renew-until cap set by KDC maxrenewlife

basics

~20 s

Kafka runs a background login thread that re-logs in or renews the Kerberos ticket before it expires. sasl.kerberos.ticket.renew.window.factor (default 0.8) sets how far into the ticket's lifetime to wait before renewing — at 80% of its life.

solid answer

~50 s

Kerberos tickets (TGTs) expire, so Kafka's KerberosLogin/AbstractLogin starts a refresh thread that renews or re-acquires the ticket before expiry, keeping long-running producers, consumers, and brokers authenticated without restarts. sasl.kerberos.ticket.renew.window.factor (default 0.8) controls timing: the thread sleeps until that fraction of the time between the ticket's issue and its renew-until/expiry has elapsed — so at 0.8 it acts at ~80% of the window. Two related knobs: sasl.kerberos.min.time.before.relogin (default 60000 ms) is the minimum wait between relogin attempts to avoid hammering the KDC, and sasl.kerberos.ticket.renew.jitter (default 0.05) adds randomness so many clients do not renew simultaneously. If the ticket is renewable, Kafka renews it; if it has hit renew-until, Kafka does a full relogin from the keytab. Renewal only works within the KDC's max renewable lifetime — beyond that a fresh login from the keytab is mandatory.

go deeper

for a junior

Know that Kafka renews Kerberos tickets automatically so long-running clients stay authenticated.

for a middle

Recall the 0.8 default and that there is also jitter and a min-time-before-relogin floor.

for a senior

Distinguish renew vs full keytab relogin, tie renew-until to KDC maxrenewlife, and tune the factor for safety margin.

for a principal

Design KDC policy (ticket/renew lifetimes) and jitter/factor settings to keep a large fleet authenticated without overloading the KDC.

## Why renewal exists Kerberos credentials are time-bounded. A **TGT** (Ticket-Granting Ticket) has a **lifetime** (e.g. 10 hours) and, if renewable, a **renew-until / max renewable lifetime** (e.g. 7 days). A Kafka process that runs for days would lose authentication the moment its ticket expired — unless something keeps it fresh. Kafka does this automatically via a background **refresh/relogin thread** in its Kerberos login implementation (historically `KerberosLogin`). ## The renewal window The thread does not wait until the very last second (risky) nor renew constantly (wasteful, KDC load). It computes a wake-up time as a fraction of the ticket's validity window: - `sasl.kerberos.ticket.renew.window.factor` (default **0.8**): renew after 80% of the time from the ticket's start to its expiry/renew-until has passed. Lower it to renew earlier (more safety margin), raise it to renew later. - `sasl.kerberos.ticket.renew.jitter` (default **0.05**): adds +/- random jitter to that time so a fleet of clients does not all hit the KDC at the same instant (thundering herd). - `sasl.kerberos.min.time.before.relogin` (default **60000 ms = 60s**): a floor on how often relogin can happen, protecting the KDC from rapid retries (e.g. when something is failing). ## Renew vs re-login - **Renew**: if the ticket is still within its **renew-until** limit, Kafka extends it via the KDC — cheaper, keeps the same login. - **Full relogin**: once the renewable lifetime is exhausted, renewal is impossible; Kafka must authenticate again from the **keytab** (which is why services need a keytab, not just a cached ticket). If you only used `useTicketCache=true` with no keytab, the process eventually cannot re-authenticate and fails. ## KDC-side limits The achievable renewal is capped by the KDC's policy: `ticket_lifetime`, `renew_lifetime` in krb5.conf `[libdefaults]`, and per-principal `maxlife`/`maxrenewlife` set on the KDC. If the principal is not flagged renewable on the KDC, Kafka's factor settings cannot extend anything; it will fall back to full relogin from the keytab each lifetime. ## Edge cases and pitfalls - Tickets not marked renewable on the KDC -> no renewal possible; relies on keytab relogin. - Clock skew beyond ~5 minutes breaks both renewal and validation. - Setting the factor too high (near 1.0) risks the ticket expiring before renewal completes under load. - With only a ticket cache (no keytab), once renew-until is reached the client cannot recover — use keytabs for services. ## Quick reference defaults - renew.window.factor: 0.8 - renew.jitter: 0.05 - min.time.before.relogin: 60000 ms

  • What happens when a ticket reaches its renew-until limit?
    Renewal is no longer possible; Kafka must do a full relogin from the keytab. A client using only a ticket cache (no keytab) cannot recover and will fail.
  • Why does sasl.kerberos.ticket.renew.jitter exist?
    It randomizes renewal timing so a large fleet of clients does not all hit the KDC at the same moment, avoiding a thundering-herd load spike.

saying these in an interview costs you the question

  • Saying Kafka must be restarted when a Kerberos ticket expires (it auto-renews/relogs in)
  • Claiming the factor renews after a fixed number of hours rather than a fraction of the ticket window
  • Thinking renewal works indefinitely without a keytab once renew-until is hit

context