skip to content

Operations and Administration

Running Kafka day to day: sizing clusters, managing configs, reassigning partitions, rolling upgrades, DR, and the admin tooling. Interviewers use this area to find out whether you have operated Kafka or only written clients against it.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

page 2 of 2

What is kafka-log-dirs used for, and how does it help with disk capacity and partition placement?

level: seniorimportance: should knowfreq 40%

basics

~10 s

kafka-log-dirs.sh reports, per broker and per log directory, how much disk each partition replica uses. You use it to find disk hot spots, balance data across JBOD log.dirs, and confirm reassignment/throttling progress.

open as a page

What is broker.rack rack awareness, and how should it influence how you add or remove brokers in a cloud deployment across availability zones?

level: seniorimportance: should knowfreq 45%

basics

~20 s

broker.rack tags each broker with a fault domain (e.g., an availability zone). Kafka then spreads a partition's replicas across different racks, so losing one rack/AZ doesn't take all copies. When scaling or decommissioning, keep replicas balanced across racks.

open as a page

Why is OS page cache central to Kafka sizing, and how do you reserve page-cache and network headroom when provisioning brokers?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Kafka serves recent reads from the OS page cache (RAM) instead of disk, so consumers that keep up stay memory-fast. Give the JVM a modest heap (~5-6 GB) and leave most RAM free for page cache. For network, size NICs above peak replication + produce + fetch traffic, which RF multiplies.

open as a page

Explain the configuration precedence hierarchy in Kafka: how is the effective value of a broker/topic config resolved across per-broker, cluster-default, static, and built-in defaults?

level: seniorimportance: should knowfreq 50%

basics

~10 s

Kafka resolves the most specific source first: per-broker dynamic, then cluster-wide dynamic default, then the static server.properties value, then the hardcoded default. Topic-level overrides win over the broker default for that topic.

open as a page

How does Kafka protect sensitive dynamic broker configs (like SSL keystore passwords), and what is the role of the password encoder secret?

level: seniorimportance: should knowfreq 35%

basics

~10 s

Sensitive dynamic configs (e.g. passwords) are stored encrypted in the metadata store, not in plaintext. Each broker needs password.encoder.secret in server.properties to encrypt/decrypt them, and kafka-configs --describe shows such values as redacted/[hidden].

open as a page

What controller-health signals should you alert on in a Kafka cluster, and how does this differ between ZooKeeper-based and KRaft clusters?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Alert if the active controller count isn't exactly 1 (0 = no controller, >1 = split brain), and on rising controller queue size or slow leader elections. In KRaft, also watch the metadata quorum: leader presence, follower lag, and unfetched metadata.

open as a page

Walk through the on-call runbook you'd follow when a single broker is down and URP has spiked across many partitions. What do you check, fix, and escalate?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Confirm the alert and identify the down broker. Check if it's process-down vs disk-full vs network-isolated. If recoverable, restart/heal it and watch URP drop as replicas rejoin ISR. If data is lost, reassign replicas. Escalate if min.insync.replicas is breached or partitions go offline.

open as a page

What is LinkedIn Cruise Control, and how do its goals and self-healing automate cluster balancing?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Cruise Control is a tool that continuously monitors a Kafka cluster's load and automatically generates and executes reassignment plans to keep it balanced. It optimizes against a prioritized list of 'goals' (hard and soft constraints) and can self-heal by reacting to broker failures or anomalies.

open as a page

In a KRaft cluster, how do feature flags and metadata.version replace inter.broker.protocol.version, and how do you upgrade metadata.version during a rolling upgrade?

level: seniorimportance: should knowfreq 45%

basics

~20 s

KRaft has no ZooKeeper, so there is no inter.broker.protocol.version to set in server.properties. Instead the cluster has a feature flag named metadata.version that gates new behavior. You roll new binaries first (metadata.version stays put), then explicitly raise it with kafka-features.sh upgrade once all nodes are on the new version.

open as a page

Which metrics and signals do you monitor to verify tiered storage is healthy in production, and what does rising remote-copy lag indicate?

level: seniorimportance: should knowfreq 35%

basics

~10 s

Watch remote copy lag (bytes/segments waiting to upload), remote copy throughput, and upload/delete error rates, plus local disk usage. Rising copy lag means uploads can't keep up, so segments stay local and disk fills.

open as a page

What is the RemoteStorageManager (RSM) plugin, and what do you need to configure to run one against an object store like S3 in production?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The RSM is the pluggable component that physically reads/writes log segments and indexes to the remote backend. To run one against S3 you configure its class name plus backend settings: bucket, region/endpoint, credentials, chunk/part size, and optional encryption.

open as a page

Design a multi-region active-passive Kafka DR strategy. Cover replication choice, metadata/snapshot handling, failover and failback, and how you avoid split-brain and duplicate processing.

level: principalimportance: should knowfreq 30%

basics

~20 s

Run a passive DR cluster in another region, replicate topics asynchronously with MirrorMaker 2 or Cluster Linking, replicate consumer offsets and ACLs/config, and keep clients pointed at the active cluster. On disaster, translate offsets, repoint clients to DR, and run only one active cluster at a time to avoid split-brain. Failback re-syncs the original direction.

open as a page

How do you decide the number of brokers in a cluster, accounting for replication factor, broker-failure headroom, and rebalance capacity?

level: principalimportance: should knowfreq 40%

basics

~20 s

Pick brokers so total disk/network/partition load fits with each broker under ~60-70% utilization, you have at least RF brokers (plus spare for min.insync.replicas), and the cluster keeps working when N brokers fail. Survivors must absorb the dead brokers' load and re-replication.

open as a page

How would you structure Kafka alerting around SLOs to avoid alert fatigue — symptom-based vs cause-based alerts, escalation tiers, and what should and shouldn't page?

level: principalimportance: should knowfreq 42%

basics

~20 s

Page on a small set of symptom alerts tied to user-facing SLOs (availability, durability, freshness) — offline partitions, write failures, SLO-breaching lag. Route cause-based and predictive signals (disk filling, URP, rising controller queue) to tickets/dashboards, not pages, unless they breach an SLO.

open as a page

How does a partition reassignment preserve availability and avoid data loss while replicas are moving?

level: principalimportance: should knowfreq 35%

basics

~20 s

Kafka adds the new target replicas and waits for them to catch up and join the ISR before removing the old replicas, so the partition always keeps enough in-sync copies. Leadership only moves to a new replica once it is in the ISR, so no acknowledged data is lost.

open as a page

How do you do capacity and cost planning when offloading Kafka segments to remote object storage, and what hidden costs should you account for?

level: principalimportance: should knowfreq 30%

basics

~20 s

Model local disk from local.retention and ingest rate; model remote cost from total retention × ingest × replication-of-1 (remote stores one copy). Add hidden costs: API request charges (PUT/GET/LIST), egress on remote fetches, and the metadata topic.

open as a page

showing 31–46 of 46