Operations and Administration
Running Kafka day to day: sizing clusters, managing configs, reassigning partitions, rolling upgrades, DR, and the admin tooling. Interviewers use this area to find out whether you have operated Kafka or only written clients against it.
part ofApache Kafkaoverview, primer and where to startread it →on this pageshowhide
explore
- Cluster Sizing and Capacity Planning5 questions
- Broker and Topic Configuration Management6 questions
- Partition Reassignment and Data Balancing5 questions
- Adding, Removing and Decommissioning Brokers5 questions
- Rolling Upgrades and Version Compatibility5 questions
- Tiered and Remote Storage5 questions
- Backup, Disaster Recovery and Retention Ops5 questions
- Admin CLI and AdminClient5 questions
- Operational Monitoring and Alerting5 questions
questions
page 2 of 2What is kafka-log-dirs used for, and how does it help with disk capacity and partition placement?
basics
~10 skafka-log-dirs.sh reports, per broker and per log directory, how much disk each partition replica uses. You use it to find disk hot spots, balance data across JBOD log.dirs, and confirm reassignment/throttling progress.
What is broker.rack rack awareness, and how should it influence how you add or remove brokers in a cloud deployment across availability zones?
basics
~20 sbroker.rack tags each broker with a fault domain (e.g., an availability zone). Kafka then spreads a partition's replicas across different racks, so losing one rack/AZ doesn't take all copies. When scaling or decommissioning, keep replicas balanced across racks.
Why is OS page cache central to Kafka sizing, and how do you reserve page-cache and network headroom when provisioning brokers?
basics
~20 sKafka serves recent reads from the OS page cache (RAM) instead of disk, so consumers that keep up stay memory-fast. Give the JVM a modest heap (~5-6 GB) and leave most RAM free for page cache. For network, size NICs above peak replication + produce + fetch traffic, which RF multiplies.
Explain the configuration precedence hierarchy in Kafka: how is the effective value of a broker/topic config resolved across per-broker, cluster-default, static, and built-in defaults?
basics
~10 sKafka resolves the most specific source first: per-broker dynamic, then cluster-wide dynamic default, then the static server.properties value, then the hardcoded default. Topic-level overrides win over the broker default for that topic.
How does Kafka protect sensitive dynamic broker configs (like SSL keystore passwords), and what is the role of the password encoder secret?
basics
~10 sSensitive dynamic configs (e.g. passwords) are stored encrypted in the metadata store, not in plaintext. Each broker needs password.encoder.secret in server.properties to encrypt/decrypt them, and kafka-configs --describe shows such values as redacted/[hidden].
What controller-health signals should you alert on in a Kafka cluster, and how does this differ between ZooKeeper-based and KRaft clusters?
basics
~20 sAlert if the active controller count isn't exactly 1 (0 = no controller, >1 = split brain), and on rising controller queue size or slow leader elections. In KRaft, also watch the metadata quorum: leader presence, follower lag, and unfetched metadata.
Walk through the on-call runbook you'd follow when a single broker is down and URP has spiked across many partitions. What do you check, fix, and escalate?
basics
~20 sConfirm the alert and identify the down broker. Check if it's process-down vs disk-full vs network-isolated. If recoverable, restart/heal it and watch URP drop as replicas rejoin ISR. If data is lost, reassign replicas. Escalate if min.insync.replicas is breached or partitions go offline.
What is LinkedIn Cruise Control, and how do its goals and self-healing automate cluster balancing?
basics
~20 sCruise Control is a tool that continuously monitors a Kafka cluster's load and automatically generates and executes reassignment plans to keep it balanced. It optimizes against a prioritized list of 'goals' (hard and soft constraints) and can self-heal by reacting to broker failures or anomalies.
In a KRaft cluster, how do feature flags and metadata.version replace inter.broker.protocol.version, and how do you upgrade metadata.version during a rolling upgrade?
basics
~20 sKRaft has no ZooKeeper, so there is no inter.broker.protocol.version to set in server.properties. Instead the cluster has a feature flag named metadata.version that gates new behavior. You roll new binaries first (metadata.version stays put), then explicitly raise it with kafka-features.sh upgrade once all nodes are on the new version.
Which metrics and signals do you monitor to verify tiered storage is healthy in production, and what does rising remote-copy lag indicate?
basics
~10 sWatch remote copy lag (bytes/segments waiting to upload), remote copy throughput, and upload/delete error rates, plus local disk usage. Rising copy lag means uploads can't keep up, so segments stay local and disk fills.
What is the RemoteStorageManager (RSM) plugin, and what do you need to configure to run one against an object store like S3 in production?
basics
~20 sThe RSM is the pluggable component that physically reads/writes log segments and indexes to the remote backend. To run one against S3 you configure its class name plus backend settings: bucket, region/endpoint, credentials, chunk/part size, and optional encryption.
Design a multi-region active-passive Kafka DR strategy. Cover replication choice, metadata/snapshot handling, failover and failback, and how you avoid split-brain and duplicate processing.
basics
~20 sRun a passive DR cluster in another region, replicate topics asynchronously with MirrorMaker 2 or Cluster Linking, replicate consumer offsets and ACLs/config, and keep clients pointed at the active cluster. On disaster, translate offsets, repoint clients to DR, and run only one active cluster at a time to avoid split-brain. Failback re-syncs the original direction.
How do you decide the number of brokers in a cluster, accounting for replication factor, broker-failure headroom, and rebalance capacity?
basics
~20 sPick brokers so total disk/network/partition load fits with each broker under ~60-70% utilization, you have at least RF brokers (plus spare for min.insync.replicas), and the cluster keeps working when N brokers fail. Survivors must absorb the dead brokers' load and re-replication.
How would you structure Kafka alerting around SLOs to avoid alert fatigue — symptom-based vs cause-based alerts, escalation tiers, and what should and shouldn't page?
basics
~20 sPage on a small set of symptom alerts tied to user-facing SLOs (availability, durability, freshness) — offline partitions, write failures, SLO-breaching lag. Route cause-based and predictive signals (disk filling, URP, rising controller queue) to tickets/dashboards, not pages, unless they breach an SLO.
How does a partition reassignment preserve availability and avoid data loss while replicas are moving?
basics
~20 sKafka adds the new target replicas and waits for them to catch up and join the ISR before removing the old replicas, so the partition always keeps enough in-sync copies. Leadership only moves to a new replica once it is in the ISR, so no acknowledged data is lost.
How do you do capacity and cost planning when offloading Kafka segments to remote object storage, and what hidden costs should you account for?
basics
~20 sModel local disk from local.retention and ingest rate; model remote cost from total retention × ingest × replication-of-1 (remote stores one copy). Add hidden costs: API request charges (PUT/GET/LIST), egress on remote fetches, and the metadata topic.
showing 31–46 of 46