skip to content

Operations and Administration

Running Kafka day to day: sizing clusters, managing configs, reassigning partitions, rolling upgrades, DR, and the admin tooling. Interviewers use this area to find out whether you have operated Kafka or only written clients against it.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

page 1 of 2

How do you create and inspect a Kafka topic from the command line, and what do the key options control?

level: juniorimportance: must knowfreq 80%

answer

  1. --bootstrap-server not --zookeeper
  2. --create / --describe / --list / --alter
  3. Leader / Replicas / Isr columns
  4. partitions = parallelism, RF = fault tolerance
  5. RF must be <= broker count

basics

~10 s

Use kafka-topics.sh with --bootstrap-server. To create: --create --topic name --partitions N --replication-factor R. To inspect: --describe --topic name, which shows partitions, leaders, replicas, and in-sync replicas (ISR).

solid answer

~40 s

The kafka-topics CLI is the primary tool for topic lifecycle. Creation: kafka-topics.sh --bootstrap-server host:9092 --create --topic orders --partitions 6 --replication-factor 3. Partitions set the unit of parallelism (one consumer per partition max within a group) and replication-factor sets fault tolerance (must be <= broker count). You can pass per-topic configs with --config, e.g. --config retention.ms=604800000 or cleanup.policy=compact. --describe --topic orders prints each partition's Leader (the broker serving reads/writes), Replicas (assigned broker IDs), and Isr (the in-sync replica set). A shrinking ISR or 'Leader: none' signals under-replication or offline partitions. --list shows all topics. In modern Kafka all of these talk to brokers via --bootstrap-server; the old --zookeeper flag is removed in KRaft clusters.

go deeper

for a junior

Know the create/describe/list commands and that --bootstrap-server points at a broker.

for a middle

Read --describe output: interpret Leader/Replicas/Isr and spot under-replicated partitions.

for a senior

Reason about partition sizing tradeoffs, RF vs min.insync.replicas, and config vs topic alter split.

for a principal

Set org-wide conventions for partition counts, RF policy, and incident runbooks using describe filters.

## What a topic is A Kafka **topic** is a named, append-only log of records. It is split into **partitions** — independent ordered logs that Kafka distributes across brokers. Partitions are why Kafka scales: producers write to many partitions in parallel, and within a **consumer group** at most one consumer reads each partition, so partition count caps a group's parallelism. Each partition is replicated for fault tolerance. **replication-factor = R** means R copies exist on R different brokers. One replica is the **leader** (handles all reads/writes); the rest are **followers** that fetch from the leader. The **ISR (in-sync replica set)** is the subset of replicas currently caught up to the leader. ## Creating a topic ``` kafka-topics.sh --bootstrap-server broker:9092 \ --create --topic orders \ --partitions 6 --replication-factor 3 \ --config retention.ms=604800000 \ --config cleanup.policy=delete ``` - `--partitions` — number of partitions. Increasing later is possible but **breaks key→partition ordering** for existing keys (the hash mapping changes), so size it up front. - `--replication-factor` — copies per partition; must be `<=` number of brokers, else creation fails. - `--config k=v` — per-topic overrides of broker defaults (retention, compaction, min.insync.replicas, etc.). ## Inspecting topics - `--list` — names of all topics. - `--describe --topic orders` — per-partition layout: ``` Topic: orders Partition: 0 Leader: 1 Replicas: 1,2,3 Isr: 1,2,3 ``` - **Leader** = broker serving that partition. `Leader: none` means no leader is available (partition offline). - **Replicas** = assigned broker IDs (preferred leader is listed first). - **Isr** = replicas in sync. If `Isr` is smaller than `Replicas`, the partition is **under-replicated** — a follower is lagging or its broker is down. - `--describe --under-replicated-partitions` / `--unavailable-partitions` filter to only problem partitions — the first thing to run during an incident. ## Altering and deleting - `--alter --topic orders --partitions 12` increases partition count (cannot decrease). - Topic *config* changes go through `kafka-configs.sh --alter --entity-type topics`, not `kafka-topics --alter` (which now only handles partition count). - `--delete --topic orders` marks it for deletion (requires `delete.topic.enable=true`, default true). ## Modern vs legacy All commands use `--bootstrap-server` to talk to brokers. The legacy `--zookeeper` flag is gone in KRaft-mode clusters; metadata now lives in the KRaft controller quorum, not ZooKeeper.

  • What does it mean when Isr is smaller than Replicas in --describe output?
    The partition is under-replicated: one or more follower replicas have fallen behind the leader (lagging beyond replica.lag.time.max.ms) or their broker is down. Durability is reduced; if min.insync.replicas can't be met, producers with acks=all start failing.
  • Can you reduce the partition count of an existing topic with --alter?
    No. Kafka only supports increasing partitions. Decreasing would require dropping data and reassigning keys, so it's not allowed — you must delete and recreate the topic (or create a new one and migrate).

saying these in an interview costs you the question

  • Saying you still need --zookeeper to manage topics (removed in KRaft).
  • Claiming you can decrease partitions with --alter.
  • Confusing Replicas (assigned) with Isr (currently in sync).
  • Thinking kafka-topics --alter changes topic configs like retention.ms (that's kafka-configs).

context

open as a page

What does Kafka's cleanup.policy control, and how do delete and compact differ? How do retention.ms and retention.bytes fit in?

level: juniorimportance: must knowfreq 70%

basics

~10 s

cleanup.policy sets how old data is removed. delete drops whole log segments once they exceed retention.ms (age) or retention.bytes (size). compact keeps the latest value per key forever. retention.ms/bytes only apply to delete.

open as a page

You start a brand-new Kafka broker with a fresh broker.id and join it to the cluster. Why does it not start serving traffic, and what must you do to put data on it?

level: juniorimportance: must knowfreq 70%

basics

~10 s

A new broker joins the cluster but Kafka never auto-moves existing partitions onto it. You must explicitly reassign replicas to the new broker using kafka-reassign-partitions to scale out.

open as a page

How do you estimate the raw disk storage a Kafka topic (or cluster) will consume given its throughput, retention, and replication settings?

level: juniorimportance: must knowfreq 75%

basics

~20 s

Storage = ingest rate (bytes/sec) x retention seconds x replication factor. So a 10 MB/s topic kept for 7 days at RF=3 needs roughly 10MB x 604800s x 3 = about 18 TB of disk across the cluster.

open as a page

How do you use kafka-configs.sh to alter a topic-level config, and what does the command look like?

level: juniorimportance: must knowfreq 65%

basics

~10 s

Use kafka-configs.sh with --bootstrap-server, --entity-type topics, --entity-name <topic>, and --alter --add-config key=value. To remove an override use --delete-config key. --describe shows current overrides.

open as a page

What is the difference between static (server.properties) configs and dynamic configs in Kafka, and why does it matter operationally?

level: juniorimportance: must knowfreq 70%

basics

~10 s

Static configs live in server.properties and only take effect after a broker restart. Dynamic configs are set with kafka-configs at runtime and apply without restarting the broker, so you can tune a live cluster.

open as a page

Why are under-replicated partitions (URP) and offline partitions the two most important things to alert on in a Kafka cluster, and how should the alert severities differ?

level: juniorimportance: must knowfreq 78%

basics

~10 s

Offline partitions mean some partitions have no leader, so reads/writes fail — page immediately (critical). Under-replicated partitions mean replicas are behind the leader, reducing durability but still serving traffic — warn, page if sustained.

open as a page

What is the kafka-reassign-partitions tool, and what are its three main phases?

level: juniorimportance: must knowfreq 70%

basics

~10 s

kafka-reassign-partitions is the CLI tool that moves partition replicas between brokers. Its three phases are generate (propose a plan), execute (apply it), and verify (check progress/completion).

open as a page

What is a rolling upgrade of a Kafka cluster, and why do you restart brokers one at a time instead of all at once?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A rolling upgrade restarts brokers one at a time, swapping in the new version on each, so the cluster keeps serving while every broker is replaced. One broker is down at a time; the rest keep handling reads and writes.

open as a page

How do you enable tiered storage in a Kafka cluster and turn it on for a specific topic?

level: juniorimportance: must knowfreq 55%

basics

~10 s

Turn it on cluster-wide with remote.log.storage.system.enable=true on each broker (plus configure a RemoteStorageManager plugin), then per topic set remote.storage.enable=true. After that, old segments are copied to the remote store.

open as a page

How do you diagnose consumer lag and group state using kafka-consumer-groups --describe?

level: middleimportance: must knowfreq 78%

basics

~10 s

Run kafka-consumer-groups.sh --bootstrap-server host --describe --group g. It shows per-partition CURRENT-OFFSET, LOG-END-OFFSET, and LAG (end minus current), plus which consumer/member owns each partition. High LAG means consumers are falling behind.

open as a page

Define RPO and RTO for a Kafka DR setup. What practical factors drive each, and what RPO does asynchronous replication imply?

level: middleimportance: must knowfreq 55%

basics

~20 s

RPO (Recovery Point Objective) = how much data you can lose, measured in time. RTO (Recovery Time Objective) = how long recovery takes. Async cross-cluster replication (MirrorMaker 2) means RPO > 0: anything not yet replicated when the primary dies is lost.

open as a page

Walk me through safely decommissioning a broker so you can take it out of the cluster permanently without data loss.

level: middleimportance: must knowfreq 65%

basics

~10 s

Move every replica off that broker.id to other brokers using kafka-reassign-partitions, wait until it hosts zero partitions, then shut it down and deregister it. Never just kill the process while it holds replicas.

open as a page

How would you design a consumer-lag alert that pages on-call only for real problems, and what makes raw lag a poor threshold?

level: middleimportance: must knowfreq 80%

basics

~20 s

Raw lag (offsets behind) is misleading because acceptable lag depends on throughput. Alert on time-to-drain (lag / consume rate) or sustained-and-growing lag, not a single absolute number, and require the condition to persist before paging.

open as a page

Explain the difference between local.retention.ms/bytes and retention.ms/bytes on a tiered-storage topic, and how to size them.

level: middleimportance: must knowfreq 60%

basics

~20 s

retention.ms/bytes is total retention across local + remote storage. local.retention.ms/bytes controls how long/how much data stays on local broker disk before it can be deleted (after upload). Local must be <= total, and local sizes your disk.

open as a page

How does kafka-consumer-groups --reset-offsets work, and what are the --to-earliest, --shift-by, and --to-datetime modes plus the dry-run/execute safety model?

level: seniorimportance: must knowfreq 70%

basics

~20 s

It rewrites a group's committed offsets so it reprocesses or skips records. Modes include --to-earliest (start of log), --to-latest, --shift-by N (relative), and --to-datetime (offsets at a time). The group must be inactive, and by default it's a dry run unless you add --execute.

open as a page

Explain how MirrorMaker 2 enables consumer failover across clusters, including offset translation and the role of MirrorCheckpointConnector.

level: seniorimportance: must knowfreq 45%

basics

~20 s

Source and target clusters have different offsets for the same record, so you can't reuse raw offsets after failover. MM2's MirrorCheckpointConnector tracks the mapping and writes checkpoints; consumers use RemoteClusterUtils (or automatic sync to __consumer_offsets) to resume at the equivalent position on the target.

open as a page

During broker turnover (adding/removing nodes), explain unclean.leader.election.enable and the durability tradeoff it represents.

level: seniorimportance: must knowfreq 60%

basics

~10 s

Unclean leader election lets Kafka elect an out-of-sync (lagging) replica as leader when no in-sync replica is available. It restores availability but can permanently lose committed messages. It defaults to false.

open as a page

What practical limits govern how many partitions a single broker (and the whole cluster) can host, and what breaks when you exceed them?

level: seniorimportance: must knowfreq 65%

basics

~10 s

Each partition replica is open files plus memory and adds controller/recovery work. A rough guide is up to ~1000-4000 partition-replicas per broker. Too many slows leader election, lengthens unclean-shutdown recovery, and increases end-to-end latency.

open as a page

How do replication throttles work during a partition reassignment, and which configs control them?

level: seniorimportance: must knowfreq 60%

basics

~10 s

Throttles cap how fast replicas copy data during a reassignment so it doesn't starve normal traffic. leader.replication.throttled.rate and follower.replication.throttled.rate (bytes/sec, per broker) set the limit; throttled-replicas lists which replicas are throttled.

open as a page

Explain the classic ZooKeeper-era two-phase Kafka upgrade using inter.broker.protocol.version and log.message.format.version. Why must the binary upgrade and the protocol bump be separate steps?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Phase 1: roll new binaries while pinning inter.broker.protocol.version (and log.message.format.version) to the OLD version, so new and old brokers still speak the old protocol. Phase 2: after all brokers are upgraded, roll again removing/raising the pins. Splitting it keeps a mixed-version cluster compatible and makes phase 1 downgradable.

open as a page

An ops team needs at least 14 days of message history retained for replay and DR backfill on a high-throughput topic. Which configs do you set and what trade-offs do you weigh?

level: middleimportance: should knowfreq 35%

basics

~20 s

Set retention.ms=1209600000 (14 days) and ensure retention.bytes is large enough (or -1) so size doesn't evict early. Keep cleanup.policy=delete. Watch disk: 14 days of high throughput can be huge, so plan capacity and consider tiered storage.

open as a page

What does controlled.shutdown.enable do during a broker restart, and what governs whether it succeeds?

level: middleimportance: should knowfreq 50%

basics

~10 s

On shutdown, the broker asks the controller to move its partition leaderships to other in-sync replicas before it exits, so leader failover is graceful instead of a sudden election. It's on by default.

open as a page

How do you choose the number of partitions for a topic based on target throughput and consumer parallelism?

level: middleimportance: should knowfreq 55%

basics

~20 s

Partitions = max(target / per-partition producer rate, target / per-partition consumer rate). Each partition is processed by at most one consumer in a group, so partition count caps consumer parallelism. Round up and leave growth headroom.

open as a page

How do you inspect the effective configuration of a topic or broker, and what is the difference between --describe and --describe --all?

level: middleimportance: should knowfreq 45%

basics

~10 s

Use kafka-configs --describe with --entity-type and --entity-name. By default it shows only explicitly-set overrides. Adding --all also lists inherited defaults and each value's source, giving the full effective config.

open as a page

When setting a dynamic broker config, what is the difference between --entity-name <id> and --entity-default, and when would you use each?

level: middleimportance: should knowfreq 40%

basics

~10 s

--entity-name <brokerId> sets a per-broker config that applies to just that broker. --entity-default sets a cluster-wide default applied to all brokers. Per-broker overrides the cluster default for that broker.

open as a page

What is leader imbalance, and how do preferred-leader election and auto.leader.rebalance.enable address it?

level: middleimportance: should knowfreq 55%

basics

~20 s

The first replica in a partition's replica list is its 'preferred leader'. After failures, leadership drifts off preferred replicas, concentrating load on a few brokers. auto.leader.rebalance.enable=true periodically moves leadership back to preferred replicas when imbalance exceeds a threshold.

open as a page

How does Kafka's bidirectional client/broker compatibility work, and what does it mean for upgrading clients vs. brokers?

level: middleimportance: should knowfreq 55%

basics

~20 s

Each Kafka API (request type) is versioned. On connect, the client asks the broker which versions it supports (ApiVersions) and uses the highest both understand. Since ~0.10.2 this works both ways: newer clients talk to older brokers and older clients talk to newer brokers. Upgrade brokers and clients independently.

open as a page

What operational safeguards (controlled shutdown, health gating, acks/min.insync.replicas) make a rolling restart safe, and how do you pace it?

level: middleimportance: should knowfreq 50%

basics

~20 s

Enable controlled shutdown so a broker moves leadership off itself before stopping. Between brokers, wait until UnderReplicatedPartitions is 0 and there are no offline partitions. Use replication factor 3 with min.insync.replicas 2 and acks=all so producers survive one broker being down.

open as a page

How do you manage topics, configs, and ACLs programmatically with the Java AdminClient, and how do its async results behave?

level: seniorimportance: should knowfreq 55%

basics

~10 s

AdminClient (Admin.create) is the Java API behind the CLI tools. You call methods like createTopics, describeTopics, alterConfigs, createAcls, listConsumerGroups. Each returns a *Result holding KafkaFutures you complete to get values or catch exceptions.

open as a page

showing 1–30 of 46