skip to content

What are controller mutation quotas (KIP-599), what do they protect, and how do they differ from request-percentage quotas?

level: principalimportance: nice to knowfreq 22%

answer

  1. KIP-599, controller_mutation_rate
  2. cost = partitions created/deleted
  3. protects controller / metadata log
  4. returns THROTTLING_QUOTA_EXCEEDED + retry
  5. vs request_percentage = data-path thread time, silent delay

basics

~20 s

Controller mutation quotas limit how fast a user/client-id can create or delete topics and partitions, protecting the controller from a flood of metadata changes. They throttle metadata-mutation rate, whereas request_percentage throttles general request CPU on data-path brokers.

solid answer

~50 s

Controller mutation quotas, added in KIP-599, cap the rate of partition mutations a principal/client-id can drive through CreateTopics, DeleteTopics, and CreatePartitions — measured in partitions mutated per second via the controller_mutation_rate quota. The cost of each request is the number of partitions it creates or deletes, accumulated over the usual sliding window. They protect the single controller (and the metadata log) from being overwhelmed by mass topic churn — e.g. a test harness or buggy automation creating thousands of topics — which could stall metadata propagation for the whole cluster. The throttling differs from data-path quotas: instead of silently delaying, the broker returns a THROTTLING_QUOTA_EXCEEDED error along with throttle_time_ms, and well-behaved admin clients retry after the delay (KIP-599 introduced this 'returned error + retry' model so admin operations fail fast and clearly rather than hanging). request_percentage, by contrast, governs IO/network thread time on ordinary produce/fetch/metadata traffic and uses cold-throttling (delayed response), not an explicit error.

go deeper

for a junior

Awareness only: there's a quota that limits how fast you can create/delete topics.

for a middle

Know it protects the controller and is measured in partition mutations per second via controller_mutation_rate.

for a senior

Explain KIP-599's error-plus-retry model (THROTTLING_QUOTA_EXCEEDED) and how it differs from cold-throttling data quotas.

for a principal

Position it within a layered tenancy control plane, set sane defaults for CI/automation, and reason about controller capacity and metadata-log protection under mass churn.

## What the controller is and why it needs protecting In a Kafka cluster, the **controller** is the special broker (in KRaft, a quorum of controller nodes) responsible for **cluster metadata**: which topics/partitions exist, leadership, and replica assignment. Every **CreateTopics**, **DeleteTopics**, and **CreatePartitions** request flows to the controller and mutates the **metadata log**. The controller is comparatively a **single point of serialization** for these changes; flooding it with mass topic creation/deletion can back up metadata propagation and degrade the whole cluster — even if data-path produce/fetch is fine. ## The controller mutation quota (KIP-599) **KIP-599** introduced the **`controller_mutation_rate`** quota. Key points: - **Unit:** partitions mutated per second. The **cost** of a request = number of partitions it creates or deletes. Creating a 50-partition topic costs 50; deleting it costs 50. - **Entities:** like other quotas, it attaches to **user**, **client-id**, or **(user, client-id)**, with defaults, set via `kafka-configs.sh --alter --add-config 'controller_mutation_rate=...'`. - **Window:** uses the standard `quota.window.*` sliding window for rate measurement. ## The crucial behavioral difference: error vs. silent delay Data-path quotas (byte-rate, request_percentage) use **cold throttling** — the broker silently **delays the response** and the operation still succeeds. That's wrong for admin operations, where a client wants a clear, fast answer. So KIP-599 uses a different model: - When the mutation quota is exceeded, the controller **rejects** the request with **`THROTTLING_QUOTA_EXCEEDED`** and includes a **`throttle_time_ms`** telling the client how long to wait. - The **AdminClient** is expected to **retry** after that delay (configurable via `retries` / `retry.backoff.ms`), so large batch operations are spread out instead of hammering the controller. - This makes the throttling **observable and recoverable**: the operation fails fast with a known error code rather than mysteriously hanging. ## Contrast with request_percentage | | controller_mutation_rate (KIP-599) | request_percentage (KIP-124) | |---|---|---| | Protects | the controller / metadata log | broker IO + network threads (data path) | | Measures | partitions created/deleted per second | % of thread time (100% = one thread) | | Triggers on | CreateTopics/DeleteTopics/CreatePartitions | any request type's CPU cost | | On exceed | returns THROTTLING_QUOTA_EXCEEDED, client retries | cold-throttle: delays response silently | They are complementary: request_percentage stops a noisy *data* client; controller_mutation_rate stops a noisy *admin/automation* client from churning metadata. ## Edge cases / design notes - A single huge CreateTopics with thousands of partitions can exceed the quota outright; the AdminClient retries with backoff, so the create eventually completes spread over time. - The quota guards against accidental DoS from CI pipelines, ephemeral-topic patterns, and buggy operators, not just malicious actors. - Because it's a rate over the window, transient bursts (a normal deploy creating a handful of topics) usually pass; only sustained mass churn is throttled. - It is enforced at the controller regardless of which broker received the admin request, since all such mutations are funneled to the controller.

  • Why does the controller mutation quota return an error instead of silently delaying like byte-rate quotas?
    Admin operations need a clear, fast result, not a hidden hang. KIP-599 returns THROTTLING_QUOTA_EXCEEDED with throttle_time_ms so the AdminClient can retry after backoff, spreading large topic operations over time while keeping the API responsive and observable.
  • How is the 'cost' of a CreateTopics request measured for this quota?
    By the number of partitions it creates (or deletes). Creating a topic with 50 partitions costs 50 against the controller_mutation_rate; the quota is expressed in partition-mutations per second over the sliding window.
  • Could request_percentage alone protect the controller from topic-churn DoS?
    Not well. request_percentage caps general request thread time on data-path brokers; it isn't tuned to metadata-mutation cost or enforced at the controller. controller_mutation_rate specifically prices partition mutations and protects the metadata log.

saying these in an interview costs you the question

  • Saying controller mutation quotas silently delay like byte quotas — they return THROTTLING_QUOTA_EXCEEDED and the client retries.
  • Measuring the quota in bytes or requests rather than partitions mutated per second.
  • Conflating it with request_percentage (different KIP, different target, different enforcement behavior).
  • Claiming it protects produce/fetch throughput — it protects the controller/metadata path.

context