skip to content

What are Kafka client quotas, and what problem do they solve in a multi-tenant cluster?

level: juniorimportance: must knowfreq 62%

answer

  1. noisy-neighbor / multi-tenant
  2. produce + fetch byte-rate, request_percentage
  3. throttle delay, never drop data
  4. per user / client-id / (user,client-id)
  5. dynamic config, kafka-configs --alter

basics

~10 s

Quotas are per-client limits that Kafka brokers enforce to cap how much produce/fetch traffic (bytes per second) or broker request time a client can use, stopping one noisy client from starving others.

solid answer

~40 s

Kafka client quotas are broker-enforced rate limits applied per principal (user), per client-id, or per (user, client-id) pair. There are three quota types: produce byte-rate (producer_byte_rate), fetch byte-rate (consumer_byte_rate), and request percentage (request_percentage, the share of broker network/IO thread time). They solve the noisy-neighbor problem in shared, multi-tenant clusters: without quotas a single runaway producer or a misconfigured consumer doing tight-loop fetches could saturate broker network or CPU and effectively DoS every other tenant. When a client exceeds its quota, the broker doesn't drop data; it computes a throttle delay and delays the response, slowing the client smoothly. Quotas are stored as dynamic configs (originally in ZooKeeper, now in the KRaft metadata log) and can be changed at runtime with kafka-configs --alter without a restart.

go deeper

for a junior

Know the one-liner: per-client limits the broker enforces to stop one client hogging resources; it throttles, not drops.

for a middle

Name the three quota types and that they attach to user/client-id, applied via kafka-configs at runtime.

for a senior

Explain throttle-delay mechanics, per-broker (not cluster) enforcement, the windowing, and when request_percentage matters over byte-rate.

for a principal

Frame quotas as a tenancy control plane: budgeting across the entity hierarchy, capacity planning per broker, and combining byte + request quotas to bound worst-case noisy-neighbor impact.

## What a quota is A **Kafka cluster** is a set of servers called **brokers** that store streams of records in **topics**. **Producers** write records, **consumers** read them. In a shared or **multi-tenant** cluster (many teams/apps using one cluster), one badly behaved client can hog broker resources and degrade everyone else — the **noisy-neighbor** problem. A **quota** is a limit the broker enforces on how much of a resource a client may consume. ## The three quota types 1. **Produce byte-rate** (`producer_byte_rate`): max bytes/second a client may produce, measured at the broker. 2. **Fetch/consume byte-rate** (`consumer_byte_rate`): max bytes/second a client may fetch. 3. **Request percentage** (`request_percentage`): a CPU-style quota limiting the percentage of broker **request-handler (IO) and network thread** time a client may use. A value of `200` means the client may use the equivalent of 2 full threads' worth of time. This catches clients that send tiny but extremely frequent requests (high request rate, low bytes) which byte-rate quotas miss. ## How quotas are identified (the entity hierarchy) Quotas attach to entities: **user** (the authenticated principal), **client-id** (a label the client sets), or the pair **(user, client-id)**. There is also a cluster-wide **`<default>`** at each level. The broker resolves the most specific match first: `(user, client-id)` → `user` → `client-id` → defaults. This lets you set a generous user budget and sub-divide it among that user's client-ids. ## What happens on violation — throttling, not dropping Kafka **never drops data** to enforce a quota. Instead the broker measures the rate over a sliding window and, when a client exceeds its limit, computes a **throttle delay** (milliseconds) needed to bring the average back under the limit. For produce/fetch the broker **delays the response** by that amount (and mutes the channel) so the client naturally slows. The delay is returned in the response's `throttle_time_ms` field, and clients surface it as the JMX metric `produce-throttle-time-avg` / `fetch-throttle-time-avg`. This is graceful **back-pressure**, not failure. ## The measurement window Rates are measured over a sliding window controlled by `quota.window.size.seconds` (default 1s) and `quota.window.num` (default 11 samples), so bursts are smoothed over ~10 seconds rather than judged instantaneously. ## Managing quotas Quotas are **dynamic configs** changed at runtime (no restart) via the `kafka-configs.sh --alter --add-config` CLI against entities. They are stored in the cluster metadata (ZooKeeper historically; the KRaft metadata log in modern Kafka). ## Edge cases - A consumer fetching a single message **larger** than its byte-rate quota is still allowed through once, then heavily throttled afterward, so it can't deadlock on an oversized record. - Quotas are **per-broker**, not cluster-wide: a `producer_byte_rate` of 1 MB/s means 1 MB/s *per broker*, so effective cluster throughput scales with how many brokers the client's partitions live on.

  • When a producer exceeds its byte-rate quota, does the broker drop the excess records?
    No. The broker delays the produce response by a computed throttle time (and mutes the channel) so the client slows down; the data is still accepted once the client backs off. Quotas apply back-pressure, never data loss.
  • Why add request_percentage quotas if you already have byte-rate quotas?
    Byte-rate quotas don't catch clients that send a huge volume of tiny requests (e.g. metadata or 1-record produce calls) which burn broker thread CPU without moving many bytes. request_percentage caps the share of IO/network thread time and stops that class of abuse.

saying these in an interview costs you the question

  • Saying Kafka drops or rejects messages to enforce quotas (it throttles via response delay instead).
  • Claiming quotas are cluster-wide totals — they are enforced per broker.
  • Thinking quotas require a broker restart — they are dynamic configs applied at runtime.
  • Confusing byte-rate quotas with request_percentage; they catch different abuse patterns.

context