skip to content

What types of client quotas does Kafka support and how do byte-rate versus request-rate quotas differ?

level: seniorimportance: should knowfreq 50%

answer

  1. 3 types: producer_byte_rate, consumer_byte_rate, request_percentage
  2. byte=network/disk, request%=CPU/thread time
  3. entities: client-id, user, (user,client-id)
  4. throttle by delaying response, sliding window
  5. per broker, not cluster-wide

basics

~10 s

Kafka has two quota kinds: byte-rate quotas (producer_byte_rate, consumer_byte_rate) that cap MB/s per client, and request-rate quotas (request_percentage) that cap the share of broker request-handler/network thread time. They apply per client-id, user, or user+client-id.

solid answer

~50 s

Kafka enforces quotas per broker, scoped to a quota entity: client-id, user (KafkaPrincipal), or the (user, client-id) pair, with a fallback default hierarchy. There are three quota types: producer_byte_rate caps incoming produce bytes/sec, consumer_byte_rate caps outgoing fetch bytes/sec, and request_percentage caps the percentage of time a client may occupy the broker's request-handler (I/O) and network threads — protecting against clients that are cheap in bytes but expensive in CPU (tiny frequent requests, metadata storms). Byte-rate quotas defend disk/network capacity; request-rate quotas defend CPU. Enforcement is by throttling: the broker computes the delay needed to bring the measured rate back under quota over a sliding window (quota.window.size.seconds × num.quota.samples) and delays the response by that much, surfacing produce/fetch-throttle-time-ms metrics. From this leaf's capacity lens, these metrics tell you when clients are being throttled and how much headroom remains.

go deeper

for a junior

Know quotas exist and that byte-rate limits MB/s per client.

for a middle

Name the three types and that quotas apply per client-id/user and throttle rather than reject.

for a senior

Explain byte (network/disk) vs request_percentage (CPU) and the sliding-window throttle mechanics.

for a principal

Design a quota strategy across users/clients, account for per-broker scaling, and tie throttle-time metrics into capacity SLOs.

**Why quotas exist (capacity view).** On a shared cluster, one misbehaving or runaway client can saturate disk, network, or CPU and starve everyone else. Quotas cap per-client resource use so the cluster's capacity is shared predictably. (Configuring them as a security/control lever is a separate concern; here we care about reading the throughput and throttle signals.) **Quota entities (who a quota applies to).** A quota is attached to one of: - **client-id** — the `client.id` a client sets (weak isolation, client-controlled). - **user** — the authenticated `KafkaPrincipal` (strong, server-controlled). - **(user, client-id)** — most specific. Kafka resolves the most specific match, falling back through a defined hierarchy down to a cluster default. Quotas are stored in the metadata (KRaft) / formerly ZooKeeper and managed with `kafka-configs.sh --alter --add-config 'producer_byte_rate=...' --entity-type clients|users`. **The three quota types:** 1. **producer_byte_rate** — max bytes/sec a client may *produce* into a broker. Defends inbound network + disk write capacity. Units: bytes/sec, per broker. 2. **consumer_byte_rate** — max bytes/sec a client may *fetch* from a broker. Defends outbound network. Per broker. 3. **request_percentage** (request-rate quota, KIP-124) — the percentage of broker thread time the client may use across the network threads and I/O (request-handler) threads. The ceiling is `n × 100%` where n = `num.io.threads + num.network.threads` (so a value of 200 ≈ two full threads). This protects **CPU**, the dimension byte quotas miss: a client sending huge numbers of tiny requests, doing aggressive metadata refresh, or causing expensive fetches consumes thread time disproportionate to its bytes. **Byte-rate vs request-rate — the key difference.** Byte quotas bound *data volume* (network/disk). Request quotas bound *processing time* (CPU). A client can be well under its byte quota yet still hammer the broker with request volume; only request_percentage catches that. Conversely a few large requests are cheap on threads but heavy on bytes. Real protection usually needs both. **Enforcement mechanics (throttling, not rejection).** Kafka does not drop requests; it *delays* responses. It measures each client's rate over a sliding window defined by `quota.window.size.seconds` (default 1s) and `num.quota.samples` (default 11) — i.e. ~11 one-second samples. When the measured rate exceeds the quota, the broker calculates the delay D that, applied to the current request, brings the average back to the quota, and holds the response for D. Modern clients honor the returned throttle time and back off; the broker also mutes the channel so the slow-down is cooperative. Observable via `produce-throttle-time-avg/max`, `fetch-throttle-time-avg/max` (client side) and per-broker throttle metrics. **Edge cases / gotchas:** - Quotas are **per broker**, not cluster-wide: a `producer_byte_rate=10MB/s` lets a client do 10 MB/s *to each broker it leads partitions on*, so effective cluster throughput scales with broker fan-out. - A too-small window or sample count makes throttling bursty/jittery. - Throttling can masquerade as latency or even client timeouts if quotas are aggressive and clients are old; rising throttle-time metrics are the tell. - request_percentage above 100 is normal — it's a percentage of *total* thread capacity (sum of threads), not a single core.

  • Why isn't a byte-rate quota enough on its own?
    Byte quotas bound data volume but not CPU. A client sending many tiny requests or doing heavy metadata/fetch work can stay under its byte quota while exhausting broker request-handler threads; request_percentage caps that thread time.
  • Does Kafka reject requests that exceed a quota?
    No. It throttles by computing a delay over the sliding window and holding the response that long, returning throttle-time so compliant clients back off. Nothing is dropped or rejected.
  • Is a producer_byte_rate quota a cluster-wide or per-broker limit?
    Per broker. The client gets that rate against each broker, so its total possible cluster throughput is the quota times the number of brokers hosting its partition leaders.

saying these in an interview costs you the question

  • Saying quotas reject/drop requests rather than throttle by delay.
  • Claiming byte-rate quotas protect CPU (they protect bytes; request_percentage protects CPU).
  • Treating a byte-rate quota as cluster-wide rather than per-broker.
  • Thinking request_percentage cannot exceed 100% (it's a share of summed thread capacity).

context