When would byte-rate quotas fail to protect a broker, and how do request-percentage quotas address that?
answer
- bytes ≠ CPU cost per request
- tiny frequent requests slip byte quotas
- request_percentage = % of thread time, 100% = 1 thread
- KIP-124, same window + throttle mechanism
- set byte + request quotas together
basics
~20 sByte-rate quotas only limit data volume, so a client sending many tiny but expensive requests can saturate broker CPU without moving many bytes. request_percentage caps the share of broker IO/network thread time a client uses, covering that gap.
solid answer
~50 sByte-rate quotas (producer_byte_rate / consumer_byte_rate) bound throughput in bytes, but broker cost isn't only bytes — it's also per-request CPU on the network and request-handler (IO) threads. A client doing tight-loop metadata requests, single-record produces, empty fetches, or constant offset commits can pin broker threads while moving almost no data, slipping under byte-rate limits entirely. The request_percentage quota (added in KIP-124) fixes this: it limits a client to a percentage of the total available IO+network thread time, where 100% equals one full thread. So request_percentage=200 on an 8-IO-thread broker means the client may use 2/8 of IO capacity before being throttled. The broker accumulates each request's thread time per client and throttles when the windowed percentage is exceeded, using the same windowing and throttle-delay mechanism as byte quotas. In a robust multi-tenant setup you set both byte-rate and request-percentage quotas so neither high-volume nor high-frequency abuse can starve other tenants.
go deeper
Know that there's a quota for request rate/CPU, not just bytes.
Explain that tiny frequent requests bypass byte quotas and request_percentage caps thread time.
Articulate the thread-pool cost model, that 100%=one thread, KIP-124 origin, and shared windowing; recommend setting both quota types.
Design layered protection across byte, request, and controller-mutation quotas, sizing request_percentage against num.io.threads and tenant SLAs.
## Why byte-rate alone is insufficient A broker's capacity isn't a single dimension. Two scarce resources matter: 1. **Network bandwidth / disk I/O** — bounded by **byte-rate quotas**. 2. **CPU on the broker's thread pools** — the **network threads** (read/write sockets) and the **request-handler / IO threads** (`num.io.threads`) that actually process requests. Every request — even a tiny one — costs thread time: parsing, authorization checks, metadata lookups, lock acquisition, response building. A client can therefore overwhelm a broker **without moving many bytes**: - A tight loop of **Metadata** requests. - **Single-record produce** calls (one record per request, thousands/sec). - **Empty or tiny fetches** polled aggressively. - Hyperactive **OffsetCommit** or **ListOffsets** calls. None of these trip `producer_byte_rate` or `consumer_byte_rate`, yet they can pin every IO thread and stall the cluster — a request-rate DoS. ## The request_percentage quota (KIP-124) Introduced in **KIP-124**, the **request quota** limits a client to a **percentage of broker thread time**. The unit: **100% = one full thread's worth of time** across the network + IO thread pools. So: - `request_percentage=100` → up to one thread fully busy. - `request_percentage=200` → up to two threads' worth. If a broker has, say, 8 IO threads + some network threads, the total available is several hundred percent; a single tenant capped at 200% can't monopolize it. ## How it's enforced The broker measures, per client entity, the **fraction of thread time** consumed, using the **same sliding window** (`quota.window.size.seconds`, `quota.window.num`) and the **same throttle-delay + cold-throttle** mechanism as byte quotas. When the windowed percentage exceeds the limit, it delays the response and reports `throttle_time_ms`. It's configured the same way: `kafka-configs.sh --alter --add-config 'request_percentage=200' --entity-type users --entity-name alice`. ## Defense in depth Mature multi-tenant clusters set **all three** quota dimensions: - `producer_byte_rate` and `consumer_byte_rate` → cap volume. - `request_percentage` → cap request-rate CPU. This closes both the high-volume and high-frequency attack surfaces. There's also a related but distinct lever, **controller mutation quotas** (KIP-599), which throttle the *rate of metadata mutations* (topic/partition create/delete) to protect the controller — a different control point from request_percentage. ## Edge cases - request_percentage can be **>100** intentionally; it's not bounded at 100. - A noisy client throttled on request_percentage but under its byte quota will still see non-zero `request-throttle-time` JMX metrics — diagnose by checking which throttle metric is elevated. - Because both quota types share the window, a client can be simultaneously throttled by bytes and by request percentage; the broker applies the larger required delay.
- On a broker, what does request_percentage=100 actually correspond to?The equivalent of one full thread's worth of network+IO request-handler time. The client may keep one thread fully busy before being throttled; 200 means two threads' worth, and so on.
- Give a concrete workload that defeats byte-rate quotas but is caught by request_percentage.A client in a tight loop issuing Metadata requests or single-record produces thousands of times per second: almost no bytes move, so byte-rate quotas don't trigger, but the per-request CPU pins IO threads — request_percentage throttles it.
saying these in an interview costs you the question
- Claiming byte-rate quotas alone fully protect a broker (they ignore per-request CPU).
- Saying request_percentage caps bytes — it caps thread time, a CPU-style metric.
- Assuming request_percentage must be ≤ 100 (values above 100 = multiple threads).
- Confusing request_percentage (KIP-124, request CPU) with controller mutation quotas (KIP-599, metadata-change rate).