How does quota throttling show up in a Kafka request's latency breakdown, and how does the broker apply the throttle?
answer
- exceed quota → delay response → ThrottleTimeMs
- byte-rate + request-rate quotas, per user/client-id
- ClientQuotaManager computes delay over sliding window
- throttle echoed to client → self-mute
- policy, not overload — raise quota, don't add threads
basics
~10 sWhen a client exceeds its configured quota, the broker delays the request's response and records that delay as ThrottleTimeMs. The delay is also sent back to the client so it can slow itself down.
solid answer
~40 sKafka enforces quotas (produce/consume byte-rate and request-rate quotas, set per user/client-id) to protect brokers from noisy clients. When a client's measured rate exceeds its quota, the broker computes a delay and holds the response for that long; this delay is recorded as ThrottleTimeMs in RequestMetrics and is also returned in the response so the client's quota manager can mute its own sends/fetches. There are distinct throttle metrics: produce-throttle-time, fetch-throttle-time, and request-throttle-time, plus per-client throttle-time in the ClientQuotaManager. Importantly, the throttling is done by delaying the *response* (the request is processed, then the answer is parked), so it adds to TotalTimeMs without consuming extra handler work. This lets you distinguish 'broker is overloaded' (high queue/local time) from 'this client is being throttled by policy' (high ThrottleTimeMs), which are very different remediations.
go deeper
Know quotas limit a client's rate and the delay appears as ThrottleTimeMs.
Explain byte-rate vs request-rate quotas and that the broker delays the response and echoes the throttle to the client.
Separate throttle-induced latency from genuine broker overload and pick the right remediation.
Design per-tenant quota policy and request-rate quotas to protect multi-tenant clusters under noisy-neighbor load.
## Why quotas exist A single misbehaving client can saturate a broker's disk or network and starve everyone else. **Quotas** cap how much a client (identified by a user principal and/or client-id) may produce, consume, or how many requests it may issue. Types: - **Produce/Fetch byte-rate quotas** — cap bytes/sec produced or consumed. - **Request-rate quotas** — cap the fraction of broker request-handler + network thread time a client may use (a percentage, e.g. 200% = 2 full threads). Quotas are configured via dynamic config on `<user>`, `<client-id>`, or both, and stored in the cluster metadata. ## How the throttle is applied When a request would push a client over its quota, the broker's `ClientQuotaManager` computes a **delay** needed to bring the client's measured rate back under the limit (over a sliding window of `quota.window.size.seconds` × `quota.window.num` samples). Crucially, the broker **processes the request normally, then delays sending the response** by that amount. The held response is parked (similar mechanism to purgatory) and released when the delay expires. ## Where it appears - **ThrottleTimeMs** in RequestMetrics — the per-request delay added by throttling. - Aggregate sensors: `produce-throttle-time`, `fetch-throttle-time`, `request-throttle-time` (in `kafka.server:type=...ClientQuotaManager` / `Request`), giving avg/max throttle per client. - The throttle value is **echoed to the client** in the response header so the client can proactively mute its connection for that duration (avoiding even more aggressive broker-side muting). ## Why this matters for diagnosis ThrottleTimeMs is part of TotalTimeMs, so a throttled client *looks* slow end-to-end. But the cause is policy, not broker overload: queue times and local times can be perfectly healthy. The remediation is to raise the client's quota, fix the client's traffic pattern, or accept the throttle — not to add broker threads or disks. ## Edge cases - A client repeatedly over quota will be muted: the broker stops reading further requests from that connection until the throttle clears. - Request-rate quotas throttle based on thread *time* consumed, which can throttle even low-byte but high-request-count workloads (e.g. metadata storms). - Older clients that don't understand the returned throttle time still get correct behavior because the broker enforces the delay regardless.
- Why does the broker process the request but delay the response, rather than rejecting it?Delaying keeps the protocol simple and lossless — the client still gets a correct response, just later, and it learns the throttle duration so it can self-pace. Rejecting would force retries and complicate client logic. Delay smoothly shapes the rate down to the quota.
- If you see high ThrottleTimeMs but healthy RequestQueueTimeMs and LocalTimeMs, what's the right fix?The broker is fine; a client is hitting its quota by policy. Either raise that client's quota, fix its traffic pattern (batching, fewer requests), or accept the throttle. Adding broker threads or disks would not help.
saying these in an interview costs you the question
- Saying throttling rejects/drops the request — it delays the response.
- Treating high ThrottleTimeMs as a sign the broker is overloaded (it's a per-client policy delay).
- Forgetting request-rate quotas exist (assuming only byte-rate quotas).