skip to content

Explain how the broker measures rate and computes the throttle delay, including the role of quota.window.size.seconds and quota.window.num.

level: seniorimportance: should knowfreq 40%

answer

  1. 11 samples × 1s ≈ 10s window
  2. delay = (observed−Q)/Q × windowMs
  3. cold-throttle: mute channel, hold response
  4. throttle_time_ms in response + JMX
  5. window configs static; quota values dynamic

basics

~20 s

The broker tracks each client's rate over a sliding window of quota.window.num samples, each quota.window.size.seconds long. When the average exceeds the quota, it computes a delay = how long the client must pause to bring the windowed average back under the limit, then delays the response by that much.

solid answer

~50 s

Kafka measures client rates with a sliding-window Rate sensor: `quota.window.num` samples (default 11), each `quota.window.size.seconds` wide (default 1s), giving roughly a 10-second observation window. When a request would push the windowed average rate above the quota Q, the broker computes a throttle delay so that, over the window, the effective rate equals Q: delay ≈ (observed − Q)/Q × windowDuration. The broker caps the delay at the full window size, returns it in `throttle_time_ms`, and mutes the client's channel for that period before sending the response (cold-throttling for produce/fetch). Because the average is windowed, a short burst is tolerated and amortized across the window rather than triggering an instant hard stop. Smaller windows react faster but allow less bursting; larger windows smooth more but let bigger bursts through. Both are broker-side configs requiring a restart to change, unlike the quota values themselves.

go deeper

for a junior

Know that throttling is measured over a short window, not instantly.

for a middle

Recall the ~10s window (11×1s) and that exceeding it delays the response.

for a senior

Explain the delay formula, cold-throttle channel muting, throttle_time_ms/JMX, and the static-vs-dynamic config split.

for a principal

Reason about tuning the window for burst tolerance vs. reaction speed and the worst-case noisy-neighbor duration that follows from window size.

## The measurement model Kafka doesn't measure an instantaneous rate; it uses a **sampled sliding window**. Two broker configs define it: - **`quota.window.size.seconds`** (default **1**): the duration of one sample bucket. - **`quota.window.num`** (default **11**): how many consecutive sample buckets are retained. Together they form an observation window of about `(num − 1) × size` ≈ **10 seconds**. The broker keeps a `Rate` metric per client entity: each incoming byte/request adds to the current bucket; old buckets age out as time advances. The measured rate is total recorded value divided by the elapsed window time. ## Computing the throttle delay When a new request arrives, the broker records its bytes, then asks the sensor whether the windowed average exceeds the quota `Q`. If it does, it computes the delay needed to bring the average back to `Q`: ``` delayMs ≈ (observedRate − Q) / Q × windowDurationMs ``` Intuitively: "the client has produced X bytes too many; at rate Q those extra bytes would take this long, so pause that long." The delay is **bounded** by the window size so a single huge spike can't cause an unbounded stall. The value is returned to the client in the response field **`throttle_time_ms`**. ## How the delay is applied (cold throttling) For produce and fetch quotas, the broker **completes the operation but holds the response**: it mutes (stops reading from) the client's network channel for `throttle_time_ms`, then sends the response. This is **cold throttling** — the client is paused at the channel level, which naturally back-pressures its send/receive loop. The client also exposes throttle time via JMX (`produce-throttle-time-avg`, `fetch-throttle-time-max`), so you can detect throttling without broker access. Request-percentage quotas throttle similarly based on accumulated thread time. ## Why a window instead of an instant limit Real traffic is bursty. An instantaneous limit would reject normal micro-bursts. The window **amortizes** bursts: a client can briefly exceed `Q` as long as the windowed average stays at or below `Q`. After a burst, subsequent requests get throttled until the average recovers. ## Tuning trade-offs - **Smaller `quota.window.size.seconds` / fewer samples** → faster reaction, tighter conformance, but less tolerance for legitimate bursts and more throttle churn. - **Larger window** → smoother, more burst-tolerant, but slower to clamp a runaway client, so worst-case noisy-neighbor impact lasts longer. ## Edge cases - A single record/fetch **larger than the quota** is allowed through once (otherwise it could never be delivered), then the client is throttled hard on the next requests. - These window configs are **static broker configs** (set in `server.properties`, restart to change), whereas the **quota values** are dynamic. People often conflate the two. - Throttle delay is capped at the window duration, so the maximum single-response delay is roughly `quota.window.size.seconds × quota.window.num` is NOT the cap — the cap is the window *size* used by the rate's delay calc; in practice delays are capped to one window length.

  • What happens if a consumer's single fetch returns a message bigger than its consumer_byte_rate quota?
    The oversized fetch is allowed through once so the client isn't permanently stuck, then the broker throttles subsequent fetches heavily to bring the windowed average back under the quota. Quotas never make a deliverable message undeliverable.
  • How can a client tell it is being throttled without access to the brokers?
    The throttle duration is returned in every response's throttle_time_ms field, and clients surface it as JMX metrics like produce-throttle-time-avg / fetch-throttle-time-max. Non-zero values mean the client is hitting its quota.
  • If you shrink quota.window.size.seconds, how does behavior change?
    Throttling reacts faster and conforms more tightly to the quota, but tolerates fewer/smaller legitimate bursts, causing more frequent throttling on spiky workloads.

saying these in an interview costs you the question

  • Saying the broker measures an instantaneous rate (it uses a sampled sliding window).
  • Claiming quota.window.* are dynamic like quota values — they are static broker configs.
  • Thinking an oversized message is rejected/dropped — it's passed once then throttled.
  • Believing throttling drops the connection rather than muting/delaying it (cold throttle).

context