skip to content

Quotas, Throttling & Fairness

How much of a shared cluster one client takes: rate and size ceilings, connection caps, the batching that sets its cost, and what a node does anyway. Asked because a shared cluster fails as one.

part ofBroker & streaming operationsoverview, primer and where to startread it →
on this pageshow

questions

25

Why does a writer group many records into one batch before sending them to a broker node?

level: juniorimportance: must knowfreq 72%

answer

  1. the cost is per request, not per byte
  2. one round trip carries many records
  3. fixed overhead amortised over the group
  4. throughput bought with a little delay
  5. the writer sets it, the cluster pays for it

basics

~20 s

Because a large part of what a request costs a node is fixed, whatever the request carries: a round trip, a queue slot, a handler, per-request bookkeeping. A batch spreads that fixed cost over many records, so the same hardware accepts far more records per second.

solid answer

~50 s

A broker node's work is not spread evenly over bytes. Each request costs it a network round trip, a slot in its bounded request queue, a handler from a finite pool, a routing decision and a response — and that cost is nearly the same for one record as for a thousand. A **batch** is the group of records the writer puts into one request, so the fixed cost is paid once instead of once per record. Where a cluster is saturated by request count rather than by bytes, this is the single cheapest throughput lever available, and it is also what gives a compressor a useful unit to work on. The price is that a record may wait in the writer for company, so its own end-to-end delay rises, and that the writer holds unsent records in memory.

go deeper

for a junior

Recall that a request costs a node a fixed amount whatever it carries, so putting many records in one request raises throughput. Also recall the price: the record that starts a batch waits for the rest before anything is sent.

for a middle

Explain the fixed costs by name — round trip, queue slot, handler, accounting — and say why a cluster saturated by request count benefits far more than one saturated by bytes. Note that the return flattens as batches grow.

for a senior

Show that you can read the ratio of request rate to byte rate per caller and spot the writer sending tiny requests. Then talk about what changes when a node must unpack the group rather than keep it whole.

for a principal

Frame it as an economics question: grouping is the cheapest headroom on a shared cluster, it is configured by people who do not pay the cluster's costs, and a platform that wants it has to make efficient defaults the easy path for new writers.

## The cost a node pays per request, not per record A broker node's work does not scale smoothly with the bytes it receives. Every request that arrives carries an overhead that is roughly the same whether it holds one record or a thousand: - a **network round trip**, with its own latency, from the writer to the node and back; - a slot in the node's bounded **request queue**, the place accepted-but-unprocessed work waits; - a **handler** from a finite pool of threads that do the real request work; - a routing decision — which stream, and which part of it, these records belong to; - per-request accounting: metrics, authorization checks, and whatever the node records about the caller; - a response written back to the client. On a cluster carrying small records, that fixed overhead — not the byte volume — is usually what saturates first. Two writers can push identical data volumes and cost the cluster wildly different amounts, purely because one sends a thousand requests a second and the other sends ten. ## What a batch is, and what it buys A **batch** is simply the group of records a writer accumulates in memory and places into a single request. Nothing about the records changes: the same content, in the same order, with the same destination. Only the packaging changes. What that packaging buys: 1. **Amortised fixed cost.** One round trip, one queue slot, one handler, one accounting entry for many records. This is the whole of the throughput argument. 2. **Fewer stalls on network latency.** A writer that waits for an answer before sending again is limited by round-trip time; carrying more per round trip removes that ceiling without touching the network. 3. **A compressible unit.** The batch is what a writer compresses, which is a separate lever with its own trade — but it only exists because the records were grouped in the first place. 4. **Steadier load.** A node handling fewer, larger requests spends less time switching between pieces of work. ## What it costs - **Delay on the records that arrive first.** A record handed to the writer early sits and waits for company. Batching raises throughput by spending latency, and the record that started the batch pays the most of it. - **Memory on the writer's host.** Unsent records live in the sending application's memory until the batch goes. - **Coarser retry granularity.** A failed request takes its whole batch with it, so a retry re-sends every record's bytes, not just the one that had trouble. - **Coarser attribution.** A single slow request now represents many records, which makes a per-record timing story harder to read. ## Where platforms differ Grouping is universal; what happens to the group after it arrives is not. | Design | What the batch is once it arrives | |---|---| | A store that keeps the stream as the writer framed it | The group survives as a stored unit and is often handed to readers the same way, so the saving reaches storage and the read path too | | A broker that unpacks on arrival and tracks each message separately | The batch was a transport optimisation only; per-record work at the node still scales with the number of records | | A node obliged to inspect each record | The round trip is saved, the per-record work is not — this is where batching disappoints people | This is why a blanket claim like *bigger batches always help* is wrong: they always help the request count, and they help the rest only where the node treats the group as a unit. ## What an operator can see, and what they can do An operator can measure it. The ratio of request rate to byte rate, per caller, tells you the average number of records per request without needing anything from the writer's team: a caller with a high request rate and a modest byte rate is sending tiny requests, and is buying a disproportionate share of the cluster's fixed costs. What an operator cannot do is change it. Batch size, the willingness to wait for a batch to fill, and the compression choice are all set in the writer's own configuration. The cluster side of this conversation is evidence and a request, not a setting — which is why an interviewer who asks *why batch* is usually working towards *and what would you do about a writer that does not*.

  • If batching is that cheap, why not make every batch enormous?
    Because the return flattens. Once the batch is large enough that the fixed per-request cost is negligible against the bytes, further growth buys almost nothing, while the delay the first record waits, the memory the writer holds, and the amount re-sent on a failed request all keep rising. The lever has a knee, and past it you are paying latency for nothing.
  • A stream carries only a handful of records a second. What does batching buy there?
    Very little. There is nothing to group: batches stay small however long the writer is willing to wait, so the request count barely falls and the main effect is added delay on every record. Batching pays where record rate is high and records are small; on a low-rate stream the honest answer is that this lever is not the one to reach for.

A courier charges for the stop, not for the envelope. Sending fifty envelopes one at a time buys fifty stops; putting them in one sack buys one. The sack is cheaper per envelope — and the envelope that went in first waited the longest for the van to leave.

saying these in an interview costs you the question

  • Says batching makes each individual record arrive sooner.
  • Thinks a node does the same amount of work per record either way.
  • Believes a bigger batch is always better, with no cost at all.
  • Assumes an operator can set a writer's batch size from the cluster.
  • Confuses grouping records for transport with a delivery guarantee.
  • Thinks grouping records by itself reduces the bytes that are stored.
open as a page

A broker node cannot process writes as fast as they arrive, with every configured ceiling respected — what can it do?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A saturated node has four responses: queue the work, delay the answer, block the writer, or refuse the request. Each shows up differently at the call site, and the quietest discards records without failing the send.

open as a page

A broker enforces a byte-rate ceiling on one client identity - what exactly is capped, and what does that client see?

level: juniorimportance: must knowfreq 75%

basics

~20 s

A byte-rate ceiling caps how many bytes per second the cluster will handle for one named principal - a client identity, user, application or tenant - not per connection. Over it, that principal is either answered more slowly or refused.

open as a page

Why is a broker node's ceiling on concurrent connections a separate ceiling from its bytes-per-second allowance?

level: juniorimportance: must knowfreq 60%

basics

~20 s

A connection cap counts connections held open, not traffic. Holding one consumes memory and bookkeeping on the node whether or not data flows, so a fleet of near-idle clients can exhaust the cap while the byte rate stays tiny.

open as a page

A single record exceeds the maximum record size a broker node accepts — what happens to that write and to everything else?

level: juniorimportance: must knowfreq 68%

basics

~20 s

The node refuses that one write with a size error and keeps serving everything else on the same connection. Retrying the identical record never succeeds: it is a deterministic check, so the record must shrink or a ceiling must rise.

open as a page

A writer waits up to fifty milliseconds for a batch to fill before sending it anyway: what does that buy and cost?

level: middleimportance: must knowfreq 62%

basics

~20 s

It buys fuller batches, so fewer requests and less fixed per-request cost at the node. It costs up to fifty milliseconds of added delay on the record that opens each batch — and at a low record rate it is nearly all cost, because the batch never fills anyway.

open as a page

Every thread in a producing application is stuck inside a send call while the broker reports no errors — why?

level: middleimportance: must knowfreq 57%

basics

~20 s

The cluster is accepting records more slowly than the application produces them, so unsent records have filled the writer send buffer. With the buffer full the client blocks the calling thread instead of failing, and broker saturation becomes an application outage.

open as a page

A writer exceeds its rate ceiling and the broker enforces it - how do a throttle delay and a throttle refusal differ?

level: middleimportance: must knowfreq 60%

basics

~20 s

A throttle delay still serves the request but holds the answer back, so the symptom is latency; a throttle refusal rejects the call and the client must repeat it, so the symptom is errors. Platforms in this class choose opposite defaults.

open as a page

A shared cluster's overall numbers look merely busy while one tenant reports slow writes — how do you find which owner is responsible?

level: middleimportance: must knowfreq 60%

basics

~20 s

Break usage down per owner on the affected nodes, over the exact interval of the complaint, and on several axes at once — bytes, request rate, request-time share, connections. Cluster-wide totals cannot separate owners, and the tenant reporting the incident is rarely the cause.

open as a page

Maximum record size is set separately at the writer, the accepting node, the copy hop and the reader — which one decides?

level: middleimportance: must knowfreq 57%

basics

~20 s

The smallest ceiling on the path decides. A record must clear the writer's own limit, the accepting node's, the copy hop's where one exists, and the reader's; raising one and leaving the others alone moves the failure rather than removing it.

open as a page

On a broker cluster shared by several teams, why do one team's writes slow down during another team's nightly burst?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Shared nodes mean shared finite resources. A tenant does not get its own slice of the machine, it gets a turn, so one owner's burst consumes the network, disk, memory and request-handling threads that every other tenant on those nodes is waiting for.

open as a page

Who pays, and who benefits, when a writer compresses each batch with a stronger but slower compressor?

level: middleimportance: should knowfreq 55%

basics

~20 s

The sending application's own processor time pays, per batch, on its own hosts. The cluster and the network collect the benefit as fewer bytes moved and, where the node keeps the batch as framed, fewer bytes held. Readers pay a smaller decompression cost.

open as a page

A writer's byte-rate ceiling is enforced over an averaging interval - why does that interval decide whether a short burst survives?

level: middleimportance: should knowfreq 42%

basics

~20 s

Because the ceiling is judged against traffic accumulated over the interval, not against a single instant. A long interval forgives a short burst and then holds the client back afterwards; a short one clips the burst while it is happening.

open as a page

A service pools eight broker connections per instance and scales to fifty instances — what does the cluster see?

level: middleimportance: should knowfreq 55%

basics

~20 s

Roughly four hundred connections, because a per-instance setting is multiplied by the instance count. The pool size is a per-process number; the cluster only ever sees the product, plus whatever a rolling deploy adds while old and new instances overlap.

open as a page

Inside a broker node, what do the connection cap, the request queue depth and the handler pool each count?

level: middleimportance: should knowfreq 52%

basics

~20 s

Three different things in three different units: connections held open at once, accepted requests allowed to wait before being processed, and requests being worked on simultaneously. They sit in series, and none of the three is measured in bytes or in requests per second.

open as a page

A node's processor use climbs while its byte and request rates stay flat: what might it be doing to arriving batches?

level: seniorimportance: should knowfreq 45%

basics

~20 s

It may have stopped passing batches through untouched and started decoding and re-encoding them. Anything that forces a node to look inside a compressed batch — a conversion, a stamp, a per-record check — turns a free transfer into processor work that scales with records, not bytes.

open as a page

A writer discards records instead of failing the send when its unsent buffer fills under cluster saturation — what has that traded?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Availability of the calling path has been bought with data, and with the evidence that the data is gone. Sends keep succeeding, error rates stay flat, and the loss is detectable only by reconciling records offered against records the cluster accepted.

open as a page

One application's write latency tripled while cluster-wide health looks normal - how do you tell a rate ceiling from a struggling cluster?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Check whether the slowdown is confined to one principal. A node deliberately holding one identity's answers leaves every other client normal and the cluster's own saturation unremarkable, while a struggling cluster slows everybody at once.

open as a page

A node's held-connection count tracks request rate rather than instance count — what does that pattern indicate?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Clients are opening a connection per unit of work instead of reusing a pooled, long-lived one. A count that rises and falls with traffic, with short connection lifetimes, is connection churn; a count that only ever rises is a leak.

open as a page

After a bursting tenant is bounded on a shared cluster, why is the contention only capped and charged rather than removed?

level: seniorimportance: should knowfreq 46%

basics

~20 s

While the hardware stays shared, every available lever is a fairness lever, not an isolation one. Bounding an owner limits how much of the pool it can take and moves most of the cost onto it, but the resources are still shared, so interference is bounded rather than eliminated.

open as a page

What does raising a node's maximum record size cost the cluster, beyond letting one large record through?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A ceiling is a reservation, not a permission. The node must be able to hold a worst-case record for every request it is willing to handle at once, so memory headroom scales with the ceiling multiplied by requests in flight — paid by every client, for the benefit of the rare large record.

open as a page

One reader's own maximum record size is smaller than the node's — what happens to a record that lands between the two?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The record is accepted, stored and perfectly healthy from the cluster's point of view, but that one reader cannot take it. Depending on the design it stalls on that record or fails to complete it repeatedly, while every other reader is unaffected.

open as a page

Your shared cluster is at capacity and the cheapest headroom lies in how writers batch and compress: how do you proceed?

level: principalimportance: should knowfreq 40%

basics

~20 s

Treat it as a negotiation, not a setting. Measure records per request and compressibility per caller, price the change for the team that would make it, and be ready to buy capacity instead when the ask would break that writer's own latency commitments.

open as a page

Across many writing applications sharing one cluster, where should overload surface, and what do you require of every writer?

level: principalimportance: should knowfreq 33%

basics

~20 s

Overload should surface in the writers, as bounded waits and explicit failures, rather than in ever-deeper queues on the cluster. Require every writer to bound its unsent buffer, cap how long a send may wait, declare whether its records tolerate loss, and have a plan for records it cannot send.

open as a page

You set default rate ceilings for every application on a shared broker cluster - which unit do you attach them to, and how do you learn one is about to bind?

level: principalimportance: should knowfreq 32%

basics

~20 s

Set defaults in more than one unit - bytes and requests per second, plus a request-handling share where it exists - and review consumption against each ceiling routinely, so no team first meets its ceiling mid-incident.

open as a page