You set default rate ceilings for every application on a shared broker cluster - which unit do you attach them to, and how do you learn one is about to bind?
answer
- one unit leaves a hole
- bytes, requests, request-handling share
- per application or per instance
- a catch-all default for the unconfigured
- review headroom before it binds
basics
~20 sSet defaults in more than one unit - bytes and requests per second, plus a request-handling share where it exists - and review consumption against each ceiling routinely, so no team first meets its ceiling mid-incident.
solid answer
~50 sA policy has four parts. The **unit**: bytes per second alone is incomplete, because a client making thousands of near-empty calls stays far under it while occupying a real share of the threads doing request work, so pair it with a request rate and, where the platform offers it, a share of request-handling capacity. The **principal**: decide whether the allowance follows one application or each of its instances separately, because if every instance carries its own identity the cluster experiences the figure multiplied by however many are running. The **default**: a catch-all that covers every principal nobody configured, or the policy only constrains clients you already knew about. And the **review**: consumption against ceiling, looked at routinely, so a team's growth is noticed before it becomes an incident. The ceiling is yours to change at any moment - the failure mode is not the limit, it is nobody knowing where they sit against it.
go deeper
The takeaway at this level is that allowances are configured deliberately by someone, including a default that catches every client nobody thought about individually.
Explain why each unit on its own has a blind spot, and be able to name which client each unit misses when it is used alone.
Do the arithmetic: state what the cluster permits if every principal sits at its ceiling at once, and say whether the identity is per application or per instance.
Own the whole shape - the default, the units, the granularity, the exception process and the review - and be able to say where a ceiling stops being the right lever and a capacity decision starts.
## What a default is actually for A default rate ceiling is not primarily a fairness device; it is a bound on how much damage any one unremarkable mistake can do. A misconfigured loop, a backfill nobody sized, a retry storm from an unrelated outage - each of these arrives as one principal suddenly consuming far more of a shared cluster than it did yesterday. A default that was set months earlier turns that from an estate-wide event into one application's problem. That reframing matters because it tells you where to set the default: not at what a well-behaved application needs today, but at a figure comfortably above that and comfortably below what would hurt everyone else. The gap between those two is the whole design space. ## Choosing the unit | unit | what it prices | what it misses on its own | |---|---|---| | bytes per second | volume moved in each direction | the client making many tiny calls | | requests per second | the cost of the call itself | one client moving enormous volume per call | | request-handling share | the fraction of a node's serving capacity used | nothing much, but not every platform offers it | The practical answer is that a single-unit policy has a hole in it, and which hole depends on which unit you picked. Bytes alone is the usual choice and leaves the high-frequency, low-volume client entirely unconstrained. Requests alone leaves a client shipping enormous payloads unconstrained. Where a platform exposes a share of request-handling capacity, that is the unit closest to what the node actually spends, and it is the one most often left unset because it is the least intuitive. ## Choosing the principal The granularity question is separate from the unit question and gets less attention than it deserves: - **One identity per application.** All its instances share one allowance. The figure you configure is close to what the cluster experiences, and growth in instance count does not quietly raise the application's share. - **One identity per instance.** Each instance gets the full figure. The cluster then experiences the configured number multiplied by however many instances are running, which means a routine scale-out silently raises what that application may consume - and nobody notices, because nothing was reconfigured. - **A shared allowance across several principals.** Where a platform supports it, several identities belonging to one team can be held against a single pooled figure. This is the option that matches how organisations actually think about ownership, and it is worth checking whether your platform has it before designing around its absence. Whichever you pick, be able to state the arithmetic: what the cluster as a whole will permit if every principal sits at its ceiling simultaneously. If that total is far above what the cluster can serve, the policy is decorative - which may be a deliberate choice, but should be a conscious one. ## Knowing before it binds The failure this policy exists to prevent is not the limit itself; it is a team meeting its limit for the first time during an incident. Three things prevent that: 1. **Published figures.** Every team can find its own allowance without asking anyone. A ceiling nobody can see is a trap rather than a control. 2. **Routine review of consumption against ceiling.** Not absolute consumption - consumption expressed as a proportion of the principal's own figure, measured over the span the allowance is evaluated over, so a spiky workload does not hide behind a comfortable average. 3. **Enforcement visible to the owner.** When a principal is held back, the team that owns it should be able to see that this happened to *them* rather than inferring it from a slow application. Because the ceilings are yours, raising one is a minutes-long change rather than a request to anybody. That is exactly why the review matters more than the figures do: the cost of being wrong is small, and the cost of nobody knowing is an afternoon of misdiagnosis. ## Exceptions, and where the lever stops Some principals legitimately need more, and the exception should carry three things: an owner, a reason, and a date to look at it again. Exceptions granted in an incident and never revisited are how a policy quietly becomes fiction. And there is a boundary worth naming out loud in an interview. A rate ceiling divides the capacity that exists; it does not create any. If the honest finding is that every principal sits near a ceiling that was set low because the cluster is too small, no allowance policy fixes that, and continuing to tune the figures is a way of avoiding the capacity conversation. Knowing where the lever runs out is as much of the answer as knowing how to pull it.
- Why is a byte-rate default alone an incomplete policy?Because it prices volume and nothing else. A principal issuing thousands of near-empty calls a second stays comfortably under a byte figure while occupying a real share of the threads that serve requests. Pairing it with a request rate, and with a share of request-handling capacity where the platform offers one, closes the gap.
- What should a granted exception carry with it?An owner, a reason and a review date. Without an owner nobody can be asked about it later; without a recorded reason the next operator cannot tell a deliberate decision from a forgotten one; without a date it survives long after the circumstance that justified it. Exceptions granted mid-incident are the ones most likely to lack all three.
- How do you tell when the allowance policy has stopped being the right lever?When most principals sit near ceilings that were set low because the cluster cannot serve more. Allowances divide existing capacity and never create any, so at that point tuning the figures moves the pain around rather than reducing it, and the honest next conversation is about capacity or about where that traffic belongs.
saying these in an interview costs you the question
- Sets one byte-rate default and calls the policy complete.
- Leaves the catch-all default unset so only known clients are bounded.
- Grants exceptions with no owner and no review date.
- Lets every team discover its ceiling by hitting it.
- Uses rate ceilings to paper over a cluster that is simply too small.
- Assumes a per-identity figure is what the cluster as a whole experiences.