A writer exceeds its rate ceiling and the broker enforces it - how do a throttle delay and a throttle refusal differ?
answer
- one verb, two outcomes
- served late against rejected outright
- latency symptom against error symptom
- who paces next: node or client
- neither loses records at the broker
basics
~20 sA throttle delay still serves the request but holds the answer back, so the symptom is latency; a throttle refusal rejects the call and the client must repeat it, so the symptom is errors. Platforms in this class choose opposite defaults.
solid answer
~50 sEnforcement takes one of two shapes and they are not interchangeable. With a **throttle delay** the node accepts and serves the request, then withholds the reply for a computed period before answering, and normally tells the client how long it was held. Nothing fails and nothing is retried, so the cost lands as latency inside the calling application, which may be holding a thread for the whole period. With a **throttle refusal** the node rejects the call outright: the request did not happen, the client sees an explicit error, and whether offered traffic actually falls now depends entirely on the client's retry behaviour - a client that repeats instantly keeps the load exactly where it was. The operational consequence is the interesting part: a delay is quiet and easily mistaken for a sick cluster, while a refusal is loud but risks a self-inflicted retry storm.
go deeper
Remember that being over an allowance does not destroy anything: the call is either answered late or turned away, and a turned-away call is the caller's to send again.
Explain both shapes and the symptom each produces, and say who sets the pace afterwards - the node when it delays, the client's retry behaviour when it refuses.
Show that you check which shape your platform uses before configuring a ceiling, and that you know a delaying node looks identical to a failing one from the calling application's seat.
Weigh visibility against stability: refusal surfaces immediately but pushes pacing onto clients you do not control, while delay is stable but silent unless you invest in making the hold visible.
## One verb, two behaviours Throttling is a single word covering two incompatible outcomes. When a node decides a principal is over its rate ceiling, it either **holds the answer back and still serves the request**, or it **refuses the request and expects the client to come back**. Both are legitimate designs, both are in production across this class of platform, and an engineer who assumes one shape is universal will misread the symptom every time they meet the other. The first thing to establish about any cluster you operate is therefore which shape it uses - and whether it uses different shapes for different units, since a byte-rate allowance and a request-handling allowance are not obliged to behave the same way. ## What a throttle delay does - The request is processed normally; nothing about it is discarded or rejected. - The reply is withheld for a computed period, typically scaled to how far over the allowance the principal is. - The client is normally told how long it was held, so a well-instrumented client can report the hold rather than simply reporting slowness. - No call fails, so no retry logic engages and the offered load is paced by the node rather than by the caller. The consequences travel upward. Latency lands inside the calling application, and if the caller is synchronous a thread is parked for the entire hold. The team that owns the application reports a slow cluster; the operator looks at the cluster and finds it healthy. From the caller's seat, a node that is deliberately answering late is indistinguishable from a node that is failing - that is the single most valuable thing to know about this shape. ## What a throttle refusal does - The call is rejected and nothing is stored; the client receives an explicit error naming the condition. - Offered traffic falls only if the client actually pauses. A client that repeats immediately re-offers exactly the same load, and the node now does rejection work on top of everything else. - A client that gives up after a fixed number of attempts converts enforcement into work lost at the application level - which is where records genuinely go missing, not at the broker. Refusal has the virtue of being loud. It appears in the application's own error counts on the first occurrence, so nobody spends an afternoon looking at the wrong system. Its hazard is that it delegates pacing to the least reliable participant. ## Side by side | aspect | throttle delay | throttle refusal | |---|---|---| | client symptom | higher latency, no failures | explicit errors | | who sets the new pace | the node | the client's retry behaviour | | resource held during enforcement | the caller's thread and the pending answer | nothing on either side | | how it is usually misread | as a struggling or failing cluster | as data loss at the broker | | records lost at the broker | none | none | The last row is worth stating plainly because both shapes are routinely misdescribed as dropping data. A delayed write is accepted, just late. A refused write was never accepted, so it is still the caller's to repeat. Loss happens when a client stops repeating or when its own unsent memory runs out, which is a different mechanism entirely. ## Operating with each shape 1. Establish which shape your platform uses **before** you configure a ceiling, because the shape decides what the incident will look like when the ceiling first binds. 2. If it delays, make the reported hold time visible in the client's own telemetry. Otherwise the enforcement is invisible and every occurrence will be reported as a cluster problem. 3. If it refuses, confirm that clients actually slow down between attempts rather than repeating in lockstep, or the refusal achieves nothing except extra work. 4. Either way, make the ceiling discoverable by the team that owns the principal, so the first person to look already knows a ceiling exists. ## The mistake that costs most The expensive mistake is not choosing wrongly between the two shapes - you rarely get to choose. It is assuming the shape you know is the only one. An engineer who has only met refusal will insist that no errors means no throttling, and will go hunting through the cluster for a fault that is not there. An engineer who has only met delay will assume any error must be a cluster fault rather than a ceiling doing exactly what it was configured to do. Naming the fork, and asking which side of it you are on, is the whole skill here.
- Which of the two shapes is harder to detect, and why?The delay. It produces no failed calls, so nothing appears in the application's error counts, and a node deliberately answering late looks exactly like a node that is struggling. It is discoverable only if you already know a ceiling exists for that principal or if the client surfaces the hold time it was told about.
- Does either shape risk losing records?Not at the broker. A delayed write is accepted and served late, and a refused write was never accepted at all. Records are lost further out: a client that stops repeating a refused call, or one whose unsent memory fills while it waits. The enforcement itself discards nothing.
- Why can an instant retry after a refusal make things worse?Because refusal moves the pacing decision to the caller. If the caller repeats immediately, the offered load is unchanged and the node now spends effort rejecting as well as serving. The traffic only actually falls when clients wait between attempts, so enforcement by refusal is only as effective as the clients' willingness to back off.
saying these in an interview costs you the question
- Says a throttled request always comes back as an error.
- Assumes a delayed answer means records were dropped.
- Thinks a refusal means the broker accepted and then lost the write.
- Treats an instant retry of a refusal as harmless.
- Restarts broker nodes because one client's writes turned slow.
- Assumes the client is never told how long it was held.