How should a team decide where to set a throttling threshold relative to a service's SLOs, and what goes wrong in production if that threshold is set too aggressively versus too loosely?
answer
- knee of the latency curve
- safety margin below measured capacity
- too low = false rejections during calm infra
- too high = throttle triggers after SLO already breached
- revisit threshold as capacity changes
basics
~20 sSet the limit based on real measured capacity, with some safety margin, not a guess. Too strict and you block normal traffic for no reason; too loose and the throttle never actually kicks in before things break.
solid answer
~50 sThe threshold should be derived from load-testing or observed capacity data showing where the service's latency/error-rate SLOs start to degrade — the 'knee' of the curve — with a safety margin below that point to account for measurement noise, dependency variability, and the fact that the throttle itself takes time to react. Set too aggressively (too low), the throttle rejects traffic the system could actually have served fine, showing up as unnecessary errors and lost business during ordinary peaks. Set too loosely (too high), the throttle doesn't trigger until the service is already past its stable operating region, so it fails to prevent the very SLO violation it exists to protect against, and by the time it engages recovery is much harder because the system may already be in a cascading or metastable failure state.
go deeper
Should understand in plain terms that the limit shouldn't be a random guess, and that setting it wrong in either direction causes a different kind of problem.
Should describe using load testing or historical traffic data to find where SLOs start degrading, and mention leaving some safety margin.
Should describe concrete production symptoms of both too-aggressive and too-loose thresholds and how to diagnose which one you're seeing from telemetry.
Should connect a too-loose threshold to system-wide cascading/metastable failure risk, discuss per-endpoint variation in threshold, and treat the threshold as something continuously revalidated against changing capacity rather than a one-time setting.
## Start from where the SLOs actually break A throttling threshold is only meaningful in relation to the SLOs it's meant to protect, so setting it correctly starts with knowing, concretely, at what load level those SLOs start to break — the point sometimes called the **knee of the latency curve**, where a system that was tracking flat, predictable latency as load increased suddenly starts to see latency (and often error rate) climb steeply as some finite resource — a thread pool, a connection pool, a downstream database — approaches saturation. **Finding that knee isn't a guess**; it comes from load testing the service (or a close proxy of it) under controlled, increasing load and observing where the SLO metrics actually start to degrade, or from historical production telemetry correlating request volume against latency and error rate during real traffic peaks or past incidents. ## Then back off by a safety margin The threshold is then set some safety margin below that measured knee, not at it, because several things eat into the margin in practice: - measurement noise means the observed knee point in a test isn't perfectly reproducible in production - downstream dependencies can have their own variable capacity that shifts the effective knee over time - and the throttle mechanism itself takes some time to detect rising load and react, during which load can keep climbing past where the throttle started reacting. ## Set too aggressively Setting the threshold **too aggressively** — meaningfully below the real capacity — produces a specific, recognizable failure mode: the service starts rejecting requests during entirely ordinary traffic, well before it's under any genuine risk of breaching its SLOs. Because the rejections look, from a monitoring dashboard, like a self-inflicted outage rather than a capacity problem, they're often confusing to diagnose — error rates spike, but infrastructure metrics like CPU, memory, and downstream latency all look calm, because the system was never actually near its real limit. The business cost of this is real: legitimate customers see failures, retries generate support load, and trust in the service erodes, even though nothing about the underlying system was actually overloaded. This failure mode is especially common right after a throttle is first introduced, when the threshold is picked conservatively 'to be safe' without load-test data to back it up, and then never revisited as the service's actual capacity grows. ## Set too loosely Setting the threshold **too loosely** — meaningfully above the real capacity, or effectively disabled — produces the opposite and more dangerous failure mode: the throttle simply never engages before the service has already crossed its knee and started degrading. At that point the mechanism has failed at its one job, because the SLO breach the throttle exists to prevent has already happened by the time any protective action is taken. This is worse than it sounds in isolation, because a service that's past its knee is usually also generating secondary load — clients timing out and retrying, upstream services piling up their own queues waiting on this one — so the system can tip from 'slow' into a self-sustaining overloaded state (sometimes called a **metastable failure**) that doesn't recover on its own even once the original traffic surge passes, because the retry traffic alone is enough to keep it pinned past capacity. A throttle engaging late doesn't just fail to prevent the initial SLO breach, it can fail to prevent the system from getting stuck in a bad state that requires manual intervention (shedding load, restarting components, or waiting out client backoff) to recover from. ## A threshold is not a one-time setting Because both directions have real costs, a mature approach doesn't treat the threshold as a single number set once. It's revisited as the service's actual capacity changes: - after infrastructure scaling - after a slow dependency is optimized or made faster - after a new expensive endpoint is added And it's validated against production reality using tight observability: alerting on the throttle's own rejection rate as a first-class signal, and correlating rejections against whether downstream health metrics (CPU, connection pool saturation, dependency latency) actually indicate the system was near its limit at the time. - If rejections are climbing while downstream health metrics stay calm, that's a strong signal the threshold is set too low - If SLO violations happen with the throttle rejection rate still near zero, that's a strong signal the threshold is set too high or too slow to react. ## Where it shows up A concrete real-world instance of getting this wrong in the 'too aggressive' direction is a service that copies a rate limit from a similar-looking API without actually load-testing its own dependencies, then rejects a routine end-of-month traffic peak that its infrastructure could have handled comfortably — a self-inflicted incident with a root cause of 'the number was never validated,' not an actual capacity problem.
- How would you tell, from production telemetry alone, that a throttling threshold is set too low rather than genuinely protecting the system?Look at whether rejections correlate with actual downstream saturation signals like CPU, connection-pool usage, or dependency latency. If the throttle's rejection rate spikes while those infrastructure signals stay calm and well within normal range, the system was never actually at risk, which means the threshold is set below real capacity rather than at it.
- Why can a throttle that engages 'too late' make an incident worse rather than simply failing to help?By the time load has already crossed the knee, the system is likely generating secondary load from timing-out and retrying clients, which can push it into a self-sustaining overloaded state that doesn't recover just because the original spike ends. A throttle engaging at that point isn't just late to prevent the SLO breach, it's now fighting a bigger problem than the one it was designed to prevent in the first place.
- Should the same threshold apply to every endpoint on a service, or does it need to vary?It generally needs to vary, because different endpoints consume different amounts of the underlying capacity per request — a cheap read and an expensive report-generation call don't hit the knee at the same request rate. A single blanket threshold across all endpoints either under-protects the expensive ones or over-restricts the cheap ones.
Like setting a car's redline based on actual engine tests rather than a guess: set it too low and you're braking on a straight highway for no reason; set it too high and you find out where the engine actually fails only after it's already blown.
saying these in an interview costs you the question
- Picks a threshold number with no reference to load-test data or observed capacity
- Doesn't mention a safety margin between the measured knee and the enforced threshold
- Assumes a throttle threshold, once set, never needs revisiting as the service changes
- Can't explain the difference in production symptoms between too-aggressive and too-loose thresholds
- Doesn't connect a too-late throttle to the risk of a self-sustaining overloaded state, not just a delayed reaction