In a throttling policy, what's the practical difference between a hard limit and a soft limit, and when would a team choose each?
answer
- ceiling regardless of load vs conditional ceiling
- credits/headroom signal
- quiet-period utilization waste
- stale signal → simultaneous bursts
- hard = business/security, soft = capacity
basics
~20 sA hard limit is a strict cap that's never crossed, even if there's spare capacity. A soft limit is a flexible cap that can be exceeded temporarily if the system has room, but tightens up when things get busy.
solid answer
~40 sA hard limit is an absolute ceiling enforced regardless of current system load — used for things like security boundaries, contractual quotas, or per-tenant cost caps, where crossing it is never acceptable even if spare capacity exists. A soft limit is a baseline that can be exceeded temporarily — bursting — when the system currently has headroom, and is only strictly enforced once overall load approaches the danger zone. Hard limits give predictability and simplicity at the cost of wasted capacity during quiet periods; soft limits improve utilization and are friendlier to bursty legitimate traffic, but need a live signal of 'how much headroom exists right now' to be safe, adding complexity and risk if that signal is stale or wrong.
go deeper
Should be able to state the basic distinction: one never bends, the other can flex for a while.
Should give at least one concrete reason to pick each — e.g., billing/security for hard, burst tolerance/utilization for soft.
Should identify what a soft limit depends on (a live headroom signal) and name the risk of that signal being stale or shared across simultaneous bursts.
Should discuss combining both in layered policy, and reason about system-wide failure when many callers burst against the same stale headroom estimate simultaneously — a capacity-planning-level concern, not just a per-caller one.
## The question a policy answers for every request A throttling policy needs to answer a basic question for every request: is this caller allowed to make this call right now? ## Hard limit versus soft limit A **hard limit** answers that question the same way regardless of what else is happening in the system: there is a fixed number (calls per second, concurrent connections, total quota per billing period) and once a caller reaches it, every further request is refused until the limit resets, full stop. It doesn't matter if the service is otherwise idle — the limit is a ceiling defined independently of current load. A **soft limit** answers the question conditionally: there's a baseline rate a caller is normally held to, but the system allows temporary excursions above that baseline — **bursting** — as long as the service currently has spare capacity to absorb it. As overall load rises and headroom shrinks, the soft limit tightens back toward (or below) the baseline, and bursting is no longer permitted. | Hard limit | Soft limit | |---|---| | a fixed number (calls per second, concurrent connections, total quota per billing period) | a baseline rate a caller is normally held to | | every further request is refused until the limit resets, full stop | temporary excursions above that baseline — bursting | | a ceiling defined independently of current load | as long as the service currently has spare capacity to absorb it | ## Why hard limits exist Hard limits exist primarily for cases where 'it depends on current load' is the wrong answer to give. - **Security and abuse boundaries** are one example: an API key limited to 1,000 calls per day for cost-control or contractual reasons should not suddenly get 5,000 calls just because the service happens to be quiet at 3 a.m. — the limit is about what the caller is entitled to, not about system health. - **Billing tiers** work the same way: a free-tier customer's quota is a business decision, not a capacity decision. Hard limits are also simpler to reason about and test, because their behavior doesn't depend on the state of the rest of the system at the moment of the call — a request either is or isn't within budget, full stop, which makes them predictable for both the operator and the caller. ## Why soft limits exist Soft limits exist to **reclaim the utilization** that pure hard limits waste. If every caller is held to a strict daily-average rate at all times, the system is provisioned for worst-case sustained load but sits mostly idle outside of peak periods, which is inefficient — plenty of real traffic is **naturally bursty** (a batch job that runs once an hour, a user who does several actions in quick succession then goes quiet). A soft limit lets that natural burstiness through when the system can absorb it, improving both the caller's experience (no unnecessary rejections during quiet periods) and the operator's effective capacity utilization, without permanently raising the baseline the system has to be provisioned for. ## The trade-off — how good is the headroom signal The trade-off is that a soft limit is only as safe as **the signal it uses** to decide 'is there headroom right now.' That signal might be: - current CPU/queue depth - a rolling average of recent traffic - or a token-style credit that accumulates during quiet periods and is spent during bursts If that signal lags reality — for instance, if it's computed on a delay, or if many callers burst at the same moment based on the same stale headroom estimate — the soft limit can let in more aggregate load than the system can actually handle, defeating its purpose. This is a real operational risk: a soft-limit scheme that looks safe under one caller's burst can fail badly when many callers burst simultaneously, because the 'spare capacity' each caller's burst decision was based on doesn't actually exist once you sum them all. ## Combining the two In practice, systems often combine both: a hard limit as the outer, non-negotiable boundary for cost/security/fairness reasons, and a soft limit inside it that manages how close to that boundary — or to overall system capacity — a caller is allowed to get at any given moment, giving the flexibility of bursting without removing the absolute guardrail. A cloud provider's compute service is a concrete real-world instance of this pattern: a virtual machine on a burstable instance type accrues CPU credits while idle and can burst above its baseline CPU allocation while credits last (soft limit tied to a headroom signal — the credit balance), but the account as a whole still has a hard cap on total number of instances or total spend that can never be exceeded regardless of credit balance. ## Failure modes - The failure mode to watch for **with hard limits** is rejecting perfectly legitimate, bursty-but-safe traffic simply because the limit wasn't designed to accommodate normal usage patterns, which shows up as support tickets and churn rather than a technical incident. - The failure mode **with soft limits** is the opposite: a headroom signal that's wrong or too slow to update lets bursts stack up simultaneously and pushes the system past its actual capacity anyway, which shows up as the very overload incident throttling was supposed to prevent — meaning a soft limit that isn't carefully bounded is really just a hard limit with extra steps and a false sense of safety.
- What signal would you use to decide how much 'headroom' exists for a soft limit, and what happens if that signal is stale?Common signals are current CPU or queue depth, a rolling short-window request rate, or an accumulated credit balance per caller. If the signal is stale — computed on a delay or shared across many callers who each burst based on the same snapshot — multiple bursts can land at once and exceed real capacity, because each caller's decision assumed headroom that others were consuming at the same moment.
- Would you use a hard limit or a soft limit for a per-tenant billing quota versus a per-second request-rate cap meant to protect the service from overload?A billing quota should almost always be a hard limit, since it reflects a contractual or cost entitlement that shouldn't flex with system load. A request-rate cap aimed at protecting the service is a better candidate for a soft limit, since its whole purpose is to track actual system health and allow more traffic when the system can safely take it.
- Can a soft limit ever be less safe than simply not throttling at all?In principle, if its headroom signal is badly wrong or gameable, a soft limit can create a false sense of protection that lets more simultaneous load through than an obviously strict, well-understood hard limit would — teams sometimes over-trust the 'smart' mechanism and skip other safeguards as a result. It's rarely literally worse than no throttling, but a poorly tuned soft limit can fail to protect the system at the exact moment protection matters most.
A hard limit is like a venue's fire-code maximum occupancy — never negotiable no matter how calm the crowd is. A soft limit is like a restaurant letting a party linger past their reserved slot when the next table hasn't shown up yet, but asking them to wrap up the moment the room fills.
saying these in an interview costs you the question
- Uses 'hard limit' and 'soft limit' interchangeably without a distinguishing definition
- Assumes soft limits are always strictly better because they 'let more traffic through'
- Doesn't mention that a soft limit needs a live signal of current headroom to be safe
- Applies a soft/bursty policy to a security or billing boundary that should never flex
- Doesn't recognize that many simultaneous bursts can defeat a soft limit's safety assumption