Why do services split rate limiting into a pre-auth hook keyed by client address and a post-auth hook keyed by principal?
answer
- the key decides the depth
- no principal exists before authentication
- protect the credential check itself
- quotas are per identity, not per address
- both inside the access log
basics
~20 sBefore authentication the only key available is the network address, and that tier exists to protect the credential-checking path itself. After authentication the limiter can key by principal and enforce the per-account quota the product actually promises.
solid answer
~50 sThe two tiers answer different questions, and their depth in the chain follows from what is known at that depth. Outside authentication, no identity has been established, so the only usable key is the network address (or a connection attribute). That tier is a blunt shield: it caps how fast anyone can hammer the credential check, which is what makes credential stuffing and unauthenticated floods expensive. Inside authentication, the principal is known, so the limiter can enforce the quota the product sells — per account, per tenant, per key — and give one heavy account no more than its share. One limiter cannot do both: place it before authentication and it cannot see identity; place it after and every rejected credential has already paid for a full verification. Both tiers sit inside the access-log hook, so their `429`s are still recorded.
go deeper
Know that a service can refuse a request for arriving too often, that the refusal is a 429, and that what the limit counts depends on whether the caller has been identified yet.
Explain why the key available before authentication is only the address, why the quota that the product sells needs a principal, and why one hook therefore cannot cover both limits.
Show operational judgment: generous coarse budgets because addresses are shared, a trusted-hop rule for forwarded addresses, both tiers inside the access log, and a deliberate fail-open or fail-closed decision.
Set the policy across services: who owns each tier, which limits belong at a shared edge and which need application identity, and how a tenant's contractual quota is stated so that it is enforced identically everywhere.
## Two limiters because there are two threats Rate limiting is usually discussed as one feature, but in a middleware chain it is naturally two hooks at two depths, because **the key a limiter can use is determined by what the chain has already established.** | Tier | Depth | Key | Protects | Typical budget | |---|---|---|---|---| | Coarse | outside authentication | client address or connection attribute | the credential check and everything cheap in front of it | generous, per address | | Fair-share | inside authentication | principal: account, tenant, API key | backend capacity and per-account fairness | tight, contractual | The coarse tier exists because verifying a credential is not free. Password verification is deliberately expensive, token verification involves signature checks and sometimes a lookup, and session validation usually touches a store. An attacker guessing credentials, or a broken client retrying in a tight loop, consumes exactly that work for every attempt. A limiter that only runs *after* authentication has, by definition, already paid for it — on every single rejected attempt. The fair-share tier exists because the limit the business actually promises is per identity: so many requests per minute per account, per tenant or per key. That limit cannot be stated before authentication, because nothing has established who is calling. ## Why one hook cannot cover both - **Placed before authentication**, a limiter has no principal. It can only count addresses, so it cannot state a quota, cannot isolate tenants, and cannot tell a paying customer's traffic from a scraper's. - **Placed after authentication**, a limiter never runs for a request that failed to authenticate — which is the traffic you most wanted to cap. Every failed attempt still costs a full credential verification. So the two are not redundancy; they are two different statements that happen to share a mechanism. ## Why address keys are weak, and what to do about it An address is a poor identity, and a candidate who says so has understood the tier: - many users share one address behind carrier-grade translation or a corporate egress, so a tight limit punishes everyone behind it; - one user's address changes as they move between networks, so a limit is easy to shed; - behind a proxy or load balancer, the immediate peer address is the proxy's; the client address comes from a forwarded header that is caller-controlled unless the hook is configured to trust only a known set of hops. The conclusions follow directly: keep the coarse tier **generous** — it is there to stop floods, not to be fair — and be explicit about where the address comes from. A limiter that reads a forwarded header without a trust rule can be bypassed by any client that sets it. ## Placement relative to the other hooks 1. **Inside the correlation-id hook**, so a rejection has an id. 2. **Inside the access-log and metrics hooks**, so every rejection is counted. A limiter placed outside them makes the one outcome you most need to see invisible: a client throttled all day looks like a client that went quiet. 3. **Inside the CORS hook**, so a browser preflight is answered rather than throttled into an opaque cross-origin failure. Counting preflights against a budget is a defensible choice; rejecting them is rarely what anyone intended. 4. **Outside the handler and outside any expensive parsing**, so a rejected request has consumed as little as possible. ## What the rejection should look like Both tiers answer with `429` and should include `Retry-After` so a well-behaved client knows when to come back rather than guessing. Keep the response body small and free of detail about the limit's internals. And keep the rejection cheap: a limiter that does expensive work in order to reject is a denial-of-service amplifier rather than a defence. ## Operating the pair - **Emit which tier rejected.** A `429` counter without a tier label cannot distinguish a flood from a customer exceeding their plan. - **Watch the coarse tier for collateral damage.** A sudden cluster of rejections from one address is often an office or a mobile carrier, not an attack. - **Roll out in observe-only first.** Count what would have been rejected before enforcing, on both tiers. - **Fail open or closed deliberately.** If the shared counter store is unavailable, the coarse tier usually fails open to avoid taking the service down with it, while a hard contractual quota may fail closed. Decide, write it down, and test it.
- Where does the pre-auth limiter get the client address when the service sits behind a proxy?From a forwarded header, but only if the hook is configured to trust a known set of hops and to take the address contributed by the last trusted one. A limiter that reads the header without a trust rule is keyed on caller-controlled input, so any client can rotate its key and shed the limit entirely.
- Should a rate-limit rejection appear in the access log?Yes, which is why both limiters sit inside the observation hooks. A throttled client that is never logged shows up only as reduced traffic, and that is indistinguishable from a client that stopped calling — the ambiguity that makes throttling incidents hard to diagnose.
- What breaks if the coarse tier is set as tightly as the per-account tier?Shared addresses collapse into one budget: an office, a school or a mobile carrier's translation pool has many legitimate users behind one key, and a tight limit refuses most of them. The coarse tier is a flood shield, so it should be generous enough that normal shared use never reaches it.
saying these in an interview costs you the question
- Puts the only limiter after authentication and calls floods handled
- Treats a network address as a reliable per-user identity
- Reads a forwarded client address with no trusted-hop rule
- Places the limiter outside the access log, hiding every rejection
- Rejects with a bare status and no indication of when to retry
- Runs an expensive check in order to reject a request cheaply