skip to content

Your secret store rate-limits issuance to protect its own capacity — why does that not bound one compromised identity?

level: seniorimportance: must knowfreq 46%

answer

  1. two limits, two different questions
  2. the service was never in trouble
  3. 4% of the ceiling is invisible
  4. flow and stock, not one of them
  5. size it from need, not from headroom

basics

~20 s

A capacity limit is shared across the fleet, so one identity can consume a small slice of it and never trip anything. Bounding a single identity needs per-identity limits: how fast it may mint, and how many of its credentials may be live at once.

solid answer

~50 s

The two limits answer different questions. A capacity limit keeps the store serving under load; it is sized for the whole estate, so 200 issuance calls a minute from one host inside a fleet of 800 is a few per cent of the ceiling and nothing degrades. A per-identity cap bounds blast radius while the store is perfectly healthy, and it needs two axes: a **flow cap** on mints per interval, and a **stock cap** on how many of that identity's credentials may be live simultaneously. Flow alone lets a patient attacker accumulate at a legal rate; stock alone lets a burst through when lifetimes are short. Derive both from what the identity legitimately needs — a collector holding one credential and renewing hourly is fine with a flow of 10 an hour, which turns a 4,000-credential incident into at most 10.

code

pseudocode · 20 lines
pseudocode
on issueRequest(parentIdentity, target):
    window = now - 1 hour
    minted = countIssued(parentIdentity, since = window)
    live   = countLive(parentIdentity)          # not yet expired, not yet withdrawn

    if minted >= parentIdentity.issuanceRatePerHour:
        emit("issuance flow cap hit", parentIdentity, minted, window)
        deny()

    if live >= parentIdentity.maxLiveCredentials:
        emit("issuance stock cap hit", parentIdentity, live)
        deny()

    credential = mint(target, lifetime = parentIdentity.credentialLifetime)
    credential.parentIdentityId = parentIdentity.id
    record(credential)
    return credential

# collector class: issuanceRatePerHour = 10, maxLiveCredentials = 4
# a loop at 200 calls/minute is denied from the 11th call of the hour

go deeper

for a junior

Recall the distinction: a limit protecting a service from overload and a limit bounding what one caller may do are different controls, even when both are described as rate limits.

for a middle

Explain the mechanics of both axes — mints per interval and credentials live at once — and why each on its own leaves a gap that the other closes.

for a senior

Show that you would derive the numbers from the consumer's real behaviour, make denials observable, and name what breaks when a legitimate scale-out meets the cap.

for a principal

Own the standard: caps belong to identity classes declared at onboarding, with a reviewed raise path, and where the store offers no per-identity ceiling, decide consciously whether a broker in front of it is worth the extra dependency.

## Two limits that share a name and do different jobs Both are "limits on issuance", which is why this question separates candidates. - A **capacity limit** exists so the store keeps serving. It caps how much work the service as a whole will accept, so that a stampede does not take down the dependency the whole estate waits on. It is sized against the store's own throughput. - A **per-identity cap** exists so that one identity cannot do unbounded damage **while the store is perfectly healthy**. It is sized against what that identity legitimately needs. The attack in question never stresses the first one. Take a fleet of 800 edge collectors and a store whose issuance ceiling is 5,000 calls a minute. One taken-over host running 200 calls a minute is **4% of the ceiling**. The store does not degrade, nothing is shed, no latency alarm moves, and the capacity telemetry is flat. A limit that only fires when the service is in trouble cannot bound an attacker who is careful to leave the service healthy. ## The two axes a per-identity cap needs | Cap | What it counts | What it catches | What it misses alone | |---|---|---|---| | **Flow cap** | mints by this identity per interval | the loop: thousands of calls in minutes | a patient attacker minting at the legal rate for hours | | **Stock cap** | this identity's credentials currently live | slow accumulation to a large live set | a burst when lifetimes are short enough to keep the live count low | Each misses what the other catches, so the usual answer is both. A third limit is worth naming: a **maximum life** on each issued credential, which bounds the tail rather than the volume and is a different control from either cap. ## Deriving the numbers from legitimate need The temptation is to set the cap from headroom — "the store can do 5,000 a minute, so 100 an identity is generous". That reproduces the original problem one level down. Set it from the consumer's actual behaviour instead: - A collector holds **one** credential against one datastore and renews it hourly. - Overlap during renewal means **two** may be live briefly. - Restarts, retries and a redeploy might legitimately mint a handful more in a bad hour. That argues for a stock cap of about **4** and a flow cap of about **10 an hour**. The collector never notices. The twenty-minute loop that would have produced 4,000 credentials now produces **at most 10**, and it produces them in the first few seconds and then starts getting denials — which is itself the signal. ## A cap that fires must be loud A cap is a control and a detector at the same time, and it is only the second one if the denial is visible: - **Deny and emit.** A denial carrying the identity, the count in the window and the cap it hit is a high-quality alert: legitimate consumers running under their cap generate none of them. - **Count per identity, not per pool.** A bucket shared by a fleet lets one member consume everyone's allowance and starve its neighbours, which turns a security control into an outage. - **Make the raise path explicit.** Teams will hit a cap during a legitimate scale-out. If the only way through is an emergency change with no reviewer, the cap gets removed the first time it hurts. ## Designs genuinely differ here Some stores let you attach a ceiling to the identity or to the role that requests issuance; some expose only a global setting; some offer a maximum life per credential but no count at all. Where the store offers neither axis, the cap has to live in front of it — a small service the workloads call, which holds the counters and forwards only the requests under the cap. That is a real architectural cost, and it is worth naming in an interview rather than assuming the store has the feature. ## The trade-off you are accepting A cap can bite something legitimate. Five hundred new collectors starting at once during a region build-out are, from the store's point of view, indistinguishable from a fleet-wide loop. The usual resolution is that the cap belongs to an **identity class** rather than to one identity, with the class's numbers declared when the workload onboards and a documented, reviewed path to raise them. That keeps the ceiling honest: it is a statement about what this kind of workload needs, which somebody owns and can be asked to defend.

  • Why is a flow cap on its own not enough?
    Because it only limits speed. An attacker that stays under the ceiling — minting at the permitted 10 an hour instead of 200 a minute — accumulates steadily and never trips it. If those credentials outlive the hour, the live set grows without limit. A stock cap on simultaneously live credentials is what bounds the total, and it is the one that makes a slow campaign visible.
  • What happens to a legitimate consumer the first time your cap fires?
    It fails to obtain a credential, which for most consumers means failing to start or failing its next renewal. That is why a cap has to be sized from real behaviour, has to deny loudly rather than silently, and needs a documented raise path. A cap nobody can raise without an emergency change is a cap that gets deleted during the first incident it causes.
  • Where does the cap live if the store has no per-identity setting?
    In front of it. A small service the workloads call holds the per-identity counters, applies both caps, and forwards only permitted requests to the store — with the store's direct issuance interface restricted to that service. It is a real cost: another dependency on the request path, its own availability story, and its own counters to keep consistent, so it is worth deciding deliberately rather than by default.

saying these in an interview costs you the question

  • A global rate limit already bounds each identity
  • The store would have degraded if the volume mattered
  • Set the cap from the store's spare capacity
  • One shared bucket across the fleet is equivalent
  • A cap can be silent as long as it denies