skip to content

Your quota counters live on a tier that removes entries under memory pressure - how far over the quota can one subject go, and how do you bound that in advance?

level: seniorimportance: should knowfreq 50%

answer

  1. absent is indistinguishable from not yet started
  2. silent on the happy path
  3. a whole quota per removal
  4. separate the four disappearances
  5. fail-open or fail-closed, decided once

basics

~20 s

Up to the full quota again for each removal, because an absent counter is indistinguishable from a subject's first call. You bound it by keeping counters out of contention for the ceiling and deciding in advance what happens when the increment fails.

solid answer

~50 s

The enforcement code treats absent as zero - it has to, because that is how a period legitimately begins. So a removed counter does not fail loudly; it grants the subject a whole fresh allowance, silently, mid-period. The worst case is not a few extra calls but the quota again per removal. You bound it before it happens rather than detecting it after: keep the counters off a tier that is also holding a large body of entries competing for the same ceiling, so they are rarely the ones chosen; watch how often entries are being removed at all, because a tier removing entries steadily is already telling you the quota is soft; and decide explicitly what enforcement does when the tier cannot answer, because the other ceiling behaviour in this class is to refuse writes rather than remove anything, and then the increment fails instead of resetting.

go deeper

for a junior

The key point to hold on to: a missing counter reads as zero, because that is exactly how a new period starts. Nothing errors, and the subject is simply allowed through again.

for a middle

Do the arithmetic out loud - a full quota per removal, not a handful of calls - and be able to say that the same zero arises from a passed deadline, a restart, and a promoted copy, not only from memory pressure.

for a senior

Show the bound being set in advance: separate the quota keyspace from whatever bulk workload competes for the ceiling, alert on the removal rate, and state the fail-open or fail-closed decision as a deliberate choice with a reason.

for a principal

Frame it as a priced degradation rather than a bug: what one over-granted subject costs, how many you tolerate per quarter, and at what cost of protection the answer changes to moving the authoritative count somewhere it cannot vanish.

## Absent means zero, and that is not a bug A period counter is created by the first call of the period. Before that call the key does not exist, and the enforcement code reads an absent key as a count of zero and allows the request. There is no way to write it otherwise: the tier cannot distinguish *this subject has not called yet* from *this subject's counter was removed twenty minutes into the period*. That is the whole failure. The removal is silent at the point where it matters. No error is raised, no branch is taken that differs from the ordinary happy path, and the subject is granted an entire fresh allowance in the middle of a period it had already exhausted. ## The size of the over-grant The intuition that a lost counter costs a few extra calls is wrong, and it is worth being explicit about the arithmetic. If the quota is *n* per period and the counter is removed *k* times during one period, the subject can be granted up to *n × (k + 1)* in that period. The tier does not remember that it removed anything, so each removal restores the full allowance rather than part of it. Worse, removals are not independent of the subject's behaviour. A subject making many calls holds a counter that is being touched constantly and is *also* contributing to the memory pressure that causes removals in the first place, which is one reason a heavy subject is a plausible repeat beneficiary rather than a statistical outlier. ## The four disappearances, and why they differ Removal under memory pressure is one cause of a zeroed counter. The others are worth separating because the answer differs: 1. **Removal under memory pressure.** Mid-period, unannounced, possibly repeated. This is the one you bound with capacity and placement. 2. **The deadline passing early.** If the lifetime was sized to the period instead of outliving it, or the clock the deadline is measured against differs from the one naming the period, the counter can end before the period does. 3. **A restart that kept nothing.** Every counter for every subject returns to zero at once - a population-wide over-grant for the remainder of the period. What survives a restart is a deployment choice: some keep nothing, some keep a periodic whole copy, some replay a log of writes, in which case the counter comes back but is behind by the loss span. 4. **A promoted copy that never received the increments.** Where a write is acknowledged before a copy holds it, a failover promotes a copy whose count is lower than the one callers were told about. Stores differ here and it is not a given: several in this class can be configured to acknowledge only once a copy has it, and some offer that per call, at the cost of latency on every increment. An answer that treats all four as *the counter was gone* misses that one is bounded by capacity planning, one by arithmetic on the lifetime, one by the durability posture, and one by the replication contract. ## Bounding it before it happens - **Keep the counters out of contention.** The counters are tiny and the pressure almost never comes from them; it comes from whatever bulk workload shares the tier. Separating the quota keyspace onto its own tier, or onto a deployment with its own ceiling, converts a shared risk into one you can size directly. - **Know which ceiling behaviour you are on.** Stores in this class do not all respond to a full tier the same way. Some remove entries to make room; some refuse writes instead. On the second kind the counter is never zeroed - the increment fails - which is a different and in some ways better problem, because it is visible. - **Decide fail-open or fail-closed, explicitly.** When the increment cannot be performed, enforcement either allows the request or refuses it. A protective quota in front of a fragile dependency usually wants to refuse; a courtesy quota on a public read endpoint usually wants to allow rather than take the product down because a tier is unwell. The point is that the decision is made once, deliberately, and not left to whatever the client library's exception path happens to do. - **Alert on the removal rate itself.** A tier steadily removing entries is telling you the quota is currently advisory. That signal is cheaper and earlier than trying to detect over-granted subjects after the fact. - **Say what the over-grant is worth.** If it is fairness or protection, one subject getting a second allowance occasionally is a priced degradation, and the honest answer is to accept and monitor it. If a second allowance means money, the counter is in the wrong home. ## What you should not claim Do not describe a counter on this tier as a hard cap. It is an enforcement mechanism whose accuracy degrades with the health of the tier, and the useful sentence in a design review is the bound - *a subject can exceed its quota by one quota per removal, we expect removals never, and we alert if any occur* - rather than a guarantee the tier does not provide.

  • How would you detect that this has happened at all?
    Not from the counters, which look normal afterwards. You detect it upstream: the tier's own removal rate, and the count of requests where the increment returned a total of one - the marker that a counter was just created. A spike of first-of-period creations in the middle of a period is the fingerprint of a batch of counters disappearing.
  • Your tier refuses writes at the ceiling instead of removing entries. Is the quota safer?
    More honest rather than safer. The counter is never silently zeroed, so no subject is over-granted by removal - but the increment fails, and enforcement has to choose between allowing an uncounted request and refusing a legitimate one. It converts an invisible correctness loss into a visible availability decision, which is usually the better trade because someone gets to make it.
  • Does splitting the keyspace across nodes change the over-grant?
    It changes its shape. One subject's counters for one period land on one node, so a single node's loss over-grants the subjects assigned to it rather than all of them, and the blast radius is a fraction of the population rather than everyone. It does not make any individual subject's over-grant smaller - that is still a whole quota.

saying these in an interview costs you the question

  • Says a lost counter costs only the few increments in flight.
  • Expects the tier to raise an error when a counter is missing.
  • Treats every cause of a zeroed counter as the same failure.
  • Calls a counter on a volatile tier a hard cap.
  • Leaves the behaviour on a failed increment to the client library's default.
  • Assumes every store removes entries rather than refusing writes when full.