A quota counter on a shared volatile tier uses one key per subject per period and never gets a lifetime - what does that cost?
answer
- growth tracks history, not traffic
- nobody reads last period's counter
- the store has no sweeper
- one key per subject per period
- attach the deadline on creation
basics
~20 sDead counters accumulate. The keyspace grows with the whole history of subjects and periods rather than with the active ones, and on a tier that removes entries under memory pressure those dead keys crowd out live ones.
solid answer
~50 sThe key encodes the subject the quota is charged to and the period being counted, which is what makes the counter self-rotating: when the period changes the key name changes, the old counter is never addressed again, and the new one starts absent, which the code reads as zero. But never addressed is not the same as removed. The store has no schema and no idea what your key name means, so nothing sweeps keys naming a period that has ended; the only general mechanism that removes an entry on its own is a lifetime attached to that entry. Without one, every subject-period pair that ever existed stays resident forever. The leak is invisible to functional tests, because nothing ever reads those keys again - it shows up only as resident size that grows and never falls.
go deeper
Remember the two halves of the key - who is being charged, and which period - and the fact that naming the period is what makes the counter roll over by itself. Then remember that rolling over is not deleting.
Be able to explain why the store cannot clean up after you: it has no schema and no meaning for your key name, so a lifetime attached at write time is the only general removal mechanism. Say where that attachment can be dropped and how you cover the gap.
Show that you know the leak is silent - no failing test, no error, only resident size that never falls - and that on a tier which removes entries under memory pressure it turns into live counters being dropped. Name the check you would put on the tier to catch it.
The question behind this one is whether the enforcement design's memory is proportional to load or to history. Take a position on whether a per-subject, per-period keyspace belongs on a shared tier at all when the subject population is unbounded and attacker-controlled.
## The shape of the thing A quota is a promise of the form *this subject may do this much in this period*. Enforcing it needs a number that every server handling that subject's requests agrees on, so the number lives on the shared volatile tier rather than inside any one process. It is raised by a **server-side increment**: the tier adds one and hands back the new total in a single operation, so two overlapping requests cannot both add one to the same starting value. Not every store in this class can do that - a store that holds an **opaque value** and only hands the bytes back cannot add one for you, and on such a store a shared counter is a materially different and weaker design. The key has to carry two things: - **the subject** the quota is charged to - the account, the client credential, the tenant, the calling network address; - **the period** being counted - the day, the hour, the minute, spelled out in the key itself. Putting the period in the key is what makes the counter self-rotating. Nothing has to reset anything. When the period changes, the key name changes, the previous counter stops being addressed, and the new one is absent - which the enforcement code reads as zero and creates on first use. ## Why nothing deletes last period's counter A volatile store of this class knows nothing about your key beyond its bytes. There is no background task that understands *keys naming a period that has ended*, and no store in this class will infer one from a naming convention. The only general mechanism that removes an entry without anyone asking is **a lifetime attached to the entry**: you say, at write time, how long this entry may live, and the store reclaims it after that. So a counter created with no deadline is permanent. Worse, it is permanent *and unread*: after its period ends nothing ever addresses it again, so no test fails, no error is logged, and no latency moves. The only symptom is that the resident size of the tier climbs and never comes down. ## What the growth actually is | | counter carries a lifetime | counter carries none | |---|---|---| | keys held at any moment | active subjects, times the one or two periods currently in play | every subject that ever appeared, times every period it appeared in | | growth driver | current traffic | elapsed time and historical subject count | | what shrinks it | the deadline passing | a person noticing and deleting by hand | | symptom when wrong | none | resident size that only ever rises | The second column is the important one: the design's memory is proportional to *history*, not to load. A service with ten thousand subjects and a per-minute period manufactures ten thousand permanent keys a minute, whose cost is per-entry overhead and the key string rather than the number it holds. On a tier configured to remove entries when it reaches its memory ceiling, the leak becomes a correctness problem rather than only a capacity one: the dead counters compete for the ceiling with live ones, and the store's choice of what to remove does not know which counters are still being charged against. ## Where the deadline gets lost Attaching the lifetime sounds like one line of code, and it is the line most often missing, because of a gap that is real on many stores: 1. The counter is usually brought into existence *by the increment itself* - the first request of the period increments a key that is absent, and the store creates it at zero and adds one. 2. Stores differ in whether an increment can carry a lifetime at all. On many it cannot: the increment raises the number, and dating the entry is a second operation. 3. Between those two operations the counter exists with no deadline. A process that dies there, or a network failure that loses the second call, leaves an undated counter behind - one that will never be reclaimed, and which can leave that subject refused for as long as it survives. The usual repair is to use what the increment returns. Where the increment hands back the new total, a total of one means *this request created the key*, and that is the moment to attach the deadline; re-applying it on a later increment is harmless and covers the lost second call. Where the store can attach a lifetime in the same operation that creates the entry, use that and the gap does not exist. ## Sizing the deadline The lifetime must **outlive the period it names**, not equal it - the counter still has to be there for the last request of the period, and clocks between callers are not identical. A fixed lifetime counted from the write, set to the period length plus a margin, is the ordinary choice. Two details survive the sizing. First, memory does not come back at the instant the deadline passes; *when* it comes back differs by store, so plan for more than one period's counters to be resident at once. Second, the deadline must not be pushed forward by each use: that is a different lifetime kind, and on a period counter it stops the period from ever ending.
- What changes if the period is a minute rather than a day?The number of dead keys per subject multiplies by the number of periods elapsed, so a minute-long period manufactures roughly fourteen hundred times the leak of a daily one for the same subject count. The live footprint barely moves - one or two keys per active subject either way - which is exactly why the undated version is the only one whose cost depends on the period length.
- The counter stores a small number. Is the memory really worth worrying about?The number is not the cost. Every entry carries the key string plus per-entry bookkeeping the store keeps for it, and in this shape the key is long - a subject identifier and a period stamp - while the value is a handful of digits. Overhead dominates, and it is paid per subject per period, forever.
- Why not let the application delete last period's counters instead?To delete them the application must know every subject that was active, which is the same unbounded set you are trying not to keep. Any sweeper also has to run, succeed, and cover the subjects seen by instances that have since been replaced. A deadline attached at write time delegates all of that to the store and fails closed: if the sweeper is what is missing, the key is already dated.
A shop that writes each day's till total on a fresh sheet of paper. Naming the sheet by date means today's total is never confused with yesterday's, and nobody ever needs yesterday's sheet again. But nothing in that habit throws the old sheets away - only an explicit rule, bin every sheet after a month, empties the drawer. The lifetime on the entry is that rule.
saying these in an interview costs you the question
- Assumes the store deletes keys once the period they name has passed.
- Says removal under memory pressure will clean the dead counters up.
- Sizes the tier from the count of currently active subjects only.
- Believes an increment always carries the entry's lifetime with it.
- Treats an unread key as costing nothing because nothing addresses it.
- Plans to sweep old counters from the application on a schedule.