skip to content

Ephemeral State Workloads

State the tier is asked to hold that is not a copy of anything else - sessions, claims, counters, dedupe records, live delivery - and what each one loses when an entry disappears.

on this pageshow

questions

21

A quota counter on a shared volatile tier uses one key per subject per period and never gets a lifetime - what does that cost?

level: juniorimportance: must knowfreq 62%

answer

  1. growth tracks history, not traffic
  2. nobody reads last period's counter
  3. the store has no sweeper
  4. one key per subject per period
  5. attach the deadline on creation

basics

~20 s

Dead counters accumulate. The keyspace grows with the whole history of subjects and periods rather than with the active ones, and on a tier that removes entries under memory pressure those dead keys crowd out live ones.

solid answer

~50 s

The key encodes the subject the quota is charged to and the period being counted, which is what makes the counter self-rotating: when the period changes the key name changes, the old counter is never addressed again, and the new one starts absent, which the code reads as zero. But never addressed is not the same as removed. The store has no schema and no idea what your key name means, so nothing sweeps keys naming a period that has ended; the only general mechanism that removes an entry on its own is a lifetime attached to that entry. Without one, every subject-period pair that ever existed stays resident forever. The leak is invisible to functional tests, because nothing ever reads those keys again - it shows up only as resident size that grows and never falls.

go deeper

for a junior

Remember the two halves of the key - who is being charged, and which period - and the fact that naming the period is what makes the counter roll over by itself. Then remember that rolling over is not deleting.

for a middle

Be able to explain why the store cannot clean up after you: it has no schema and no meaning for your key name, so a lifetime attached at write time is the only general removal mechanism. Say where that attachment can be dropped and how you cover the gap.

for a senior

Show that you know the leak is silent - no failing test, no error, only resident size that never falls - and that on a tier which removes entries under memory pressure it turns into live counters being dropped. Name the check you would put on the tier to catch it.

for a principal

The question behind this one is whether the enforcement design's memory is proportional to load or to history. Take a position on whether a per-subject, per-period keyspace belongs on a shared tier at all when the subject population is unbounded and attacker-controlled.

## The shape of the thing A quota is a promise of the form *this subject may do this much in this period*. Enforcing it needs a number that every server handling that subject's requests agrees on, so the number lives on the shared volatile tier rather than inside any one process. It is raised by a **server-side increment**: the tier adds one and hands back the new total in a single operation, so two overlapping requests cannot both add one to the same starting value. Not every store in this class can do that - a store that holds an **opaque value** and only hands the bytes back cannot add one for you, and on such a store a shared counter is a materially different and weaker design. The key has to carry two things: - **the subject** the quota is charged to - the account, the client credential, the tenant, the calling network address; - **the period** being counted - the day, the hour, the minute, spelled out in the key itself. Putting the period in the key is what makes the counter self-rotating. Nothing has to reset anything. When the period changes, the key name changes, the previous counter stops being addressed, and the new one is absent - which the enforcement code reads as zero and creates on first use. ## Why nothing deletes last period's counter A volatile store of this class knows nothing about your key beyond its bytes. There is no background task that understands *keys naming a period that has ended*, and no store in this class will infer one from a naming convention. The only general mechanism that removes an entry without anyone asking is **a lifetime attached to the entry**: you say, at write time, how long this entry may live, and the store reclaims it after that. So a counter created with no deadline is permanent. Worse, it is permanent *and unread*: after its period ends nothing ever addresses it again, so no test fails, no error is logged, and no latency moves. The only symptom is that the resident size of the tier climbs and never comes down. ## What the growth actually is | | counter carries a lifetime | counter carries none | |---|---|---| | keys held at any moment | active subjects, times the one or two periods currently in play | every subject that ever appeared, times every period it appeared in | | growth driver | current traffic | elapsed time and historical subject count | | what shrinks it | the deadline passing | a person noticing and deleting by hand | | symptom when wrong | none | resident size that only ever rises | The second column is the important one: the design's memory is proportional to *history*, not to load. A service with ten thousand subjects and a per-minute period manufactures ten thousand permanent keys a minute, whose cost is per-entry overhead and the key string rather than the number it holds. On a tier configured to remove entries when it reaches its memory ceiling, the leak becomes a correctness problem rather than only a capacity one: the dead counters compete for the ceiling with live ones, and the store's choice of what to remove does not know which counters are still being charged against. ## Where the deadline gets lost Attaching the lifetime sounds like one line of code, and it is the line most often missing, because of a gap that is real on many stores: 1. The counter is usually brought into existence *by the increment itself* - the first request of the period increments a key that is absent, and the store creates it at zero and adds one. 2. Stores differ in whether an increment can carry a lifetime at all. On many it cannot: the increment raises the number, and dating the entry is a second operation. 3. Between those two operations the counter exists with no deadline. A process that dies there, or a network failure that loses the second call, leaves an undated counter behind - one that will never be reclaimed, and which can leave that subject refused for as long as it survives. The usual repair is to use what the increment returns. Where the increment hands back the new total, a total of one means *this request created the key*, and that is the moment to attach the deadline; re-applying it on a later increment is harmless and covers the lost second call. Where the store can attach a lifetime in the same operation that creates the entry, use that and the gap does not exist. ## Sizing the deadline The lifetime must **outlive the period it names**, not equal it - the counter still has to be there for the last request of the period, and clocks between callers are not identical. A fixed lifetime counted from the write, set to the period length plus a margin, is the ordinary choice. Two details survive the sizing. First, memory does not come back at the instant the deadline passes; *when* it comes back differs by store, so plan for more than one period's counters to be resident at once. Second, the deadline must not be pushed forward by each use: that is a different lifetime kind, and on a period counter it stops the period from ever ending.

  • What changes if the period is a minute rather than a day?
    The number of dead keys per subject multiplies by the number of periods elapsed, so a minute-long period manufactures roughly fourteen hundred times the leak of a daily one for the same subject count. The live footprint barely moves - one or two keys per active subject either way - which is exactly why the undated version is the only one whose cost depends on the period length.
  • The counter stores a small number. Is the memory really worth worrying about?
    The number is not the cost. Every entry carries the key string plus per-entry bookkeeping the store keeps for it, and in this shape the key is long - a subject identifier and a period stamp - while the value is a handful of digits. Overhead dominates, and it is paid per subject per period, forever.
  • Why not let the application delete last period's counters instead?
    To delete them the application must know every subject that was active, which is the same unbounded set you are trying not to keep. Any sweeper also has to run, succeed, and cover the subjects seen by instances that have since been replaced. A deadline attached at write time delegates all of that to the store and fails closed: if the sweeper is what is missing, the key is already dated.

A shop that writes each day's till total on a fresh sheet of paper. Naming the sheet by date means today's total is never confused with yesterday's, and nobody ever needs yesterday's sheet again. But nothing in that habit throws the old sheets away - only an explicit rule, bin every sheet after a month, empties the drawer. The lifetime on the entry is that rule.

saying these in an interview costs you the question

  • Assumes the store deletes keys once the period they name has passed.
  • Says removal under memory pressure will clean the dead counters up.
  • Sizes the tier from the count of currently active subjects only.
  • Believes an increment always carries the entry's lifetime with it.
  • Treats an unread key as costing nothing because nothing addresses it.
  • Plans to sweep old counters from the application on a schedule.
open as a page

Why does a lease key that lets only one worker run a job carry a lifetime instead of living until its holder releases it?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A lifetime bounds a crash. If the holder dies between claiming and releasing, an entry that only a release deletes blocks the job forever; the lifetime lets the store drop the claim so another worker can take it.

open as a page

Why must per-user request state leave the application instance once a second instance exists, and what does the move cost?

level: juniorimportance: must knowfreq 72%

basics

~20 s

State held in one instance's memory is readable only by that instance, so the next request may land on a stranger. A shared tier makes all instances equivalent, at the price of a round trip per authenticated request.

open as a page

A price feed is broadcast to recipients attached to a shared volatile tier; one drops for two seconds — what does it receive on reconnect?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Nothing from those two seconds. A broadcast reaches only the connections attached at the instant it is sent, and is then forgotten, so the reconnecting recipient starts from the next message — and neither side can tell a gap happened.

open as a page

Why must a deduplication record be written before the guarded work rather than after it, and what must the store offer?

level: middleimportance: must knowfreq 68%

basics

~20 s

A retry arrives while the first attempt is still running, so a record written after the work is absent exactly when it is needed. Claim it first, with a conditional create: a write that succeeds only if the key is absent.

open as a page

How long must a deduplication record live, and why is the first retry delay the wrong number to size it against?

level: middleimportance: must knowfreq 64%

basics

~20 s

Size a deduplication record against the retry window - the longest span over which the same work can arrive again, including a stalled consumer, a nightly reconciliation resend or an operator replay - never against the sender's first retry delay.

open as a page

In a shared volatile tier, how can a worker deleting its own lease key release a claim that another worker now holds?

level: middleimportance: must knowfreq 60%

basics

~20 s

A blind delete removes whatever is under the key, not the claim you made. If your lease reached its deadline mid-job and a second worker claimed that key, your release frees the second worker's claim.

open as a page

A worker takes an item from a short-lived work channel on the volatile tier and dies mid-job — what happens to that item?

level: middleimportance: must knowfreq 58%

basics

~20 s

It is gone. The take removed it and the store keeps no record that it was handed out, so no acknowledgement is owed and none is missing — nothing in the system can detect that the work was never finished.

open as a page

Which requirements tell you a fan-out workload has outgrown the volatile tier's broadcast and belongs on a durable log with a broker?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Four: something must be kept for a recipient absent at send time; someone must know a recipient handled it; a recipient needs a resumable position; or history must be replayable. Any one means a broker.

open as a page

What does a deduplication record buy a payment service, and what does a customer see if a restart emptied the store holding it?

level: juniorimportance: should knowfreq 58%

basics

~20 s

A deduplication record marks a request identifier as already handled, so a repeated arrival of the same work is skipped instead of charged again. If a restart empties the store, the record is gone and the customer is charged twice.

open as a page

A daily quota per customer: how does anchoring the period to a shared clock rather than the customer's first call change the store?

level: middleimportance: should knowfreq 55%

basics

~20 s

A shared clock puts the period in the key, so no extra state is stored and every subject resets at the same instant. First-call anchoring puts the period start in the counter's own lifetime, making the deadline load-bearing.

open as a page

Per-user state on a shared tier uses an access-extended lifetime rather than a fixed one — what does that cost per request?

level: middleimportance: should knowfreq 58%

basics

~20 s

An access-extended lifetime is pushed forward by each use, so every authenticated request must write to the tier as well as read from it. A read-mostly workload becomes roughly one write per request per active user.

open as a page

Your quota counters live on a tier that removes entries under memory pressure - how far over the quota can one subject go, and how do you bound that in advance?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Up to the full quota again for each removal, because an absent counter is indistinguishable from a subject's first call. You bound it by keeping counters out of contention for the ceiling and deciding in advance what happens when the increment fails.

open as a page

Your deduplication records sit on a tier that removes entries under memory pressure and the same payout arrives twice, so what happens and where should the guard live?

level: seniorimportance: should knowfreq 52%

basics

~20 s

The second arrival finds nothing, is treated as new work, and the payout goes out twice: once-only handling degrades to at-least-once. A record the tier can remove is an optimisation, so a payout that must never repeat needs a durable uniqueness rule.

open as a page

How can two workers end up holding the same lease after a failover, when each of them won a conditional create?

level: seniorimportance: should knowfreq 48%

basics

~20 s

The first claim was acknowledged by a node that failed before any copy held the write. The promoted copy has no such key, so the second worker's conditional create succeeds — and neither worker is told.

open as a page

A job outruns its lease's deadline while still working — what does renewing the lease buy, and what can renewal not prevent?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Renewal replaces one guess with a smaller one: the lifetime need only exceed the worst pause between renewals, not the worst run. It cannot stop a stalled holder losing a claim it still believes it holds.

open as a page

Sign-out deletes a user's state entry from the shared tier — which copies of that state does the deletion fail to reach?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Deleting the entry ends the copy on the node that took it. A copy inside the tier that has not received it yet, one an application instance still holds in memory, and the client's own copy all survive.

open as a page

A recipient attached to the store's broadcast surface reads slower than messages are sent — where does the undelivered backlog sit, and what ends it?

level: seniorimportance: should knowfreq 46%

basics

~20 s

On the server, in the delivery queue the store holds for that recipient's connection, drawing on the same memory as the data. Stores that bound those bytes close the slow recipient's connection: the tier is protected, the recipient sacrificed.

open as a page

Your product both throttles free users and bills paid users by call volume - which of those counts may live only on the volatile tier?

level: principalimportance: should knowfreq 38%

basics

~20 s

The throttle count may; the billed count may not. The test is what a silent return to zero costs: an over-granted allowance is a priced degradation, while a number an invoice is computed from must stay reconstructible.

open as a page

Two workers must never both charge a customer — what may a lease in a shared volatile tier be responsible for here?

level: principalimportance: should knowfreq 44%

basics

~10 s

It may be responsible for cost, not for correctness. A lease keeps the common case to one worker, but it cannot promise a single holder, so the guard against a second charge belongs downstream.

open as a page

Your operations console keeps panes live via transient fan-out on the volatile tier and you will not add a broker — how do you design so a missed message is a named degradation rather than a wrong screen?

level: principalimportance: should knowfreq 40%

basics

~20 s

Make the broadcast a hint, never the truth. Panes read state from the durable source of record on attach, on reconnect and on a periodic refresh; the broadcast only says read sooner. The miss then costs bounded staleness you can state.

open as a page