A job clears each key's entry with a per-key scheduled callback. How does that bound the retained set, and how can the callbacks themselves leak?
answer
- registration removes nothing
- the handler must delete
- pending wake-ups are entries too
- cancel, or round the deadline
- a frozen clock never fires it
basics
~20 sThe callback is a wake-up registered against one key and one moment; what bounds the retained set is the handler that deletes the entry when it fires. Pending callbacks are stored entries too, so one per record leaks just as badly.
solid answer
~50 sA per-key scheduled callback is a wake-up the job registers against one key and one future moment, which fires later and runs a handler. It bounds the retained set — everything the job still holds between records — only because that handler explicitly deletes the key's entry; the registration on its own removes nothing. The leak is that a pending callback is itself a stored entry, snapshotted and restored like any other. Register a fresh one on every record to extend a deadline, and unless you cancel the previous one, or the runtime collapses duplicates for the same key and moment, you accumulate one pending callback per record while the entry you were trying to bound is still there. Two further leaks: a deadline measured against the job's progress assertion never fires while that assertion is frozen, and a handler that deletes nothing is bookkeeping.
go deeper
Recall that a callback is a wake-up registered against one key and one moment, and that the entry disappears only if the handler you wrote deletes it.
Explain that pending callbacks are stored entries too, and give the discipline that stops one registration per record: cancel the previous one, or round deadlines so they collapse.
Show the judgment for choosing callbacks over an interval rule — you need to emit something at the moment of removal — and account for the case where the clock driving the deadline stops advancing.
Set the convention: hand-written cleanup handlers are reviewed for the deletion line and for the registration discipline, because both defects are silent in test and fatal in month three.
## What a per-key scheduled callback is A **per-key scheduled callback** is a wake-up the job registers against one key and one moment. When that moment arrives the runtime invokes a handler, under that key, with access to that key's entries. It is the imperative alternative to a declarative expiry rule, and it is what you reach for when the removal condition is not "a fixed interval since the last write" but something the author computes: a burst of activity from one user is over when no event for them has arrived for thirty minutes; an order's buffered record can go once the matching confirmation has been seen or the deadline passed. ## What actually bounds the retained set The callback does not remove anything. The **handler** does, and only if the author wrote the deletion: 1. A record arrives for key K; the step writes or updates K's entry. 2. The step registers a callback for K at a computed moment. 3. At that moment the handler runs under K. 4. The handler emits whatever result is owed **and deletes K's entry**. Drop step 4 and you have a job that wakes up punctually to do nothing, with a retained set that grows exactly as fast as it did before. This is the single most common defect in hand-written cleanup, and it is invisible in tests, because a test that runs for a minute never notices the entries nobody deleted. ## Three ways the callbacks leak - **One registration per record.** The natural way to extend a deadline is to register a new callback each time a record arrives. Pending callbacks are stored entries in their own right — they are snapshotted, restored and counted like any other — so a thousand records for one key can leave a thousand pending wake-ups for that key. Some runtimes collapse two callbacks registered for the same key at the same moment into one; others keep both, and then the deduplication depends entirely on your arithmetic. The disciplines that work are to cancel the previous callback when you register the next, or to round every deadline to a coarse grain so re-registrations collide and collapse, or to register one callback and, when it fires, check the entry's own last-touched field and re-register only if the key is still active. - **A callback that can never fire.** If the deadline is measured against the job's assertion of how far the input has progressed, and that assertion stops advancing because the input went quiet, the callback stays pending indefinitely and the entries it would have cleared stay with it. Why the assertion stalls is another node's subject; the consequence for the retained set is this node's. - **Keys that never come back.** Registering a callback far in the future for a key that will never be seen again is fine for one key and ruinous for a key space that turns over — every dead key leaves a pending wake-up until its moment arrives, so the pending set tracks key churn over the whole horizon, not the active key count. ## Callback against declarative expiry | | per-key scheduled callback | entry expiry | |---|---|---| | removal condition | anything the author can compute, including emitting a final result first | a fixed interval after last write or last read | | who deletes | your handler, explicitly | the runtime | | can emit before removing | yes — this is its main advantage | no, removal is silent | | its own footprint | pending callbacks are stored entries and can leak | none beyond a timestamp on the entry | | failure mode when wrong | wakes and deletes nothing, or accumulates registrations | removes too early or too late, silently | The reason to accept the extra machinery of callbacks is almost always that **something must be emitted at the moment of removal** — the closing summary of a burst of activity, a final count, an alert that a matching record never arrived. A declarative expiry rule cannot do that: it drops the entry and tells nobody. ## What varies between engines Whether callbacks exist at all, whether they can be cancelled, whether two registrations at the same key and moment collapse, and whether a pending callback survives a restart are all points on which this class of systems genuinely differs. Where a continuous input is processed as repeated small finite jobs, deadlines are evaluated once per chunk, so a callback fires at the first chunk boundary at or after its moment rather than at the moment itself. In the two-phase disk-to-disk batch model nothing is retained between runs, so the mechanism has no meaning there. State which model you mean before asserting a behaviour. ## Interview framing The strong answer separates three things that beginners merge: the deadline, the handler, and the deletion. The deadline decides *when*; the handler decides *what is emitted*; and only an explicit deletion decides *whether anything actually shrinks*.
- When is a per-key scheduled callback worth its extra machinery over a declarative expiry rule?When something must be emitted at the moment of removal — the summary of a burst of activity, a final count, or an alert that the matching record never arrived. A declarative rule drops the entry silently and tells nobody. It is also the answer when the removal condition is computed from the entry's own contents rather than from a fixed interval.
- How do you extend a key's deadline without accumulating a pending callback per record?Cancel the previous registration before making the new one where the runtime allows it; otherwise round every deadline to a coarse grain so repeated registrations land on the same moment and collapse. A third option is to register once, and when it fires compare the entry's last-touched field against the deadline: delete if idle, re-register once if not.
- Is a pending callback carried across a restart?Where callbacks are part of the retained set, yes — they are written into the durable snapshot and restored with everything else, which is why a large pending set slows restart like any other entry. Whether that holds is one of the things engines differ on, so confirm it for the runtime you are on rather than assuming it.
saying these in an interview costs you the question
- Thinks registering a callback removes the entry by itself.
- Registers one per record and never cancels the previous registration.
- Assumes pending callbacks cost nothing and are not stored anywhere.
- Believes a deadline always fires even when the clock driving it is frozen.
- Uses a callback where a plain interval rule would do, adding failure modes.