skip to content

A batch load writes ten million entries with the same lifetime, so every deadline lands in the same second — what happens inside the store?

level: seniorimportance: nice to knowfreq 38%

answer

  1. the deadline itself schedules nothing
  2. a dead population, not a spike of work
  3. the drain is paced, the death is not
  4. spread the lifetimes to spread the work

basics

~10 s

At that second, ten million entries stop being servable and nothing else happens. The reclaim work appears at once but drains gradually, so memory slopes down - or stays flat if nothing reads them.

solid answer

~50 s

The deadline itself is free: the entries become unservable and no work is scheduled. What lands is a large **dead population**, and how it drains depends on the routes the store has. A sampled background sweep now has far more to find, but its budget per cycle is fixed, so it works the backlog down over time and reported memory falls as a slope rather than a cliff — while it drains, that reclaim work is competing with request serving. If nothing reads those entries and the tier is under no pressure, nothing drains at all and the full batch stays resident past its deadline. Either way, headroom must cover the whole batch well beyond the moment it dies. The practical remedy is to vary the lifetime slightly per entry so the deaths, and therefore the reclaim work, spread out.

go deeper

for a junior

Remember that deadlines landing together do not free anything together. The entries stop being servable at once; the memory comes back later and gradually.

for a middle

Explain the drain: a sampled pass has a fixed budget per cycle, so a large dead population is worked off over many cycles rather than cleared in one.

for a senior

Bring the sizing conclusion — a batch's footprint outlives its deadline, so waves of loads can be resident simultaneously — and offer spreading the lifetimes as the cheap fix.

for a principal

The call to own is whether deadlines may be used as the disposal mechanism for bulk loads at all, given that no store in this class will commit to returning the space by a given time.

## The moment itself costs nothing It helps to start with what does **not** happen. At the second all those deadlines pass, the store does no work. There is no wake-up, no queue drained, no pass triggered. Ten million entries simply become unservable: any read of one now reports a miss. Reported memory does not move, and serving latency does not change. What has changed is the store's internal accounting of reality: the dead population went from near zero to ten million in one instant, and every route that returns memory now has far more to do. ## How the population drains | route | behaviour under a mass deadline | effect on the memory graph | |---|---|---| | reclaim on access | fires only for entries someone reads; a batch nobody reads gets nothing | none | | the background sweep | finds dead entries far more often per sample, but the per-cycle budget is fixed | a slope, over many cycles | | reclaim under allocation pressure | fires when the store needs the room, and dead entries are given up first | a step, whenever pressure arrives | Two shapes fall out of that table: 1. **Sweep-driven store, entries nobody reads.** The pass now finds something dead almost every time it samples, which is as efficient as it gets, but it still only does its budget's worth of work per cycle. The backlog drains over many cycles, and while it does, reclaim work is competing for the same machine that answers requests. The memory graph slopes down over minutes rather than dropping at the deadline. 2. **Access-driven store, or no sweep at all.** Nothing drains. The full batch stays resident indefinitely, and the memory comes back only when the store needs the room. Neither shape is a defect. They are what the class does, and which one you get is a property of the store you chose — which is exactly why an answer that asserts one of them as "what happens" is wrong somewhere real. ## Why this is a sizing question, not a latency question The thing to take to a capacity plan is that **a batch's footprint outlives its deadline**. If ten million entries with a one-hour lifetime fit comfortably, that tells you nothing about whether a second batch written an hour later fits, because the first batch may still be entirely resident. On a tier that loads in waves, the worst case is several waves resident at once, all of them dead. The secondary effect is on the store's own work. Where reclaim competes with serving, a mass of dead entries means a period of elevated internal work that you did not schedule and cannot pace — which is why spreading the deaths out is worth doing even when the memory is affordable. ## What to do about it - **Vary the lifetime per entry.** Adding a small random spread — a few percent of the lifetime — turns one enormous dead population into a stream of small ones. The total reclaim work is identical; its distribution over time is not, and the distribution is the whole problem. - **Delete explicitly when the batch is genuinely finished with**, rather than letting a deadline do it. That returns memory on your schedule and makes the cost visible where you chose to pay it. - **Size for the batch surviving well past its deadline**, not for it disappearing at one. - **Do not reach for a shorter lifetime as the fix.** Entries die sooner and are reclaimed by the same unchanged routes, so the resident footprint barely moves. ## The answer that misses the point The most common wrong turn is to answer with what happens *outside* the store — the load that falls on whatever produced the data once its entries stop being servable. That is a real concern and a different subject with a different owner. The question here is about the store's own behaviour: a dead population appearing in one instant, draining at a pace nothing about the deadline controls, and a memory profile that stays at its peak long after the data stopped being useful.

  • Why does adding a small random spread to each entry's lifetime help, when the total work is unchanged?
    Because the problem is distribution, not volume. A few percent of spread turns one instantaneous dead population into a stream of small ones, so whatever route reclaims them does a steady trickle of work instead of facing a backlog it can only drain at its fixed pace. The memory profile flattens for the same total cost.
  • Does a mass deadline make reads slower at that instant?
    Not by itself. The deadline schedules no work, and a read of a dead entry is an ordinary lookup that ends in a miss. What can show up afterwards is elevated internal reclaim work on stores where that work shares the machine with request serving, spread over however long the backlog takes to drain.
  • How would you tell whether a given tier drains such a backlog at all?
    Load a batch with a short lifetime on a tier under no pressure, stop reading it, and watch the store's reported memory after the deadline passes. A gradual decline means a background pass is finding them; a flat line until you press the tier means reclaim there depends on readers and on scarcity.

saying these in an interview costs you the question

  • Says the store frees ten million entries at the instant the deadlines pass.
  • Answers with the load on whatever produced the data instead of the store's own reclaim work.
  • Assumes a background pass drains the backlog faster simply because there is more of it.
  • Expects the memory graph to show a cliff at the deadline.
  • Believes a shorter lifetime would reduce the resident footprint.