skip to content

Why do test recipient identities need a scheduled expiry sweep when per-case cleanup already runs reliably?

level: seniorimportance: nice to knowfreq 22%

answer

  1. Cleanup assumes the run reaches the end
  2. What runs when nothing runs?
  3. Some claims have a cooling period
  4. Sweep by owner and age, not by run

basics

~20 s

Per-case cleanup only covers cases that reach their own end. A cancelled run, a process killed on a time limit, and a notification arriving after the case finished all leave identities behind. Some claims cannot be released on demand.

solid answer

~50 s

Deleting an identity at the end of its case handles the happy path and nothing else. **Runs die**: a cancelled pipeline, a process killed on a global time limit, a machine reclaimed mid-suite — cleanup is never scheduled, and there is no second chance because the only thing that knew about the identity is gone. **Notifications arrive late**: a delivery can land well after the case has passed, so an eagerly deleted mailbox destroys the evidence you would want for the next intermittent failure. **Some claims cannot be handed back on request** — a leased number with a cooling period, an address the delivery side will not forget immediately. A sweep that expires identities by owner, age and lease status is stateless: it does not care which run created them, or whether that run survived.

code

pseudocode · 12 lines
pseudocode
# on a timer; it never learns which run created anything
for identity in registry.where(owner = "automated-suite"):
    if identity.lease_expires > now:
        continue                       # a live case still holds it
    if now - identity.created_at < RETENTION:
        continue                       # inside the late-arrival window

    release(identity.mailbox)          # or the leased number
    forget(identity.address)           # consent and confirmation records
    registry.remove(identity)

emit_count(reclaimed)                  # a sudden jump means something upstream stopped cleaning up

go deeper

for a junior

Recall that cleanup code only runs when the case reaches it. Be ready to say what happens to whatever a test claimed when the run is cancelled halfway through.

for a middle

Explain how a sweep selects: by owner, age and lease status rather than by run, so it needs no knowledge of what created an identity and still works after that run has vanished.

for a senior

Demonstrate the judgement. Name the outcomes per-case deletion cannot cover, derive the retention window from the observed tail of late arrivals, and say how a live case protects its identity from the sweep.

for a principal

Own the cost and the blast radius: the sweep is shared infrastructure whose whole job is deletion. Decide who runs it, what its ownership filter is allowed to match, and what happens the day test and real recipients share one store.

## The assumption inside per-case cleanup Deleting an identity at the end of the case that created it is correct, cheap, and worth doing. It also rests on an assumption that fails often enough to matter: that the case reaches its own end. This is not the familiar point that cleanup should still execute when an assertion fails partway through — assume that is already handled, and that every case which finishes tidies up after itself. The gap is larger. There are outcomes in which *no code of yours executes at all* after some point, and there are claims that cannot be handed back at the moment you would like to. - **The run is cancelled.** Someone stops the pipeline, or a newer change supersedes it. The process is signalled and gone; identities created ten seconds ago are now orphans nobody has a record of. - **The process is killed.** A global time limit, an out-of-memory kill, an evicted machine. Teardown is not skipped by a bug — it is never scheduled. - **The message arrives late.** The case passed, teardown ran, and the notification lands thirty seconds afterwards. The evidence you would need to explain a later intermittent failure has nowhere to land. - **The claim has a cooling period.** A leased number often cannot be re-leased immediately, and a delivery side will not always forget an address on request. "Delete it now" is not available. - **The identity is shared with a later step.** A case that hands its recipient to a follow-on case cannot delete it at its own end without breaking the second one. The common shape is that per-case deletion needs the run to be alive and cooperative at exactly the moment cleanup is due. A scheduled sweep needs nothing of the sort. ## What the sweep is A sweep is a small process on a timer that walks a registry of identities and reclaims the ones that are old enough and not in use. Crucially it is **stateless with respect to runs**: it never learns which run created an identity, and it does not care whether that run succeeded, failed or vanished. Its inputs are ownership, age and status. | | Per-case deletion | Scheduled expiry | |---|---|---| | Needs the run to survive | yes | no | | Reclaims immediately | yes | only within its window | | Covers a cancelled or killed run | no | yes | | Preserves late arrivals for debugging | no, deletes the evidence | yes, for the retention window | | Requires a registry outside the run | no | yes | | Typical failure | the run dies first | the ownership filter is wrong | The two are complementary, not alternatives. Per-case deletion keeps the steady-state population small so quota is never the binding constraint; the sweep is the floor beneath it that guarantees the population is bounded even when nothing else runs. ## Choosing the window, and protecting live work Two decisions make or break a sweep. 1. **The retention window.** Too short and it deletes an identity while a notification is still in flight, or while someone is mid-investigation with a paused suite. Too long and the population grows until quota bites. Derive it from the observed tail of arrival times plus the longest case duration, with margin — not from a round number that felt safe. 2. **In-use protection.** Age alone is not sufficient, because a legitimately long case can outlive the window. The standard answer is a lease: a case marks its identity in use with an expiry it renews while running, and the sweep skips anything whose lease is live. A lease that is *not* renewed expires by itself, so a dead run cannot hold an identity forever. Then guard the blast radius, because a sweep is a process whose entire job is deletion: - Filter on an explicit ownership marker, never on "everything in this store". The moment real recipients share the store with test ones, an over-broad filter deletes customer state. - Make it idempotent and re-runnable; it will overlap with itself eventually. - Give it a dry-run mode that reports what it would remove, and read that output before widening any rule. - Have it emit counts. A sudden jump in reclaimed identities means something upstream stopped cleaning up, and that is worth knowing before quota tells you. ## Why the schedule is the honest design The deeper reason to prefer a schedule is that it stops the suite from pretending. Per-case deletion encodes a claim — *every identity I create, I destroy* — that the suite cannot actually honour, because the failure modes that leak identities are precisely the ones in which the suite is not running. A scheduled expiry states the truth instead: identities exist for a bounded time and are reclaimed by something whose survival does not depend on the run. That is a property you can rely on, and it is the reason the same pattern shows up wherever a system hands out resources it may never get back.

  • How would you choose the retention window, and what makes one too short?
    Derive it from the observed tail of arrival times plus the longest case duration, with margin. Too short and it removes an identity while a notification is still in flight or while someone is mid-investigation with a paused suite. Too long and the population grows until quota becomes the binding constraint. A round number that felt safe is not a derivation.
  • How does the sweep avoid deleting an identity a running case still needs?
    Age alone is not enough, because a legitimately long case can outlive the window. Give each identity a lease the running case renews while it works, and have the sweep skip anything whose lease is still live. A lease that stops being renewed expires by itself, so a dead run cannot hold an identity forever.
  • If a sweep exists, why keep per-case cleanup at all?
    Because the sweep is a floor, not a strategy. Per-case deletion keeps the steady-state population small so quota never binds and so a shared capacity is not exhausted by concurrent runs. Without it the retention window itself becomes a capacity limit, and you end up shortening it for the wrong reason.

A hotel does not rely on every guest handing back the key. It rotates the locks on a schedule, because the guests who never came back are exactly the ones still holding one.

saying these in an interview costs you the question

  • Assumes cleanup always runs, so nothing can leak
  • Deletes the mailbox before a late notification could arrive
  • Sweeps purely by age, removing identities a live case holds
  • Thinks a sweep replaces per-case cleanup entirely
  • Believes every claim can be released on demand
  • Runs the sweep without an explicit ownership filter