skip to content

In a persisted document registry, how do you decide when an old operation entry is safe to delete?

level: seniorimportance: nice to knowfreq 23%

answer

  1. Age is the wrong axis
  2. Ask who can still send it
  3. Silence is not death
  4. Last-seen plus client release
  5. Withdraw and alert before removing

basics

~20 s

By client usage, never by age. An entry is safe to remove once no client build still in the field can send its identifier — proven from per-entry last-seen data and the releases you still support.

solid answer

~50 s

Deleting an entry some shipped build still sends causes the same total failure as publishing a manifest late, except you cannot fix it by rolling forward — that client is already installed. So the decision needs data the registry has to collect: record on every lookup the identifier, the timestamp, whether it was found, and which client release sent it. Then retain the union of the manifests for every release at or above a supported floor, and consider an entry only when it falls below the floor **and** has gone unused for longer than the worst realistic gap between one client's uses. Intermittently connected clients matter here — a warehouse handheld can be off the network for weeks, so silence is not proof of death. Bias toward keeping, since entries are kilobytes of text, and where you do delete, withdraw first: keep serving the entry but alert on every hit before removing it.

code

pseudocode · 18 lines
pseudocode
# on every registry lookup
entry = registry.get(request.documentId)
usage.record(request.documentId,
             at      = now(),
             release = request.header("X-Client-Release"),   # who is still calling
             found   = entry != null)

# retention sweep, run weekly
for id, entry in registry:
    if entry.introducedIn >= SUPPORTED_FLOOR:        continue   # still supported
    if usage.lastSeen(id) > now() - 180.days:        continue   # someone is calling

    if entry.state == LIVE:
        entry.state = WITHDRAWN                                 # keep serving it
        alertOnEveryHit(id)
    else if entry.withdrawnFor > 60.days:
        archive(entry.text)                                     # restorable
        registry.delete(id)

go deeper

for a junior

Know that entries in a persisted document registry are not junk to be swept on a schedule: some shipped client may be the only thing that knows an identifier is still needed.

for a middle

Be able to name the data the decision needs — last-seen time per entry and the client release that sent it — and explain why a purely time-based retention window is unsafe for installed clients.

for a senior

Demonstrate the policy end to end: retain the union of manifests above a supported release floor, treat an offline fleet's silence as inconclusive, and withdraw with alerting before removing anything.

for a principal

Own the tradeoff between registry growth and upgrade discipline. Without an enforceable minimum client version the retention set only grows, and that call sits with product, not with the platform team.

## Deleting an entry is a client-breaking change with no roll-forward Everything that makes a late-published manifest painful makes a prematurely deleted entry worse. The failure is identical — an identifier the server cannot resolve, every request from that build failing, no text to fall back on — but the usual fix is unavailable: the client is already installed, already offline, already in a warehouse. Restoring the entry is the real remedy, and it only starts once somebody notices. So "when can this row go?" is not a housekeeping question. It is "can anything still send this key?", and the registry has to be built to answer it. ## Collect the evidence at lookup time The only place that knows an identifier is still alive is the lookup path. Record, on every resolution: the identifier, the timestamp, whether it was found, and — the part people leave out — **which client release sent it**. A build already knows its own version; pass it in a request header, or namespace the identifiers by release. Without it you learn that something still uses an entry but not what, which tells you nothing about whether that something can be upgraded. With that data the registry becomes a set of entries tagged by the release that introduced them and by the releases still calling them, and a retention policy can be written in terms of releases rather than dates. ## The policy Retain the union of the manifests for every client release at or above a **supported floor**, and consider an entry for removal only when it is below the floor *and* has gone unused for longer than the worst realistic gap between one client's uses. That second clause is where purely time-based retention fails. A warehouse handheld on a freight-tracking graph may scan all day for three weeks and then sit in a charging cradle in a depot for a month; a driver's device syncs when it finds signal. Thirty days of silence from such a fleet is not evidence of death, it is evidence of a quiet month. Any window you pick has to be longer than the longest legitimate quiet period, which for installed software is measured in months. The floor itself only means something if the client enforces it. An app that checks a minimum supported build at launch and blocks below it makes the floor real; without that, the floor is whatever the slowest user in the field does, and the retained set only grows. That is a product decision as much as a platform one, and naming it is what separates a senior answer from a principal one: you are choosing between registry growth and upgrade friction, and calling the choice technical does not make it go away. ## Delete in two stages Withdraw before you remove. Mark the entry withdrawn but keep serving it, and alert on every hit with the release that sent it. A quiet period in that state is far better evidence than a quiet period in a usage log, because it is measured against a decision you have already made and are actively watching. Hard-delete only once a withdrawn entry has gone untouched across a full sync cycle for the slowest client you support. Do it in batches you can undo. Keeping the withdrawn text somewhere restorable means a mistake costs a re-publish rather than an emergency rebuild of an old client. ## Bias toward keeping — but know the real cost A registry entry is a few kilobytes of query text and an index row. Eight thousand four hundred of them, accumulated over a couple of years of releases, is a rounding error in storage and no measurable difference in lookup time. Anyone who justifies deletion by registry size has not measured it. The costs that are real are subtler. Retained documents **pin the schema surface they select**: a field no current build asks for still has live callers for as long as an old build's document is registered and reachable, so the retention set quietly constrains what you can change. And per-operation dashboards fill with entries nothing shipped uses, which makes usage data harder to read at exactly the moment you want to read it. Both argue for deleting eventually and deliberately — with a supported floor, evidence from the lookup path, and a withdrawn state in between — rather than either purging on a schedule or never touching the registry at all. "We never delete, and here is why" is a defensible interim position. "We delete after 90 days", with no idea who is still calling, is the answer that gets marked down.

  • What makes a supported release floor enforceable rather than aspirational for an installed app?
    A version gate in the client itself: the app checks a minimum supported build at launch and blocks below it with an upgrade prompt. Without that, the floor is whatever the slowest user in the field does and the retained set only ever grows. It also gives the unknown-identifier error somewhere useful to land — map it to the same upgrade prompt rather than a generic failure message.
  • Is there any real cost to never deleting anything from the registry?
    Storage is not it — several thousand entries of query text is trivial to hold and to look up. The real costs are that retained documents pin the schema surface they select, so a field no current build asks for still has live callers, and that per-operation dashboards fill with entries nothing shipped uses, which makes the usage data harder to read.
  • How do you verify a deletion is safe before you make it?
    Withdraw the entry: mark it retired but keep serving it, and alert on every hit with the release that sent it. A quiet period in that state is much stronger evidence than a quiet period in a usage log, because it is measured against a decision you are actively watching. Hard-delete only after a withdrawn entry has gone untouched across a full sync cycle for your slowest client.

It is closer to retiring a key from a hotel's lock system than to clearing out old files: nothing goes wrong until the one guest still carrying that key comes back to the door.

saying these in an interview costs you the question

  • Deletes entries on a fixed age window
  • Assumes an unused entry has no live caller
  • Thinks a stale client can re-register the document text
  • Keeps only the current release's manifest
  • Treats registry storage size as the main cost
  • Ignores offline clients that sync in bursts

context