A shared in-memory tier keeps growing and holds many entries with no lifetime attached — what audit and what contract do you put in place?
answer
- growth with no failed operation
- the store keeps state, not history
- share per namespace, trended
- intentional needs an owner and a ceiling
- enforcement belongs in the write path
basics
~20 sMeasure the share of entries with no lifetime attached per namespace and trend it, rather than counting once. Then make it a contract: each namespace declares its posture, permanent ones need an owner and a deletion path.
solid answer
~50 sAn audit alone cannot fix this, because the store keeps only the current lifetime state and not the history of what set or cleared it — you cannot tell an entry born without a lifetime from one that lost it in a later write. So the audit's job is to show shape, not blame: sample the keyspace rather than enumerate it on a live tier, group by namespace prefix, and report **the share** of entries with no lifetime attached per group, trended over releases. An absolute count means nothing on its own. The contract that goes with it is that every namespace states whether its entries carry a lifetime, the default being that they do; a namespace holding permanent entries needs a named owner, a bounded cardinality and a deletion path. Enforcement belongs in the shared client, because an audit only ever finds these afterwards.
go deeper
Recall that an entry with no lifetime attached stays until something deletes it explicitly, and that a tier can grow steadily without any operation ever failing or any error being logged.
Explain why reading the lifetime state per entry gives the current picture only, and why grouping by namespace prefix is the only way to turn that picture into something a team can act on.
Show the measurement you would actually build — sampled, grouped, reported as a share, trended — and the ceiling consequence that makes permanent entries an availability problem rather than an untidiness problem.
Own the contract. Decide what consuming teams may write without a lifetime, where that is enforced, what a permanent namespace must declare, and what you do when a team's intentional namespace has unbounded cardinality.
## Why this is a contract question and not a cleanup A tier that grows with no failed operation anywhere is not a bug you fix once. Each permanent entry arrived because some write path did not carry a lifetime, and those write paths belong to teams that do not read your memory graph. Deleting the entries buys a few weeks; the same paths refill the tier. The durable answer is a contract about what may be written without a lifetime, plus a measurement that shows whether the contract is holding. ## What the audit can and cannot tell you The store holds the *current* lifetime state of an entry. It does not hold the history of how that state was reached, so these two entries are indistinguishable: - one created by a path that never attached a lifetime at all; - one created with a lifetime and later rewritten by a path that dropped it. That is not a gap you can close by looking harder. It means the audit tells you *where* permanent entries accumulate, and the write path is the only place the cause lives. Say this out loud before proposing a tool, because a plan that promises to distinguish the two is promising something the store cannot supply. ## Measuring the right number - **Sample, do not enumerate.** On a live tier of any size, a full pass is a cost you pay against the same resource you are investigating. A sample large enough to be stable per namespace is enough for a trend. - **Group by namespace prefix**, because the prefix is the only ownership signal a keyspace with no schema carries. A number for the tier as a whole tells nobody what to do. - **Report a share, not a count.** "Four hundred thousand permanent entries" is meaningless without the total; "this namespace went from two per cent to thirty-one per cent permanent across the last two releases" names a release and a team. - **Trend it.** The signal is the change, and it is what connects a growth curve to the change that caused it. - **Pair it with the read that distinguishes the two non-duration answers** — no lifetime attached, and no such entry — so the audit does not count names that are already gone. ## Deciding which permanent entries are intentional Not every permanent entry is a defect. Sorting them is the part that needs judgment: | category | what it looks like | what the contract requires | |---|---|---| | intentional and bounded | a lookup table loaded at deploy, a small registry, a feature-decision set | a named owner, a known maximum size, a path that replaces it | | intentional but unbounded | a per-user or per-object entry deliberately kept forever | reclassify: either give it a lifetime or move it to a durable engine | | accidental | anything the owning team is surprised to see | fix the write path, then delete | The middle row is the one teams argue about, and the argument is usually settled by remembering what this tier is: a volatile component that may be emptied by a restart. Data that must not be lost has no business being the only copy here, and data that may be lost does not need to be permanent. ## What the contract says 1. **A write carries a lifetime by default.** Omitting one is a deliberate act, not the easy path. 2. **Each namespace declares its posture** — timed or permanent — in a registry a human can read, alongside its owner and its expected cardinality. 3. **A permanent namespace names how entries leave.** If the answer is "they do not", the size ceiling has to be stated instead, and monitored. 4. **Enforcement lives in the shared client**, which can require a lifetime argument or an explicit opt-out, because that is the only place that sees every write before it happens. 5. **The audit reports against the registry**, so a namespace that is permanent by declaration is not noise and a namespace that is permanent by accident stands out. ## The consequence at the ceiling One operational fact makes the case for the contract without any appeal to tidiness. Where the tier is configured to remove entries under memory pressure but only considers entries that carry a lifetime, permanent entries are simply not candidates: the pressure falls entirely on the timed entries, which are the ones the application actually depends on being there. A tier can therefore be thrashing, or refusing writes, while a growing fraction of its memory is held by entries nothing will ever choose to remove. That is the moment a slow accumulation stops being a graph and becomes an incident. ## How to answer Lead with the contract, not the script. State plainly that the store cannot tell you how an entry lost its lifetime, propose a sampled per-namespace share that is trended, give the three categories and what each one owes, put enforcement in the write path, and close on the ceiling consequence — it is the argument that moves an organisation.
- A team insists their permanent namespace is intentional. What do you ask them for?An owner, a maximum cardinality with the reasoning behind it, and the path by which an entry leaves — replacement at deploy, explicit deletion, or a bounded rewrite. If the maximum grows with users or objects, it is not bounded, and the namespace either takes a lifetime or moves to an engine that is meant to hold data permanently.
- Why is the total count of permanent entries the wrong headline number?Because it moves with the size of the tier and with traffic, so it cannot be compared across weeks, and it names no owner. A share per namespace, trended, does both: it stays meaningful as the tier grows and it points at the prefix, and therefore the team, whose behaviour changed.
- Can the audit run continuously, or should it be a periodic job?Periodic and sampled. The measurement competes for the same resource it is measuring, and the signal is a trend across releases rather than a live number, so a run per deploy or per day is sufficient. Alert on a change in a namespace's share, not on a threshold crossing for the tier.
saying these in an interview costs you the question
- Proposes deleting the entries and calls it fixed
- Claims the store can show which entries lost a lifetime and when
- Reports one absolute count for the whole tier
- Enumerates the entire keyspace on a live tier to measure it
- Treats every permanent entry as a defect regardless of owner
- Assumes memory pressure will eventually remove permanent entries anyway