skip to content

An expiry rule can measure its interval from an entry's last write or its last read. How do the two differ, and which suits deduplication?

level: middleimportance: must knowfreq 56%

answer

  1. two clocks, not one
  2. reads refresh or they do not
  3. a read is a duplicate arriving
  4. hot keys become immortal
  5. interval from the redelivery tail

basics

~20 s

Last-write expiry gives an entry a fixed lifetime from its last store or update; last-access expiry restarts the clock on reads too, so hot entries never leave. Deduplication wants last-write: a read there means a duplicate arrived.

solid answer

~50 s

Entry expiry is a removal rule that drops an entry from the retained set — everything the job still holds between records — a fixed interval after a chosen event. With a **last-write** clock, the interval runs from the moment the entry was stored or updated, so every entry has a bounded lifetime whatever the reads do. With a **last-access** clock, a read also refreshes the entry, so hot entries survive indefinitely and only genuinely cold ones fall out; that is cache behaviour, right for a lookup entry you want kept while in use. For a deduplication set it is exactly wrong: a read means a duplicate for that id has just arrived, so a poison retry loop refreshes its own entry forever and the one key you most wanted bounded becomes immortal. Engines differ in which clocks they offer, and in whether an in-place update counts as a write.

go deeper

for a junior

Know that an expiry rule drops an entry after an interval, and that the interval can run from the last write or from the last touch. Remember that reading an entry may or may not restart that clock.

for a middle

Explain why a last-access clock makes a repeatedly hit deduplication entry immortal, and state the clock source — wall clock or the job's own progress assertion — in every claim you make about when removal happens.

for a senior

Demonstrate that you derive the interval from the longest redelivery or replay horizon, and that you write the resulting weaker guarantee into the contract the consumers of the output rely on.

for a principal

Own the interval as an organisational default with an owner: a number chosen once, documented alongside the guarantee it implies, and revisited when the source's redelivery behaviour changes.

## What an expiry rule actually is **Entry expiry** is a removal rule that drops an entry from **the retained set** — everything a running job is still holding between records — a fixed interval after some chosen event. It is the cheapest of the removal rules to adopt because it needs no per-key bookkeeping from the author: you state an interval and a clock, and the runtime removes entries that have outlived it. Everything interesting is in the choice of clock. ## The two clocks | | last write | last access | |---|---|---| | the clock restarts on | storing or updating the entry | storing or updating it **and** reading it | | an entry read constantly | still expires on schedule | never expires while the reads continue | | an entry never touched again | expires one interval after its last write | expires one interval after its last touch | | what it models | "this fact is stale after N" | "keep this while it is in use" | | natural fit | deduplication entries, results with a staleness horizon, anything with a business retention rule | lookup-style values, a cached lookup copy, an accumulator you want kept while a key is active | ## Why deduplication specifically wants last write Read the deduplication logic against the last-access clock and the failure is immediate. The step looks the id up; a hit means a duplicate has arrived; the lookup itself counts as an access; the entry's clock restarts. So the ids that keep arriving — precisely the retried, replayed, looping ones — refresh their own entries forever, while the well-behaved ids that appeared once expire on time. The rule bounds the part of the set that was never the problem and leaves the pathological part permanently resident. With a last-write clock, the guarantee the job offers becomes a statement you can write down: **no duplicates within the expiry interval**. That is a weaker guarantee than "no duplicates ever", and the correct response is to choose the interval from the longest realistic redelivery or replay horizon rather than from how much memory happens to be free. ## Traps in both clocks - **Does an update count as a write?** For an accumulator that is incremented on every record, a last-write clock behaves almost like a last-access one, because nearly every touch is a write. Whether an in-place update refreshes the clock is a real difference between runtimes, and the answer changes what your rule bounds. - **Which clock source drives the interval?** The interval may be measured against the worker's own wall clock, or against the job's own assertion of how far the input has progressed. The second is reproducible across a replay of history and will not remove entries while the input is quiet; the first keeps removing regardless of whether records are flowing and gives different results when you reprocess the same data. - **Removal may be lazy.** Some stores remove the entry the moment it is due; others mark it and reclaim the space only when the entry is next touched or during the state store's background file merge. With lazy removal, a set of entries that are never touched again can remain on disk long after they logically expired, which is a nasty surprise if you added expiry hoping to shrink a snapshot this week. - **The interval is usually not free to change later.** Changing it changes what the job holds, so the first run after the change removes a large batch of entries at once, which is a load spike, and shortens the guarantee you had been giving consumers. - **Expiry is not a bug filter.** An entry removed early is a correctness change, not a cleanup: the next record for that key behaves as though the key were new. ## What varies across engines Some runtimes expose expiry as a declarative property of the stored value with a choice of clock; others offer no declarative form at all and expect the author to implement removal with a **per-key scheduled callback** — a wake-up registered against one key and one moment. Where a continuous input is processed as repeated small finite jobs, removal is evaluated once per chunk rather than once per record, so an entry can outlive its due moment by up to one chunk. In the two-phase disk-to-disk batch model, which keeps nothing between runs, there is no entry to expire. A candidate who states which model they are describing before making a claim is doing the thing this subject rewards. ## Choosing the interval Derive it from the tail of the real world, then round up: the longest redelivery window the source can produce, the longest outage after which you will replay, the longest gap between two events that must still be treated as the same unit of activity. If that number makes the retained set too large to snapshot, the honest conclusion is that the deduplication horizon must be argued down with the people who depend on it, or moved out of the job — not that the interval should be quietly set to something the memory likes.

  • Why might expiry measured against the job's own progress assertion be preferred over wall clock?
    Because it makes removal a function of the data rather than of when you happened to run. Reprocessing the same history produces the same removals, so a rerun matches the original output. The cost is that a quiet input stops advancing the assertion, so entries stop being removed while nothing arrives — the reason that clock stalling is a monitored condition in its own right.
  • What does the expiry interval do to the guarantee the job offers its consumers?
    It converts "exactly once per id" into "no duplicate within the interval". That is a contract worth writing down, because a consumer that assumed the stronger version will silently double-count a record redelivered after the interval. Pick the interval from the longest redelivery or replay horizon you will actually accept, and tell the consumer what it is.

saying these in an interview costs you the question

  • Picks last-access expiry for deduplication, so a retried id never expires.
  • Assumes any expiry rule guarantees the set stops growing, whatever the clock.
  • Thinks an expired entry always frees its bytes the moment it is due.
  • Sets the interval from available memory rather than the redelivery horizon.
  • Believes expiry is cleanup with no effect on the job's output.