skip to content

How long must a deduplication record live, and why is the first retry delay the wrong number to size it against?

level: middleimportance: must knowfreq 64%

answer

  1. longest repeat, not the first one
  2. enumerate senders, take the maximum
  3. a configured lifetime is a ceiling
  4. memory pressure shortens retention silently
  5. durable home when the span is long

basics

~20 s

Size a deduplication record against the retry window - the longest span over which the same work can arrive again, including a stalled consumer, a nightly reconciliation resend or an operator replay - never against the sender's first retry delay.

solid answer

~40 s

The record has to outlive the **retry window**: the longest span over which the same unit of work can arrive again. The first retry delay is a few seconds, but the same work can also come back from a consumer that stalled for hours, a reconciliation job that resends anything with no recorded outcome, or a person replaying a batch days later. Enumerate every party that can resend, take each one's last possible attempt, take the maximum, and add margin. The second half of the answer is that the deadline you configure is a **ceiling on retention, not a floor**: removal under memory pressure and a restart both shorten it silently, so if you cannot say what the tier does at its memory ceiling, you cannot state your real retention either.

go deeper

for a junior

Remember the direction of the rule: the record has to last longer than the longest gap before the same work can come back, not longer than the first retry.

for a middle

Derive the number out loud - list every party that can resend, take the maximum of their last possible attempts, add margin - instead of naming a habitual duration.

for a senior

Say that the configured deadline is a ceiling on retention and name what shortens it, then state what you would actually measure to know your true retention.

for a principal

Weigh a month of per-request keys in memory against a durable uniqueness rule with a short volatile front door, and say which layer is carrying the guarantee.

## The rule, stated once A deduplication record must outlive the **retry window** - the longest span over which the same work can arrive again. Not the average gap between duplicates, not the request timeout, and not the sender's first backoff step. The last moment a duplicate can plausibly appear is the number the lifetime has to clear. The common failure is sizing against the wrong end of the distribution. Somebody sees that the client retries after two seconds, rounds up generously to five minutes, and ships it. Five minutes is an enormous multiple of the number they looked at, and it is still far too small, because the duplicate that actually costs money almost never comes from the client's own immediate retry. ## Where repeats actually come from | Source of the repeat | Who triggers it | Span it implies | |---|---|---| | The sender's own retry schedule | the client library, automatically | seconds to a few minutes, to its final attempt | | A consumer that stalled and later resumed | a stuck worker, a paused deployment | as long as the stall, often hours | | A scheduled reconciliation that resends anything with no recorded outcome | a nightly or hourly job | up to that job's interval, commonly a day | | A person: support replaying a stuck batch, a customer pressing submit again | a human | days, bounded only by policy | | A partner re-driving from their side after their own incident | someone else's runbook | whatever their runbook says | Deriving the number is mechanical once the sources are listed: 1. List every party that can send this work - including the ones inside your own company. 2. For each, write down the last moment it can send it again. 3. Take the maximum, not the median. 4. Add margin for clock differences between the parties and for the replay nobody documented. 5. Compare that number against what the tier will actually retain, and if the tier cannot promise it, decide deliberately where the real guard lives. Step five is the one candidates skip, and it is the interesting one. ## A configured lifetime is a ceiling, not a floor Attaching a deadline to the entry says when it stops being valid. It does not promise the entry will live that long. Two things shorten it without appearing in any code review: - **Removal under memory pressure.** A tier that is configured to make room by removing existing entries when it reaches its memory ceiling will remove some of yours. On a shared tier that is driven by other people's keys and other people's traffic peaks, so the retention of your records is coupled to workloads you do not own. Stores differ here in an important way: some refuse further writes at the ceiling instead of removing anything, and on those your records are safe but your writes start failing - a different failure that you would rather find out about in an interview answer than in production. - **A restart.** Some stores in this class keep nothing across one. Others reload a copy from disk that typically lags the most recent writes, which means the records for work currently in flight are the least likely to survive. So effective retention is at most the lifetime you configured, and your sizing is a request rather than a contract. ## The other direction: no deadline at all The mirror-image mistake is leaving the record permanent. These entries are minted one per request, so the key count grows with traffic forever, and on this tier the number of keys is usually what hurts before the bytes do. Permanent records also make removal under memory pressure likelier for every other workload sharing the tier, and they are among the hardest entries to clean up later, because from the outside you cannot tell which ones are still inside somebody's retry window. ## When the honest window is longer than you will pay for Suppose the honest retry window is thirty days, because a partner can re-drive a month of work after an incident, and you will not hold thirty days of per-request keys in memory. The answer is not to lie about the window. It is to split the job: - Put the durable guard where the effect is recorded - the store that writes the payout can declare the request identifier unique, so a second attempt fails when it tries to record the effect. - Keep a **short** volatile record sized to the common repeat, the sender's own retries, so the overwhelming majority of duplicates are recognised without touching the durable store at all. That arrangement is honest about what each layer is for: the volatile record is an optimisation with a stated hit rate, and the guarantee lives somewhere that cannot silently forget.

  • Two senders can each replay the same work, on very different schedules. How does that change the number?
    The record has one lifetime, so it must cover the slowest replayer, not the average of the two. Take the maximum. If that number is uncomfortably large, the cheaper fix is often to shorten the slow replayer's window by agreement rather than to hold every record in memory for its benefit.
  • What does it cost to make the lifetime far longer than necessary?
    Memory held for records nobody will ever look up again. These keys are minted one per request, so key count grows with traffic, and on a tier that makes room by removing entries your surplus records raise the pressure for everyone sharing it. Over-long is safer than too short, but it is not free.
  • How would you find out what retention the tier actually delivered?
    Not from the lifetime you configured. Measure it: how often entries leave before their deadline, and whether the process restarted. If you cannot measure either, treat the retention as unknown and put a guard that cannot forget behind the volatile record.

saying these in an interview costs you the question

  • Picks a round number such as one hour because it looks tidy.
  • Sizes the record against the first retry delay rather than the last possible attempt.
  • Assumes the entry survives to its deadline whatever the tier is doing.
  • Leaves the record permanent, growing the key count with every request forever.
  • Forgets that an operator or a partner can resend the same work days later.