skip to content

Your deduplication records sit on a tier that removes entries under memory pressure and the same payout arrives twice, so what happens and where should the guard live?

level: seniorimportance: should knowfreq 52%

answer

  1. repeated, not lost
  2. degrades to at-least-once
  3. state the stake before judging
  4. entries leave for unrelated reasons
  5. guarantee where the effect is recorded

basics

~20 s

The second arrival finds nothing, is treated as new work, and the payout goes out twice: once-only handling degrades to at-least-once. A record the tier can remove is an optimisation, so a payout that must never repeat needs a durable uniqueness rule.

solid answer

~50 s

With the record removed, the repeat is indistinguishable from a first arrival, so the payout goes out again - the work is not lost, it is repeated, which is what **degrades to at-least-once** means. The question that decides the design is whether this record is an optimisation or a guarantee. It can only be an optimisation, because entries leave this tier for reasons unconnected to your request: memory pressure driven by other people's keys, a restart, and - depending on the store and how it is configured - a promoted copy that never received the write. If a duplicate is merely wasteful, that is entirely enough. If the payout must never repeat, the guard belongs in the durable store that records the payout and can declare the request identifier unique, with the volatile record kept in front as a cheap filter.

go deeper

for a junior

Get the direction right: a missing record means the work runs again, not that the work is lost. Duplicates are what you go looking for afterwards.

for a middle

Explain why the entry can vanish for reasons unrelated to your request, and connect that to the cost of the duplicate it was supposed to prevent.

for a senior

State the stake first, then place the guard. Where the effect must not repeat, the uniqueness rule belongs with the durable record of the effect, with the volatile entry kept as a filter.

for a principal

Put a number on the degradation - how often, how many, what each duplicate costs - and decide whether it is a priced degradation or an unbounded liability wearing a design's clothes.

## What happens, precisely The tier reaches its memory ceiling and makes room by removing existing entries. Some of those are deduplication records, chosen for reasons that have nothing to do with your payout - the pressure came from another team's keys during their peak. Later, the same payout request arrives again, because the original sender never learned the outcome. The service looks for the record, finds nothing, and correctly concludes - on the evidence available to it - that this is new work. It sends the payout a second time. Name the outcome carefully, because the name drives the response. Nothing was lost, no request went unprocessed, and there is nothing to replay. One thing happened twice. Once-only handling has **degraded to at-least-once**. After such an incident you go looking for duplicate effects to reverse, not for missing work to redo. ## The question that decides everything: optimisation or guarantee Asked without a stated stake, *is this record enough?* has no answer. With the stake stated it has an easy one. | What a duplicate costs | What the record has to be | Where the guard belongs | |---|---|---| | Wasted computation, a repeated import, a second identical write of the same value | an optimisation | the volatile record alone is fine | | A duplicate notification: annoying, reversible, cheap to apologise for | an optimisation with monitoring | the volatile record, plus a count of how often it misses | | A second payout, a second charge, a second shipment | a guarantee | a uniqueness rule in the durable store that records the effect | The common failure in an interview is to answer the third row with the first row's design and not notice the difference. ## Why this tier cannot carry the guarantee Three causes take the record away independently of anything your code does, and it is worth being exact about which are universal and which depend on the store: - **Removal under memory pressure.** This is only in play if the tier is set up to make room by removing entries; some stores instead refuse further writes when they reach their ceiling, which protects existing records but breaks new ones. Either way, you are not in control of the outcome for one key. - **A restart.** Some stores in this class keep nothing at all; others reload a copy from disk that lags the newest writes, so the records for work in flight are the likeliest casualties. - **A promoted copy that never received the write.** This is not a property of the class. Several stores acknowledge a write only once a copy holds it, and some let a caller ask for that per call; on those, this cause is closed, at a latency cost. Where the write is acknowledged before any copy has it, the record can simply not exist on the node that takes over. None of these is a bug. They are the terms on which a volatile tier is fast, and they are why an entry here is evidence rather than proof. ## The two-layer arrangement When a duplicate is intolerable, the arrangement that actually holds is: 1. **The durable guard.** The store that records the payout declares the request identifier unique, and the identifier is written in the same durable operation that records the effect. A second attempt then fails when it tries to record, rather than being waved through by a missing marker. That store is somebody else's subject - what matters here is that the guarantee lives somewhere that cannot silently forget. 2. **The volatile front door.** The deduplication record stays, sized to the common repeat, and absorbs the overwhelming majority of duplicates without a durable round trip. It is now explicitly an optimisation, and saying so out loud is the point: nobody later mistakes it for the thing keeping customers from being paid twice. This also reframes a miss. When the volatile record is gone, the duplicate is not a second payout, it is one extra durable lookup that fails on the uniqueness rule. The failure mode of the optimisation is cost, not correctness. ## Pricing the degradation The senior move is to put a number on it rather than argue in the abstract. Roughly: how often does the tier remove entries or restart, how many requests are inside their retry window when it happens, what fraction of those are actually retried, and what does one duplicate cost - in money, in support time, in regulatory exposure? A design where the answer is *a handful of duplicate notifications a quarter* is a priced degradation and can ship. A design where the answer is *some number of duplicate payouts we cannot bound* is not a degradation, it is an outage that has not been noticed yet, and the record must stop being the guard. That framing is also the honest one for the failure you cannot prevent: on this tier the goal is never that the entry can never disappear, it is that the disappearance has a named, bounded cost.

  • If the durable store already enforces uniqueness, why keep the volatile record at all?
    Because it keeps the overwhelming majority of duplicates off the durable path, which matters when retries arrive in bursts and the durable store is the expensive resource. Its failure mode is now cost rather than correctness: a miss means one extra durable attempt that is rejected, not a second payout.
  • Would replicating the record to a second node make it a guarantee?
    It narrows one cause, not all of them. A copy does nothing about a deadline passing, about removal under memory pressure, or about a restart that empties both. It only helps against a failover, and only where the write is acknowledged after a copy holds it - which several stores offer and others do not.
  • The duplicate is a repeated internal recomputation that produces the same value. What now?
    Then the volatile record is entirely sufficient and a durable guard would be overpriced. The stake decides the design: an occasional repeat that wastes some processing and writes the same value again is a cost, not an incident, and the cheap optimisation is the right size of answer.

saying these in an interview costs you the question

  • Treats an entry on a volatile tier as proof the payout cannot repeat.
  • Says data was lost, when the work was repeated rather than dropped.
  • Answers is this enough without asking what a duplicate actually costs.
  • Tries to fix it by making the lifetime longer and stopping there.
  • Assumes every store in this class acknowledges a write only after a copy holds it.
  • Removes the volatile record entirely once a durable guard exists.