When is a call safe for a retry decorator to repeat, and how do you make an unsafe one safe?
answer
- Repeating is not always harmless
- A timeout tells you nothing at all
- The second call must be recognisable
- One key per logical operation, minted once
- Server-side dedupe or a conditional write
basics
~20 sOnly repeat calls whose second run leaves the same end state: reads, key-addressed overwrites, upserts. Appends, increments and sends are not. Make one safe by minting an operation key once, above the retry loop, so the far side can recognise the repeat.
solid answer
~50 sA call is safe to retry when running it twice leaves the world in the same state as running it once. Reads and key-addressed overwrites qualify; an append, a counter increment, a charge or a notification does not. The trap is that a failure rarely tells you which happened: a connection refused proves the request never landed, but a read timeout means the far side may well have processed it and only the reply was lost, so the retry duplicates work. The fix is a deduplication key — a `uuid.uuid4()` string minted **once per logical operation, above the retry wrapper** — that the far side records and recognises, or a conditional write guarded by an expected version. If the key is generated inside the wrapped function it changes on every attempt and dedupe is impossible. When you cannot make the call idempotent, do not retry it: fail fast and surface the ambiguity.
code
python · 39 linesimport functools
import time
import uuid
def retry(attempts=3, exceptions=(TimeoutError,)):
def decorate(func):
@functools.wraps(func)
def wrapper(*args, **kwargs):
for attempt in range(1, attempts + 1):
try:
return func(*args, **kwargs)
except exceptions:
if attempt == attempts:
raise
time.sleep(0.001)
return wrapper
return decorate
committed = {}
received = []
@retry()
def append_batch(chat_id, messages, operation_id):
received.append(operation_id)
if len(received) == 1:
raise TimeoutError("no acknowledgement from the archive service")
if operation_id in committed: # the far side recognises the repeat
return committed[operation_id]
committed[operation_id] = len(messages)
return committed[operation_id]
operation_id = str(uuid.uuid4()) # once per logical operation
print(append_batch("chat-42", ["a", "b", "c"], operation_id))
print("attempts seen:", len(received), "distinct keys:", len(set(received)))
print("batches committed:", len(committed))go deeper
Recall that repeating a call can repeat its effect: asking for data twice is harmless, sending the same message twice is not. Be able to sort a few example operations into safe and unsafe to repeat.
Explain idempotency as 'same end state after two runs', and give the standard fix: an operation identifier sent with every attempt so the far side can recognise and ignore a repeat. Know why the identifier must not be generated per attempt.
Demonstrate that you reason from what the failure proves — refused versus timed out versus rejected — and that you design the safe repeat yourself: deterministic keys, conditional writes on an expected version, or a deliberate decision not to retry at all.
Own where the exactly-once illusion is paid for in your architecture, which call paths are allowed retries at all, and how you make 'this operation is safe to repeat' a documented property of an interface rather than folklore in each caller.
## The property, stated precisely An operation is idempotent when performing it more than once leaves the same observable state as performing it once. Note what that does *not* say: it does not say the operation is free, and it does not say every response is identical. A read is idempotent but still costs the far side work. A delete is idempotent in state even though the second call may report "not found". An append is not idempotent at all — two calls mean two records. A retry decorator is a machine for performing an operation more than once. So the decorator's correctness is not a property of the decorator; it is a property of what you point it at. ## The ambiguous failure The interesting cases are not the clean ones. Group failures by what they prove: - **Proof the work never started**: the connection was refused, DNS did not resolve, the request was rejected before the body was read. Retrying is safe even for a non-idempotent operation, because nothing happened. - **Proof it will never succeed**: a validation rejection, an authorisation failure, a malformed payload. Retrying is pointless. - **No proof either way**: a read timeout, a connection reset mid-flight, a crash after the far side committed but before it answered. This is the dangerous middle, and it is also the most common failure in practice. A timeout is not evidence of failure. It is evidence that you do not know. Treating "no answer" as "did not happen" is the single most expensive misconception in this area, because it is right most of the time and catastrophic when it is not. ## Making an unsafe call safe Four techniques, roughly in order of preference: **Make it naturally idempotent.** Write to a deterministic address rather than appending: store the chat-transcript batch under a key derived from its content or its range, so the second write overwrites the first with identical bytes. Nothing needs to remember anything. **Carry a deduplication key.** Mint an identifier once per *logical* operation — `uuid.uuid4()` is fine — and send it with every attempt. The far side records processed keys and returns the original result for a repeat. The load-bearing detail is *where the key is created*. If it is minted inside the function the decorator wraps, every attempt carries a fresh key and every attempt is a new operation as far as the far side is concerned; the dedupe is worse than useless because it looks like it is working. The key must be created above the retry wrapper, or passed in as an argument, so all attempts of one logical operation share it. **Guard with a condition.** Send the version or state you expect and let the write fail if it no longer matches. The second attempt of a retried write sees that the state has already advanced and declines instead of double-applying. **Do not retry.** Perfectly respectable. Fail fast, hand the ambiguity to a caller that can reconcile it, or move the work onto a durable queue where an at-least-once consumer is already designed to see duplicates. ## The boundary bug that hides here Cursor-based work is where this goes wrong quietly. Suppose an archiver reads transcript messages from a cursor and asks for "everything after offset N". A retried batch that already committed on the far side leaves your local cursor behind by one batch; if the range is treated as inclusive on one side and exclusive on the other, each retry either re-archives one message or skips one, and the corruption is exactly one record wide. It survives review because a one-record discrepancy per retry looks like noise until someone counts. Two defences: make the range half-open and say so in the name, and have the far side report the offset it actually committed rather than trusting the client's arithmetic after an ambiguous failure. ## Reads are safe, not free "Idempotent" is not permission to retry without thought. Retrying a read against a dependency that is already overloaded adds load precisely when it can least afford it, and an expensive read retried three times by many callers can be the thing that keeps the dependency down. Idempotency governs *correctness*; the delay schedule and attempt bound govern *load*. You need both. ## What to say in an interview Lead with the property and the ambiguity: "safe to repeat means the same end state, and the hard part is that a timeout doesn't tell me whether it ran." Then name the fix you would actually build — a key minted once per operation, recognised on the far side — and say explicitly where in the code the key is created. Finish with the honest option: some calls should simply not be wrapped in a retry.
- A write times out. What do you actually know about whether it happened?Nothing. The request may have been lost on the way out, processed and lost on the way back, or still be in flight. That is different from a refused connection, which proves the work never started, and from a validation rejection, which proves it never will. Because the outcome is unknown, a retry of a non-idempotent write is a coin flip between a lost operation and a duplicated one — which is exactly why the dedupe key exists.
- Where must the deduplication key be created for a retry decorator to work?Above the retry wrapper — in the caller, or passed in as an argument — so every attempt of one logical operation carries the same value. If the wrapped function mints it, each attempt looks like a distinct operation to the far side and the dedupe silently does nothing while appearing to be correct. A useful review habit is to check that the identifier is an argument to the decorated function, not a local inside it.
- Is retrying a read always safe because reads are idempotent?Correct on state, wrong on load. A repeated read leaves nothing changed, but it still consumes the dependency's capacity at the moment it is least able to spare it, and an expensive read retried by many callers can keep an overloaded service down. Idempotency licenses the retry for correctness; the delay schedule, the ceiling and the total deadline are what keep it from being an attack.
- How do you keep a cursor-based archiver from duplicating or losing a record when a batch is retried?Define the range half-open and name it that way, so an inclusive/exclusive mix-up cannot silently shift a boundary by one. Then have the far side report the offset it actually committed and advance the local cursor from that report rather than from local arithmetic, so an ambiguous failure resynchronises on the next attempt instead of drifting by one record each time.
saying these in an interview costs you the question
- Assumes a timeout means the request did not happen
- Retries an append with no deduplication key
- Mints the idempotency key inside the wrapped function
- Treats idempotent as free and ignores the added load
- Plans to reconcile duplicates manually afterwards
- Says every retry is fine as long as one attempt succeeds