Where should retry logic live in a Python codebase, and when is a hand-rolled decorator the wrong choice?
answer
- One policy, or one policy per layer
- Attempts multiply when wrappers nest
- Who can still tell transient from permanent
- The deadline belongs to the caller
- An injected sleep makes it testable
basics
~20 sPut one retry layer on each call path, in the thinnest adapter that still understands the transport's exceptions. Nested decorators multiply attempts. Hand-rolled is fine for one policy; a maintained library earns its place once you need async variants, deadline propagation and telemetry.
solid answer
~50 sDecide the owning layer first. Retries belong where the code can still tell a transient failure from a permanent one — the adapter that speaks the transport — because once an exception has been translated into a domain type, the information needed to classify it is often gone. Then enforce **one** layer: a decorator on a helper that is itself called by a decorated function turns three attempts into nine, and the backoff windows compound into a latency nobody budgeted for. Make attempts, caps and the deadline configuration rather than literals, inject the sleep so tests assert a schedule without spending real seconds, and emit one warning per retried attempt plus a counter, so retry volume is observable rather than folklore. Hand-rolling is right when the policy is a dozen lines used in one place; reach for a maintained library when you need sync and async variants, deadline propagation, per-call-site policies and tested jitter — and accept the dependency's upgrade cost as part of the trade.
code
python · 37 linesimport functools
calls = 0
def retry(attempts):
def decorate(func):
@functools.wraps(func)
def wrapper(*args, **kwargs):
for attempt in range(1, attempts + 1):
try:
return func(*args, **kwargs)
except TimeoutError:
if attempt == attempts:
raise
return wrapper
return decorate
@retry(attempts=3)
def fetch_page():
global calls
calls += 1
raise TimeoutError("archive service did not answer")
@retry(attempts=3)
def fetch_transcript():
return fetch_page()
try:
fetch_transcript()
except TimeoutError:
pass
print("calls made:", calls)go deeper
You are not expected to design this, but know that retry code should exist in one place rather than be sprinkled around, and that a retry hides a failure from the caller while it is happening.
Be able to explain why two nested retry wrappers multiply the calls made, and why attempts, delays and the exception set should be parameters rather than literals baked into each call site.
Show that you place the retry in the layer that can still classify the failure, that you bound the path by a deadline you pass down, and that you can test the policy without real sleeps and observe retry volume in production.
Own the standard: a single house policy and exception classification, rules about which call paths may retry at all, the build-versus-adopt call with its maintenance cost stated, and the metrics that stop retries from masking a degrading dependency.
## Choose the owning layer The first decision is not *how* to retry but *who*. Retry logic needs two things to be correct: the ability to classify a failure as transient or permanent, and the knowledge of whether the operation is safe to repeat. Those two facts rarely live in the same place by accident, so you have to put them there. Classification lives lowest. The adapter that talks to a dependency sees the transport's real exception vocabulary — refused connections, timeouts, resets, error codes. By the time that has been translated into an application-level error, the distinction between "the network hiccupped" and "the request was rejected" is often collapsed into one type, and a retry above that layer is guessing. Safety-to-repeat lives highest. Only the caller knows whether this particular operation may run twice, and whether it carries an identifier that makes a repeat harmless. The practical resolution is a thin adapter that owns the mechanism and takes the policy — including "do not retry" — from the call site. What you must not do is let both ends retry independently. ## One layer, because attempts multiply Nesting is the failure mode that reaches production most often, and it is arithmetic, not opinion. A decorated helper called from a decorated function, each with three attempts, is nine calls to the dependency; add a third layer and it is twenty-seven. The delays compound too: the outer layer's backoff restarts the inner layer's full schedule, so the worst-case wall time is a product, not a sum, and it always exceeds whatever the caller thought the timeout was. The defences are cheap. Keep retries out of general-purpose helpers so they cannot be composed accidentally. Make the policy an explicit argument so a call site can pass "no retries" when it knows an outer layer owns them. And bound the whole path with a deadline that is captured once and passed down, so an inner layer knows how much of the budget is left rather than starting a fresh schedule of its own. ## Time, not attempts, is the contract Attempt counts are a convenient knob and a poor guarantee. What a caller — a request handler, a scheduler, another service — actually cares about is when it gets an answer. Express the policy as a deadline and treat the attempt count as a secondary safety limit. This also makes retry policy composable: a caller with 800 ms left can hand that down, and a layer with no budget left simply fails immediately instead of politely sleeping past it. ## Configuration, testability, observability Three properties separate a retry layer people trust from one they eventually delete. **Configuration.** Attempts, base delay, cap and deadline are operational knobs. They belong in configuration that can be changed for one dependency during an incident, not as literals in a decorator argument list spread over forty modules. **Testability.** Nothing kills a test suite faster than real backoff. Make the sleep a parameter that defaults to `time.sleep`, or replace it in tests with `unittest.mock.patch` at the module path where it is looked up. Then assert the interesting things: how many attempts happened, which exceptions were re-raised untouched, and — with a seeded `random.Random` — that the delay schedule grows and stays under the cap. The tests run in milliseconds and they test the policy rather than the clock. **Observability.** Retries are invisible by default: the call succeeds and nobody knows it took four attempts. Emit a counter per retried attempt, labelled by dependency and exception class, and log once at the final failure with the traceback. Retry rate is a leading indicator — it rises before error rate does. It is also how you find the paths where retries are papering over a real defect. Be careful about what the denominator is: if a cache in front of a dependency absorbs 83% of reads, a retry rate measured against all reads looks harmless while the origin path is retrying constantly, so measure against calls that actually reach the dependency. ## Hand-rolled or a library Hand-rolling is a legitimate answer. A dozen lines, one exception tuple, one schedule, no dependency, and everyone on the team can read it. If that describes your need, write it and move on. A maintained library earns its keep when the requirements grow past that: sync and async variants of the same policy, deadline propagation, per-call-site overrides, predicates that inspect the exception rather than just its class, hooks for metrics, and jitter someone else has already got right. Weigh it honestly — a third-party dependency is a supply-chain surface, an upgrade obligation and another API for reviewers to learn — but "we wrote our own" stops being a virtue once the home-grown version has grown its fourth keyword argument. The organisational move that outranks both: one internal module with the house policy and the house exception classification, so every service treats the same dependency's failures the same way, rather than forty call sites each re-deriving which errors are transient.
- Two layers on the same call path each retry three times. What does the caller experience?Up to nine calls to the dependency and a worst-case wait that is the product of the two schedules, not the sum — the outer backoff restarts the inner one from scratch each time. The caller's timeout is blown, the dependency sees three times the load it was meant to, and no single piece of code looks wrong in review. Pick one owning layer and have the others fail fast, ideally by passing an explicit no-retry policy.
- How do you unit-test a retry policy without the tests taking real seconds?Make the sleep an injectable parameter defaulting to `time.sleep`, or replace it with `unittest.mock.patch` at the module path where it is looked up rather than where it is defined. The test then records the delays instead of sleeping, and asserts the attempt count, the growth and cap of the schedule with a seeded `random.Random`, and that a non-retryable exception propagates on the first attempt.
- What should a retry layer log and measure so retries are not invisible?One warning per retried attempt carrying the dependency, the attempt number and the exception class, and one error with the traceback when attempts are exhausted — not a log line per attempt at info, which floods during an incident. Alongside it, a counter of retried attempts per dependency, measured against calls that actually reach that dependency rather than all calls, since a cache in front makes the ratio look far healthier than it is.
- When would you decide not to add retries to a call path at all?When the operation cannot be made safe to repeat and the ambiguity matters more than the extra success rate; when the caller has no latency budget left, so a retry only converts a fast failure into a slow one; or when a durable queue already provides at-least-once delivery and a second mechanism just duplicates work. Failing fast with a clear error is often the better engineering answer than improving a success-rate graph.
saying these in an interview costs you the question
- Adds a retry decorator at every layer of the call stack
- Hardcodes attempts and delays as literals across the codebase
- Tests the policy with real multi-second sleeps
- Retries everything by default to improve the success rate
- Treats adopting a retry library as a cost-free decision
- Bounds only attempts, never the caller's total latency