Should a Python retry loop wrap the entire `with` block or live inside it?
answer
- The failure may have damaged something
- Ask what the block is holding
- Fresh resource or reused resource
- Generator context managers are single-use
- Sleep outside, re-enter per attempt
basics
~20 sIt depends on whether the failure damaged what the block is holding. If the resource carries protocol or transaction state, wrap the whole with statement so each attempt acquires a fresh one; retrying inside reuses something the failure may already have poisoned.
solid answer
~50 sAsk what the context manager is holding. A connection mid-message, an open transaction the far side already rolled back, or a file object at an unknown offset is not usable for a second attempt, so the retry belongs *outside* the `with`: each attempt calls the factory again, gets a clean resource, and `__exit__` runs per attempt to close or roll back. Retrying inside is right only when the resource is genuinely reusable and stateless, and it saves a handshake per attempt. Two mechanics matter. A generator-based context manager created by `contextlib.contextmanager` is single-use: the instance is consumed by its first `with`, and entering it a second time raises instead of re-running, so the loop must call the factory rather than re-enter a saved instance. And the backoff sleep goes outside the block: sleeping while holding a lock or a connection turns one slow dependency into contention for everything else.
code
python · 21 linesimport contextlib
@contextlib.contextmanager
def session():
print("open")
try:
yield {"open": True}
finally:
print("close")
cm = session()
with cm:
print("first attempt")
try:
with cm:
print("second attempt")
except Exception as exc:
print("reuse failed:", type(exc).__name__)go deeper
Know that a with statement acquires something on entry and releases it on exit, and that the release still happens when the body raises. That alone tells you a failed attempt does not silently leak the resource.
Be able to place the loop correctly and say why: outside the with when a fresh resource is needed per attempt, inside only when the object is genuinely reusable. Know that a contextlib.contextmanager instance cannot be entered twice.
Reason about resource state after the specific failure — a half-read socket, an aborted transaction — and show that you check cleanup cost, __exit__ suppression, and that the backoff sleep happens while holding nothing.
Decide how resource lifetime and retry boundaries are expressed across a codebase: whether acquisition is hidden behind a factory that is safe to call per attempt, and how you prevent teams from wrapping retries around blocks that quietly hold pool capacity.
## The question behind the question "Inside or outside" is really "is the thing this block acquired still usable after the failure?" Everything else follows from the answer. ## When the resource is poisoned Most failures worth retrying are transport failures, and transport failures tend to damage exactly the object you are holding. A socket that timed out mid-response may still have unread bytes queued, so the next request on it reads the tail of the previous reply and everything after that is misaligned. A database session whose statement failed may be in an aborted transaction that rejects every subsequent statement until it is rolled back. A file object written to by a partially-failed batch sits at an offset nobody can name. Retrying inside the block reuses that object. The best case is that every subsequent attempt fails immediately with a different, more confusing exception — "connection already closed" instead of the original timeout — and the retry burns its attempts learning nothing. The worst case is a half-success: the reused resource accepts the second attempt and writes it in the wrong place. So the default shape is the retry *around* the whole statement: ```python for attempt in range(1, attempts + 1): try: with open_session() as session: return session.flush(batch) except (ConnectionError, TimeoutError): if attempt == attempts: raise time.sleep(backoff(attempt)) # holding nothing ``` Each attempt calls the factory again, so each gets a fresh resource. `__exit__` runs on every failed attempt with the exception information, which is what lets a session roll back or a connection close rather than leaking. ## When inside is right Retrying inside is correct when the resource genuinely survives the failure and re-acquiring it is expensive. A pooled handle that validates itself, an in-memory structure guarded by a lock, or a client object whose failure mode is "this one call did not work" can all be reused. The gain is real: you skip a connection setup and any authentication on every attempt. Just make the claim explicitly rather than by accident — "this object is safe to reuse after this exception" is an assertion about a library's behaviour, and it should be one you have checked. A middle ground exists: retry inside for failures the resource survives, and let the enclosing loop re-acquire for failures it does not. That is two exception tuples and two loops, and it is worth the complexity only on a hot path. ## Single-use context managers A mechanic that catches people out: the object returned by a function decorated with `contextlib.contextmanager` wraps a generator, and a generator runs once. Enter it, exit it, and that instance is spent — entering it again raises rather than re-running the body (on CPython 3.14 an `AttributeError`, because the instance discards its stored arguments on first entry). So this is broken: ```python cm = open_session() # one instance for attempt in range(3): with cm: # raises on the second pass ... ``` and the fix is to call the factory inside the loop so each attempt gets a new instance. Class-based context managers may or may not be reusable; a `threading.Lock` is, an arbitrary session object usually is not. Never assume. ## What `__exit__` does on the failure path When an attempt raises inside the block, `__exit__` is called with the exception type, value and traceback before the exception continues outward. Two consequences. First, `__exit__` can do real work on that path — roll back, abort, close, flush a buffer — and if that work is expensive or has side effects, you are paying it on every attempt, not once. Second, if `__exit__` returns a truthy value it *suppresses* the exception, so an enclosing retry never sees a failure and the loop simply returns whatever comes next. A context manager that swallows errors makes retries silently pointless. Also note that an exception raised by `__exit__` itself — a close that fails — happens outside the block's body, so a retry written to guard only the body will not catch it. If closing can fail in a way you want to retry, the retry has to be outside the `with`. ## Where the sleep goes Outside the block, always, when the block holds anything contended. Sleeping through a doubling backoff while holding a lock means every other worker waiting on that lock waits with you, and the wait grows with each attempt. Compute the delay, leave the block, sleep, then re-acquire on the next attempt. The same reasoning applies to open connections and to any resource drawn from a bounded pool: holding one idle during the backoff removes it from everyone else at the exact moment the system is under stress.
- Which kinds of resource are actually safe to reuse across retry attempts?Ones that carry no per-call protocol or transaction state: a `threading.Lock`, an in-memory structure, a stateless client object whose failure mode is confined to the single call. Anything holding a half-read response, an aborted transaction, or a file position determined by a partial write is not safe, because the next attempt starts from an unknown state. Treat reusability as a property you verified in the library's documentation, not as a default.
- Does `__exit__` still run when an attempt fails inside the `with` block?Yes — it is called with the exception type, value and traceback before the exception propagates, which is how a session rolls back or a connection closes on the failure path. Two things follow: any expensive cleanup happens on every attempt rather than once, and if `__exit__` returns a truthy value it suppresses the exception entirely, so an enclosing retry loop never sees a failure and stops retrying for the wrong reason.
- Why should the backoff sleep never happen inside the `with` block?Because the block is holding something. Sleeping through a doubling delay while holding a lock blocks every other worker waiting on it, for progressively longer each attempt; holding a pooled connection idle removes it from the pool exactly when the system is under stress. Compute the delay, exit the block, sleep holding nothing, and acquire again on the next attempt.
saying these in an interview costs you the question
- Retries inside the with, reusing a connection the failure poisoned
- Re-enters one spent generator-based context manager instance
- Sleeps the backoff while still holding a lock or connection
- Returns True from __exit__, so the retry never sees the failure
- Assumes cleanup in __exit__ is free to repeat per attempt
- Ignores that re-acquiring costs a handshake on every attempt