A flight-schedule differ stacks a retry decorator over one that opens an output file; how does that order decide whether handles leak?
answer
- Retry re-invokes everything below it
- Count how often the resource is acquired
- One leak, or one leak per attempt
- Acquisition belongs inside the retry
- finally in the acquiring wrapper
basics
~20 sRetry re-invokes everything below it, so with retry on top the file wrapper runs once per attempt. That is correct only if it closes in a finally; otherwise each failed attempt of the 6,800-row diff strands a handle.
solid answer
~50 sRetry belongs **above** the wrapper that acquires the resource, so each attempt gets a freshly opened file and a failed attempt cannot leave a half-written one in play. That placement is only safe if the acquiring wrapper releases in a `finally` — or uses a `with` block — because retry converts one leak into one leak per attempt: three retries over a 6,800-row batch strand three handles instead of one, and a nightly job hits the process file-descriptor limit within days rather than never. The inverted stack, with retry beneath the acquisition, opens the file once and re-runs only the body, so every attempt appends to the same handle and a partially written diff is silently reused. Either order can be made correct; only one is correct by accident, and the acquiring wrapper must own its cleanup regardless.
code
python · 37 linesimport io
attempts = []
def opens_output(f):
def wrapper(rows):
handle = io.StringIO()
try:
return f(rows, handle)
finally:
handle.close()
return wrapper
def retry(f):
def wrapper(*args, **kwargs):
for attempt in (1, 2, 3):
try:
return f(*args, **kwargs)
except OSError:
if attempt == 3:
raise
return wrapper
@retry
@opens_output
def diff_schedules(rows, handle):
attempts.append(len(attempts) + 1)
if len(attempts) < 3:
raise OSError("transient")
handle.write(f"{rows} rows")
return handle.tell()
print(diff_schedules(6800), attempts)go deeper
Recall that the outer decorator re-calls the inner one, so a retry on top means everything below it runs again — including any setup work that wrapper performs.
Trace how many times each layer executes for three attempts, and explain why cleanup that is not in a finally is skipped when the wrapped call raises.
Diagnose from the symptom: decide whether leaks scale with attempts or with runs, choose per-attempt versus per-operation acquisition deliberately, and fix the missing cleanup rather than only the order.
Argue the structural fix — resources managed at the point of use instead of implied by decorator order — and set the standard that any wrapper acquiring on a callee's behalf ships with a released-on-exception test.
### The two candidate stacks The job diffs two flight schedules and writes the differences out. Two cross-cutting wrappers are involved: one acquires and releases the output file, the other retries transient failures. There are only two orderings, and they mean genuinely different things. **Retry outermost:** ``` @retry @opens_output def diff_schedules(rows, handle): ... ``` Retry's wrapper calls the acquisition wrapper, which opens a file, calls the body, and (correctly written) closes in a `finally`. A failure propagates up through the acquisition wrapper — running its cleanup on the way — and reaches retry, which calls the whole thing again. **Each attempt gets a fresh, fully scoped resource.** The file is opened three times for three attempts. **Retry innermost:** ``` @opens_output @retry def diff_schedules(rows, handle): ... ``` The file is opened once, and retry re-runs only the body, handing it the *same* handle each time. The second attempt appends to whatever the first attempt had already written before it failed, so a 6,800-row diff can emit a truncated block followed by a complete one and no error anywhere says so. ### Where the leak actually comes from Ordering does not create the leak; it **multiplies** it. The leak comes from an acquisition wrapper that opens without a `try`/`finally`: ```python def opens_output(f): def wrapper(rows): handle = open("diff.txt", "w") # no finally result = f(rows, handle) handle.close() # skipped when f raises return result return wrapper ``` When the body raises, `close()` is never reached. On CPython the handle is usually reclaimed when the last reference goes away, but that is a refcounting accident, not a guarantee: the exception's traceback holds the frame alive, a reference cycle or an exception stored for later reporting extends it further, and on any implementation without refcounting the file stays open until a collection runs. Under a retry decorator this stops being a rare, self-correcting event: three attempts mean three abandoned handles per run, every run, and a nightly batch reaches the process descriptor limit on a schedule. The order therefore has a diagnostic signature. If the leak count per failed run scales with the retry count, acquisition is inside retry and cleanup is not in a `finally`. If exactly one handle is stranded per failed run no matter how many attempts occurred, acquisition is outside retry. ### The rule to apply **Acquire inside the retry, release in a `finally`.** Stated more generally: any wrapper holding state that must not survive a failed attempt belongs *below* the retry decorator, and any wrapper whose work must happen once per logical operation belongs *above* it. That distinction settles the other layers in the same stack: - **A logging or audit wrapper above retry** logs one line per logical operation; **below retry** it logs one per attempt. Both are legitimate; pick deliberately, and note that the below-retry position is the one that lets you count transient failures. - **A timing wrapper above retry** measures total wall time including all attempts and back-off sleeps; **below retry** it measures a single attempt. A dashboard built on the wrong one will not show the retries at all. - **A transaction or lock-acquisition wrapper** is in the same family as the file: it must be inside retry, because retrying inside an already-failed transaction or while still holding a lock is worse than not retrying. - **A wrapper that mutates shared accumulators** — a partial-results list, a counter — must either be inside retry so each attempt starts clean, or be explicitly reset by retry between attempts. ### Making the intent explicit Because both orders parse and both run, the safest structure does not rely on the stack to express the invariant at all. Move acquisition into the body with a `with` block, or into a context manager the body enters, and the resource's lifetime becomes visible at the point of use rather than implied by two lines of decorator order that a later edit can silently swap. A decorator that acquires a resource for its callee is the fragile pattern; a decorator that retries a callee which manages its own resource is the robust one. If the acquisition really must be a decorator — because many functions in the differ share it — then its `finally` is not optional, and a test that asserts the resource is closed after the wrapped callable raises is the test that keeps it. ### What to say when handed the symptom Given "the nightly differ runs out of file descriptors", the sequence is: confirm the leak scales with failures rather than with rows; read the decorator stack top to bottom and identify which layer acquires and which retries; check the acquiring wrapper for a `finally`; then decide whether per-attempt or per-operation acquisition is the intended semantic before changing the order. Reordering without fixing the missing `finally` moves the leak; adding the `finally` without deciding the semantics leaves the truncated-output bug in place.
- Where would you put a timing wrapper in that same stack, and what changes?Above the retry decorator it measures the whole logical operation including every attempt and any back-off; below it, one attempt at a time. Neither is wrong, but a dashboard fed by the inner position will never show that retries are happening, while the outer position hides how slow a single attempt is. Pick one and name it in the metric.
- How would you prove the ordering is the cause rather than guessing?Instrument the acquiring wrapper to record each acquire and release, then force a transient failure and count. If acquisitions equal the attempt count, acquisition is inside retry; if it stays at one, retry is inside acquisition. Then assert in a test that a raising callable still leaves the resource closed — that test fails loudly on a missing `finally` regardless of ordering.
saying these in an interview costs you the question
- Says the order cannot matter because both wrap the same function
- Relies on CPython refcounting to close abandoned handles
- Puts acquisition outside retry so attempts share one handle
- Reorders the stack without adding a finally
- Cannot say how many times each layer runs per attempt
- Assumes retry resets state the inner wrappers accumulated