A threading.Semaphore caps concurrent calls in a fraud-scoring service, but permits leak until throughput reaches zero. What causes that, and how do you prevent it?
answer
- A count that never comes back
- The failure path skipped a step
- Context manager guarantees the pairing
- The bounded variant turns the mirror bug into ValueError
basics
~20 sSome path acquires a permit and never releases it — typically an exception swallowed by a broad except between acquire() and release(). The counter only falls, so concurrency drops until every caller blocks. Acquire with with sem: or try/finally, and use BoundedSemaphore to catch the mirror bug.
solid answer
~60 sA `threading.Semaphore` is a counter, not an owned lock: `acquire()` decrements it and blocks at zero, `release()` increments it, and nothing ties a release to the thread that acquired. If a worker acquires a permit and then hits an exception that a broad `except` swallows before `release()` runs, that permit is gone for the life of the process. With a 17-service dependency graph behind the scorer, one flaky dependency raising on a rare path bleeds the limiter one permit at a time — latency creeps up, then every request parks in `acquire()` and throughput hits zero with no error in the logs. The fix is to make the pairing structural: use `with sem:`, which releases in `__exit__` even when the body raises, or an explicit `try`/`finally`. Use `threading.BoundedSemaphore` so the symmetric bug — releasing more than you took — raises `ValueError` at the point of the defect instead of silently widening the limit. Bound the wait too: `acquire(timeout=...)` returns `False` rather than parking forever, so an exhausted limiter sheds load instead of hanging.
code
python · 14 linesimport threading
sem = threading.Semaphore(2)
def leaky_call():
sem.acquire()
try:
raise RuntimeError("downstream dependency failed")
except RuntimeError:
return
leaky_call()
leaky_call()
print("permit available within 0.1s:", sem.acquire(timeout=0.1))go deeper
Recall the shape of the fix even if the diagnosis is new to you: acquire a semaphore with a with block so the release cannot be skipped, and never put release() only at the end of the happy path.
Explain the counter mechanics — acquire decrements and blocks at zero, release increments, nobody owns a permit — and why that makes a skipped release permanent. Know what BoundedSemaphore adds and what acquire's timeout returns.
Show that you can read the symptom backwards: rising queueing latency with healthy dependencies, then every worker parked in the same acquire frame. Talk about narrowing the except clause, timing out acquisitions, and instrumenting saturation before capacity hits zero.
Own where limits live at all — in-process semaphores versus a connection pool, a bulkhead per dependency versus one shared budget across a large dependency graph — and the failure mode you accept in each. Argue for shedding load over unbounded queueing as a system policy.
### What a semaphore is here `threading.Semaphore(value)` wraps an integer and a condition. `acquire()` waits until the counter is positive, decrements it, and returns `True`; `release(n=1)` increments it and wakes waiters. That is the whole contract, and two properties of it drive every real bug: * **No ownership.** Unlike `threading.Lock` — and unlike `threading.RLock`, which raises when a non-owner releases — a semaphore does not record who took a permit. Any thread may release, and a thread may release a permit it never acquired. Nothing raises. * **No self-healing.** The counter is pure state. A permit that is never released is not reclaimed when the thread dies, when the request finishes, or when the garbage collector runs. It is simply gone. Used as a concurrency limiter — "at most eight outbound calls in flight at once" — those two properties mean the limiter's effective capacity is a monotonically non-increasing function of how many code paths forget to release. ### The failure, in order A scoring request fans out across a dependency graph of seventeen services. Each outbound call takes a permit. One dependency starts returning a malformed payload on a rare input, the parse raises, and the worker's `except Exception:` logs nothing useful and returns a default score. The request looks *successful* from the outside — that is what makes a swallowed exception so expensive here — but the `release()` on the line after the parse never ran. Now: capacity 8 becomes 7, then 6. Queueing delay in front of `acquire()` grows superlinearly as capacity shrinks against constant arrival rate, so the first symptom is a latency graph bending upward with no corresponding change in downstream latency. Eventually the counter reaches zero, every worker blocks in `acquire()`, and the service stops. There is no exception, no crash, no restart — just threads parked in a wait. A process dump shows every worker thread with the same stack ending in `acquire`, which is the fingerprint of a leaked-permit stall as opposed to a slow dependency. ### The fixes, strongest first **1. Make release structural, never a statement.** `threading.Semaphore` implements the context-manager protocol: `with sem:` acquires on entry and releases in `__exit__`, and `__exit__` runs whether the body returns, raises, or breaks. If you cannot use `with` — because you need a timeout — write the explicit form: ```python if not sem.acquire(timeout=2.0): raise TimeoutError("scorer at capacity") try: call_dependency() finally: sem.release() ``` The `finally` is the entire point. A `release()` sitting at the end of the happy path is a bug waiting for its first exception. **2. Use BoundedSemaphore for anything that models capacity.** `threading.BoundedSemaphore(value)` remembers its initial value and raises `ValueError` if `release()` would push the counter above it. It cannot detect a *missing* release — nothing can, since a permit legitimately stays out for as long as the work takes — but it catches the mirror-image defect: a double release, or a release on a path that never acquired, which silently *raises* the concurrency limit and defeats the protection you thought you had. Since a bounded semaphore fails at the defective call rather than at some later victim, it should be the default for a limiter. **3. Never acquire without a bound.** `acquire(timeout=...)` returns `False` instead of raising when it gives up, and `acquire(blocking=False)` returns immediately. Either turns "hang forever" into a decision you can make — shed the request, return a degraded score, increment a saturation counter. A limiter that can only queue is a limiter that converts a capacity problem into an availability problem. **4. Stop swallowing exceptions.** The permit leak is a symptom; the broad `except` that returns a default score while discarding the traceback is the disease. Catch narrowly, log with the exception attached, and let anything you did not anticipate propagate — then the `finally` still runs and the alert fires on the real cause. **5. Instrument it.** A semaphore exposes no supported public counter to read, so track saturation yourself: count timed-out acquisitions and the time spent waiting, and alert when either trends up. Both move well before the counter reaches zero. A last design note: because a semaphore has no owner, it is not a substitute for a lock and is not reentrant. A thread that takes two permits from a limiter of size one to guard a recursive call will deadlock, and no exception will tell you so.
- What does threading.BoundedSemaphore give you that threading.Semaphore does not?It records the initial value and raises `ValueError` if a `release()` would take the counter above it. That converts a double release or a release on a path that never acquired — which silently widens the concurrency limit — into an immediate failure at the defective line. It cannot detect a missing release, so it complements, rather than replaces, `with`-based acquisition.
- What does threading.Semaphore.acquire(timeout=1.0) return when it gives up?`False`, with no exception raised. That is the same contract as `threading.Lock.acquire`, and it means the return value must be checked — code that ignores it proceeds as though it holds a permit and will then release one it never took, inflating the limit. Pair the check with a shed-load or degrade path.
- Why can a threading.Semaphore not be used as a reentrant lock?It has no owner and no recursion count. A thread that already holds the only permit and acquires again blocks against itself exactly as a plain `Lock` would, and because any thread may release, a stray `release()` elsewhere can let a second thread into a section you believed was exclusive. Use `RLock` for re-entrancy and a semaphore only for counting capacity.
Permits are keys on a hook by the door. Nothing forces a borrower to hang the key back, so every borrower who leaves through a window shrinks the pool by one — until nobody can get in and no alarm has sounded.
saying these in an interview costs you the question
- Pairs acquire with release only on the success path
- Catches an exception and returns without releasing
- Uses plain Semaphore where over-release should be an error
- Calls acquire with no timeout in request-serving code
- Thinks a permit is reclaimed when the thread exits
- Restarts the process instead of finding the unbalanced path