A geocoding batch using asyncio.to_thread() starts timing out unrelated connections - why?
answer
- One pool behind every offload
- Its cap is smaller than people assume
- The loop is a customer of it too
- Name resolution waits behind your batch
- Dedicated executor, named threads, explicit shutdown
basics
~20 sThe batch saturated the loop's single default thread executor. That pool is shared, capped near min(32, CPUs + 4), and the loop also uses it for name resolution, so every free worker held by the batch stalls new connections. Give the batch its own bounded executor.
solid answer
~40 s`asyncio.to_thread()` always submits to the *loop's default executor* - one shared `ThreadPoolExecutor` whose cap is `min(32, CPUs + 4)`. Fire a 340-case regression pack of blocking geocoding calls through it and every worker is busy; anything else that needs a thread simply queues. The sharp part is that the loop resolves hostnames in that same pool, so outbound connections cannot even start and surface as connect timeouts - which is why the timing artefacts get misread as clock skew rather than as pool starvation. The fix is isolation: create a dedicated `concurrent.futures.ThreadPoolExecutor(max_workers=..., thread_name_prefix=...)` for the batch and submit through the loop's `run_in_executor()`, keeping the default pool free. Bound the in-flight batch with an `asyncio.Semaphore`, use `functools.partial` for keyword arguments, and shut the executor down yourself - `asyncio.run()` only shuts down the default one.
code
python · 20 linesimport asyncio
import functools
from concurrent.futures import ThreadPoolExecutor
def geocode(address, *, retries=2):
return address.upper() # stands in for a blocking client call
async def main():
loop = asyncio.get_running_loop()
with ThreadPoolExecutor(max_workers=8, thread_name_prefix="geocode") as pool:
calls = [
loop.run_in_executor(pool, functools.partial(geocode, a, retries=3))
for a in ("12 pine st", "9 elm rd")
]
print(await asyncio.gather(*calls))
asyncio.run(main())go deeper
Take away one fact: the threads behind an offload come from a shared pool with a limited number of workers, so offloading is not free and not unlimited.
Explain where those workers come from, roughly how many there are, and how to submit to a pool of your own through the loop's run_in_executor(), including the positional-arguments-only restriction.
Diagnose the coupling out loud: a saturated shared pool starves the loop's own name resolution, so unrelated connects time out. Then isolate with a bounded named executor, bound fan-out, and own the shutdown.
Set the standard: which dependencies get dedicated pools, how their sizes relate to downstream capacity, and what the service exposes so the next saturation is attributable in minutes rather than argued about as clock skew.
## The shape of the incident A batch job geocodes addresses through a synchronous client, correctly offloaded with `await asyncio.to_thread(client.geocode, addr)`, and fans out across a 340-case regression pack. The batch itself completes. What breaks is everything *else* in the process: unrelated outbound calls start failing at connect time, health checks flap, and the latency histogram fills with values that look impossible against wall-clock timestamps - which is why the first hypothesis is usually a clock-skew artefact in the timing logs. It is not. It is one saturated thread pool. ## One pool, shared by everything `asyncio.to_thread()` takes no executor argument. It submits to the running loop's **default executor**, created lazily as a `concurrent.futures.ThreadPoolExecutor` the first time anything needs it. Every `to_thread` call anywhere in the process shares it, and its worker cap is the pool's own default: `min(32, CPUs + 4)`, with the CPU count coming from `os.process_cpu_count()` since 3.13 - so a container pinned to two CPUs gets six workers, not thirty-two. Submissions beyond the cap sit in an unbounded queue, which is why saturation shows up as latency rather than as an error. The part that turns a slow batch into an outage is that the loop uses the same pool for its own work. Address resolution (`getaddrinfo`) is a blocking C call, so the default event loop performs it in the default executor. With every worker parked in the geocoding client, a new connection cannot resolve its host; it waits behind the batch and eventually reports a connect timeout. The failure lands on components that have nothing to do with the batch, which is exactly what makes it hard to attribute. ## Isolation is the fix Give the noisy dependency its own bounded pool and route it through the loop's `run_in_executor()`: ```python loop = asyncio.get_running_loop() pool = ThreadPoolExecutor(max_workers=8, thread_name_prefix="geocode") await loop.run_in_executor(pool, functools.partial(geocode, addr, retries=3)) ``` Three properties come with that. The default pool stays free for name resolution and for short offloads elsewhere. The thread name prefix makes stack dumps and thread listings legible during the next incident. And the pool size becomes an explicit capacity decision per dependency instead of an accident of CPU count. `run_in_executor()` accepts positional arguments only, so keyword arguments go through `functools.partial`. It is also the door to a `ProcessPoolExecutor` when the offloaded work is CPU-bound, at the cost of picklable arguments and no lambdas or closures. ## Bounding, and who cleans up A pool cap limits concurrency but not the queue: submitting all 340 cases at once still allocates 340 futures and their argument graphs immediately. Bound the fan-out on the async side too - an `asyncio.Semaphore` around each `run_in_executor` await, or batching the input - so memory tracks in-flight work rather than total work. Cleanup is yours. `asyncio.run()` shuts down the *default* executor before returning, waiting for its threads; it knows nothing about pools you created. Use the executor as a context manager, or call `shutdown(wait=True, cancel_futures=True)` explicitly on the way out. Note that a synchronous `shutdown(wait=True)` called from the loop thread blocks the loop while it drains, so at shutdown time do it after the loop's work is done, or offload the shutdown itself. ## Sizing the dedicated pool For blocking I/O, worker count is a concurrency budget, not a CPU budget: the right number is roughly the target in-flight requests, capped by what the downstream dependency and its connection pool tolerate. Oversizing turns your outage into the dependency's. Undersizing shows up as queueing delay in *your* pool, which is at least attributable, because the wait is now behind a named prefix instead of behind name resolution. ## The diagnostic move to remember When unrelated connections start timing out while a blocking batch runs, dump thread names and stacks. If you see the whole default pool parked in one library's frames, you have found it: the coincidence is not clock skew, it is a shared, bounded, invisible resource with two very different classes of user on it.
- How would you confirm the default thread pool is the bottleneck rather than the dependency?Dump every thread's name and stack while it is happening - the default pool's threads carry an asyncio prefix, and seeing all of them parked in one client's frames is the confirmation. Cross-check that the failing operations are *connects* rather than reads, since name resolution is what shares the pool. If a dedicated executor for the batch makes the unrelated timeouts disappear, the attribution is settled.
- How do you size a dedicated thread pool for blocking I/O?Treat workers as a concurrency budget, not a CPU budget: start from the in-flight request target the downstream dependency and its own connection pool can absorb, and cap there. Oversizing exports your queue into someone else's service; undersizing shows up as queue delay inside a pool you named, which is at least attributable. Measure queue wait, not just call latency.
- Who shuts down a custom executor passed to the loop's run_in_executor()?You do. `asyncio.run()` shuts down only the loop's default executor before returning. Use the pool as a context manager or call `shutdown(wait=True, cancel_futures=True)` on the exit path. Be aware that a blocking shutdown called from the loop thread stalls the loop while workers drain, so do it once the loop's real work is finished.
saying these in an interview costs you the question
- Assuming asyncio.to_thread() creates threads on demand without a cap
- Not knowing the default executor is shared process-wide per loop
- Blaming clock skew or the dependency before checking pool saturation
- Passing keyword arguments straight to run_in_executor()
- Expecting asyncio.run() to shut down a custom executor
- Sizing an I/O thread pool by CPU count