What does asyncio.to_thread() do, and when should a coroutine use it?
answer
- One thread runs every coroutine
- Blocking one task blocks them all
- Hand the call to a worker thread
- Awaitable wrapper over the default thread pool
- Only helps when the GIL is released
basics
~20 sasyncio.to_thread() runs a blocking function on a worker thread and returns an awaitable for its result, so the event loop keeps serving other tasks meanwhile. Await it for any synchronous call that would otherwise freeze the loop.
solid answer
~40 sAn asyncio event loop is a single thread running one task at a time; a blocking call inside a coroutine stops **every** other task until it returns. `asyncio.to_thread(func, *args, **kwargs)`, added in 3.9, hands `func` to the loop's default thread pool and gives you an awaitable, so the loop stays responsive: `rows = await asyncio.to_thread(cursor.fetchall)`. It copies the caller's context variables into the worker thread, and it is the right tool when the blocking call releases the GIL - socket and file I/O, `time.sleep()`, most C-extension work. It is the wrong tool for pure-Python CPU work, which still contends for the GIL; that belongs in a process pool. Two edges matter: cancelling the awaiting task does not interrupt the running thread, and the pool is shared and bounded, so it can be saturated.
code
python · 22 linesimport asyncio
import time
def blocking_lookup(n):
time.sleep(0.2) # stands in for a synchronous client call
return n * 2
async def ticker():
for _ in range(3):
print("loop still alive")
await asyncio.sleep(0.05)
async def main():
offloaded = asyncio.gather(*(asyncio.to_thread(blocking_lookup, i) for i in range(4)))
results, _ = await asyncio.gather(offloaded, ticker())
print(results)
asyncio.run(main())go deeper
Be ready to say why one blocking call inside a coroutine hurts the whole program, and to show the one-line fix: await asyncio.to_thread(func, arg). Remember it returns something you must await.
Explain the mechanics: the coroutine submits to the loop's default ThreadPoolExecutor, context variables are copied, and only work that releases the GIL genuinely runs in parallel with the loop.
Show that you treat the default pool as a shared, bounded resource: reason about its worker cap, what happens when it saturates, that cancellation cannot interrupt a running thread, and where timeouts must actually live.
Own the policy: which dependencies are allowed to use the shared pool, which get dedicated bounded executors, and how the team keeps blocking calls from re-entering coroutines as the codebase grows.
## The problem it solves An asyncio program runs its event loop on one OS thread. Coroutines are *cooperatively* scheduled: control returns to the loop only at an `await` that actually suspends. If a coroutine calls something that blocks the thread - a synchronous HTTP client, a database driver's `execute()`, `time.sleep()`, a large `open(...).read()` - the loop cannot run timers, cannot read ready sockets, and cannot start any other task. One 400 ms blocking call in one coroutine is 400 ms of latency added to every other in-flight request. This is the single most common way an async service ends up slower than the threaded one it replaced. ## What to_thread actually does `asyncio.to_thread(func, /, *args, **kwargs)` was added in Python 3.9. It is a coroutine function: calling it produces a coroutine object that you must `await` (or wrap in a task). When awaited it does three things: 1. Takes a snapshot of the current context variables, so anything stored in a `contextvars.ContextVar` - a request id, a correlation token - is visible inside `func`. 2. Submits `func` to the running loop's **default executor**, a `concurrent.futures.ThreadPoolExecutor` created lazily on first use. 3. Wraps the resulting `concurrent.futures.Future` so the awaiting coroutine suspends until the thread finishes, then returns the value or re-raises the exception. The loop is free the whole time. Other tasks run; timers fire; sockets get serviced. ## The GIL boundary Offloading only helps when the work actually releases the Global Interpreter Lock. Blocking I/O in the standard library and in C extensions releases it, so the worker thread genuinely waits in the kernel while the loop thread executes Python. Pure-Python CPU work does **not** release it: moving a tight numeric loop into a thread just interleaves it with the loop thread and can make latency worse, because the loop now competes for the GIL. CPU-bound work belongs in `concurrent.futures.ProcessPoolExecutor` (or a subinterpreter-based pool). The free-threaded build became officially supported in 3.14 (PEP 779) and removes that GIL constraint, but it is a separate build with a 5-10% single-threaded penalty and is not what a default `python3.14` gives you. ## The pool is shared and bounded There is exactly one default executor per loop, and everything using `to_thread` shares it. Its worker cap is the `ThreadPoolExecutor` default - `min(32, CPUs + 4)`, where since 3.13 the CPU count comes from `os.process_cpu_count()`, which respects CPU affinity and the `-X cpu_count` / `PYTHON_CPU_COUNT` overrides. Fill those workers with slow calls and further `to_thread` awaits simply queue. `asyncio.run()` waits for the default executor's threads to finish before returning, so a runaway offloaded call also delays shutdown. ## Cancellation does not reach into the thread Cancelling the task that awaits `to_thread` raises `asyncio.CancelledError` in the *awaiting coroutine* only. Python cannot interrupt a thread that is inside a blocking C call, so `func` keeps running to completion and its result is discarded. Anything with a real timeout requirement needs the timeout pushed down into the blocking call itself (a socket timeout, a driver timeout); wrapping the await in `asyncio.timeout()` bounds *your* wait, not the thread's work. ## Thread-safety is yours to prove `func` now runs concurrently with other workers and with the loop thread. Shared mutable state it touches needs a `threading.Lock`, not an `asyncio.Lock` - the latter protects coroutines on one thread and does nothing across threads. Equally, `func` must not touch loop objects: creating tasks or putting to an `asyncio.Queue` from a worker thread is a data race with the loop. ## Choosing between to_thread and run_in_executor `to_thread` is the ergonomic default: no loop handle, keyword arguments supported, context propagated. Reach for the loop's `run_in_executor()` when you need a *different* pool - a bounded pool dedicated to one dependency so it cannot starve the shared one, a pool with a naming prefix for readable stack dumps, or a process pool for CPU work. That call takes positional arguments only, so keyword arguments go through `functools.partial`.
- Where do the threads that asyncio.to_thread() uses come from, and how many are there?They come from the running loop's default executor, a `concurrent.futures.ThreadPoolExecutor` created on first use and shared by every `to_thread` call on that loop. Its worker cap is the pool default, `min(32, CPUs + 4)`, with the CPU count taken from `os.process_cpu_count()` since 3.13. Once those workers are busy, further offloads queue rather than spawning new threads.
- Would you use asyncio.to_thread() for a CPU-bound function on a default CPython 3.14 build?No. Pure-Python CPU work holds the GIL, so the worker thread and the loop thread take turns instead of running in parallel, and loop latency gets worse rather than better. Use `concurrent.futures.ProcessPoolExecutor` for that, accepting the pickling and startup cost. Only a C-level call that releases the GIL, or the separately built free-threaded interpreter, changes that answer.
- How do you pass keyword arguments to a function offloaded with the loop's run_in_executor()?Wrap it: `functools.partial(func, arg, timeout=5)`. `run_in_executor(executor, func, *args)` accepts positional arguments only. `asyncio.to_thread()` does not have that limitation - it forwards both `*args` and `**kwargs` - which is one reason it is the friendlier default when the shared pool is fine.
The event loop is a single clerk serving a queue of customers. A blocking call is the clerk walking to the archive himself while everyone waits; to_thread sends a runner to the archive so the clerk keeps serving.
saying these in an interview costs you the question
- Claiming to_thread gives CPU-bound Python code real parallelism
- Thinking to_thread spawns a fresh thread per call, unbounded
- Believing cancelling the await stops the running thread
- Calling to_thread without awaiting the coroutine it returns
- Guarding thread-shared state with asyncio.Lock instead of threading.Lock
- Assuming any blocking call inside a coroutine only slows that coroutine