skip to content

Why does asyncio.sleep as a synchronisation device make an async test flaky?

level: middleimportance: must knowfreq 58%

answer

  1. It asserts on machine speed
  2. Timing is not ordering
  3. Loaded runner breaks the guess
  4. Let the code signal readiness
  5. Await an Event, bound by timeout

basics

~10 s

A sleep guesses how long another coroutine needs, and a loaded machine breaks the guess. Await an asyncio.Event the code under test sets instead: the handshake resumes as soon as the state is real.

solid answer

~50 s

`asyncio.sleep(0.1)` promises only 'not before 100 ms'; it says nothing about whether the other coroutine has made progress. On a loaded CI runner it has not, so the test fails intermittently, and raising the duration trades flakiness for a slow suite that is still flaky. Replace the guess with an explicit handshake: have the code under test (or an injected collaborator) call `asyncio.Event.set()` at the point of interest and have the test `await event.wait()`, which resumes on the very next loop iteration. Bound it with `async with asyncio.timeout(1):` (3.11+) so a broken build fails in a second instead of hanging. `await asyncio.sleep(0)` is different — it is a scheduling checkpoint that yields once, not a timing assumption — but an event is still preferable because it does not encode how many times the loop must turn.

code

python · 23 lines
python
import asyncio
from contextlib import suppress


async def build_pick_list(ready: asyncio.Event, lines: list[str]) -> None:
    lines.append("AISLE-04/BIN-12")
    ready.set()                     # signal the exact moment the state is real
    await asyncio.sleep(3600)       # the builder keeps working afterwards


async def main() -> None:
    ready = asyncio.Event()
    lines: list[str] = []
    task = asyncio.create_task(build_pick_list(ready, lines))
    async with asyncio.timeout(1):  # a hang fails fast instead of blocking
        await ready.wait()
    assert lines == ["AISLE-04/BIN-12"]
    task.cancel()
    with suppress(asyncio.CancelledError):
        await task


asyncio.run(main())

go deeper

for a junior

Be ready to say why a test that sleeps before asserting is unreliable, and to name asyncio.Event as the thing you await instead of a fixed delay. Knowing that time.sleep must never appear in a coroutine is expected.

for a middle

Explain the mechanics: a sleep is a lower bound on wall-clock time and says nothing about another coroutine's progress, while Event.wait resumes on the first loop iteration after set() — deterministic and faster.

for a senior

Show production judgement: convert a flaky suite by replacing every non-zero sleep with a signal, bound each wait with asyncio.timeout so hangs become fast failures, and build deliberate interleavings to pin races.

for a principal

Own the standard: sleeps in tests are a systemic flake source that erodes trust in CI. Decide whether the codebase bans them outright, what synchronisation seams new async code must expose, and who enforces it.

### Why the sleep is a guess asyncio is cooperative and single-threaded: a coroutine keeps the event loop until it hits an `await` that actually suspends. That gives you an unusual amount of determinism *about ordering* — control changes hands only at suspension points — but no determinism at all about **wall-clock timing**. `await asyncio.sleep(0.1)` says "resume me no earlier than 100 ms from now". It does not say "the other coroutine has finished by then", and it cannot: whether the other coroutine has made progress depends on the machine, the CI runner's load, other tests running in parallel, and any blocking work that stalls the loop. Consider a warehouse pick-list builder whose test starts the builder as a task, sleeps 100 ms, and asserts that three batches were emitted. On the developer's idle laptop this passes every time. On a shared CI runner where a neighbouring job has a 2.4 GB working set and the machine is paging, the builder gets its first slice of the loop 300 ms in, and the test fails with two batches — or worse, fails on Tuesdays only. The reflex fix is to raise the sleep to 500 ms, which converts a flaky test into a slow test that is *still* flaky, just less often. Do that fifty times and the suite takes twenty minutes and no one trusts a red run. ### The replacement: an explicit handshake Instead of guessing how long the other coroutine needs, make it **tell you** when it has reached the state you want to assert on. `asyncio.Event` is the standard vehicle: - The code under test (or a fake collaborator injected into it) calls `Event.set()` at the exact point of interest. - The test does `await event.wait()`, which resumes on the very next loop iteration after `set()` — no earlier, no later, and never a millisecond wasted. The handshake is *deterministic* rather than *probable*: the test cannot observe the state before it exists, and it does not linger once it does. It is also faster than the sleep it replaces, which is why converting a suite from sleeps to events usually cuts its runtime as well as its flake rate. Other legitimate handshakes: awaiting the task itself when the assertion is about its result; awaiting `asyncio.Queue.get()` when the code under test publishes; awaiting a `Future` the fake collaborator resolves; and `asyncio.Barrier` when you need *two* coroutines to arrive before either proceeds. ### Always bound the wait An `Event.wait()` with nothing around it converts a failure into a hang: if the code under test never sets the event, the test blocks until the whole suite is killed, and the output tells you nothing. Wrap it: ```python async with asyncio.timeout(1): await ready.wait() ``` `asyncio.timeout` (added in 3.11) raises `TimeoutError` at the boundary, so a broken build fails in one second with a clear traceback instead of hanging. The one-second budget is not a timing assumption — it is an upper bound on "obviously broken", and nothing about the test's correctness depends on its exact value. ### The one honest use of sleep `await asyncio.sleep(0)` is not a timing guess. It yields to the event loop exactly once, letting callbacks that are already scheduled run before your coroutine resumes. Used deliberately — "let the task I just created reach its first await" — it is a *scheduling checkpoint*. It is still worse than an event when you can have an event, because it encodes how many times the loop must turn before the state you want exists, and that count changes when the implementation adds an `await`. Treat `sleep(0)` as acceptable glue and any non-zero sleep as a defect. ### Forcing a race on purpose The same machinery lets you do the opposite: instead of removing nondeterminism, manufacture one specific interleaving. Give the code under test a collaborator that awaits an event you control, start both coroutines, then release them in the order that reproduces the bug. Because interleaving only ever changes at await points, a test built this way pins an exact schedule — for instance, proving that two concurrent builders writing a locale-dependent header both read the format before either wrote it. That test is deterministic and fails one hundred percent of the time before the fix and zero percent after, which is exactly what a regression test for a race should do — the "run it a thousand times and hope" approach is not a test, it is a lottery. ### The short version for an interview Sleeping asserts on the machine's speed; awaiting an event asserts on the program's state. Replace every non-zero sleep with a signal from the code under test, bound the wait with `asyncio.timeout` so hangs become fast failures, and reserve `asyncio.sleep(0)` for deliberately pumping the loop once.

  • How do you deliberately force one specific interleaving to reproduce a concurrency bug?
    Give the code under test a collaborator that awaits an event you control, start both coroutines, then release them in the order that triggers the bug. Because control changes hands only at await points, the schedule you build is exact and repeatable, so the regression test fails every time before the fix and never after — unlike running it a thousand times and hoping.
  • When is await asyncio.sleep(0) legitimate in a test?
    It yields to the event loop exactly once so already-scheduled callbacks can run — a scheduling checkpoint, not a timing guess. It is acceptable glue for 'let the task I just created reach its first await', but it encodes how many turns the loop needs, so it breaks when the implementation adds an await. Prefer an event whenever one is available.
  • What stops an Event handshake from hanging the suite forever when the code under test is broken?
    Wrap the wait in `async with asyncio.timeout(1):`, which raises `TimeoutError` at the boundary. The budget is an upper bound on 'obviously broken', not a timing assumption, so its exact value never affects correctness — it just turns an unbounded hang into a one-second failure with a traceback.

Sleeping is telling a colleague 'I'll assume you're done in ten minutes'. Awaiting an event is asking them to shout when they are done — always correct, and usually sooner.

saying these in an interview costs you the question

  • Raises the sleep duration until CI stops failing
  • Believes asyncio.sleep guarantees the other coroutine finished
  • Uses time.sleep inside a coroutine to wait for progress
  • Claims asyncio can switch coroutines at any line
  • Waits on an event with no timeout, so hangs replace failures
  • Asserts on measured elapsed time to prove ordering

context