skip to content

If tasks in a runtime switch only at explicit suspension points and share one carrier thread, can you still get race conditions? Explain what is and is not guaranteed.

level: seniorimportance: must knowfreq 42%

answer

  1. no data races, still atomicity violations
  2. every await = arbitrary other code ran
  3. check-then-act across a pause is broken
  4. never suspend with an invariant broken
  5. re-entrancy: same handler live twice

basics

~20 s

Yes. Single-threaded cooperative scheduling removes low-level data races — no two tasks touch memory at the same instant — but not logical races. Any suspension point is a place where arbitrary other tasks ran, so check-then-act sequences spanning an await can still be violated.

solid answer

~50 s

Two different hazards get conflated. **Data races** — two threads touching the same location concurrently with at least one write, giving torn or stale reads. On a single carrier thread with cooperative switching, these cannot happen: only one task executes at a time and switches happen at known points. **Atomicity violations / logical races** — a sequence that must be all-or-nothing gets interleaved. These absolutely still happen. Every suspension point is a place where other tasks ran arbitrary code, including code that mutates the same state you are midway through updating. The practical rule: *treat each suspension point as if everything else in the process ran there*. Invariants must hold whenever you pause. Check-then-act across an await is broken (the check may be stale). Re-entrancy is real: the same handler can be entered again while an earlier instance is suspended. Also, if the runtime multiplexes tasks over multiple carrier threads, low-level data races come back too, so single-threadedness is a property you must verify, not assume.

code

text · 10 lines
text
Task A                          Task B
----------------------------    ----------------------------
check cache -> miss
start fetch, SUSPEND
                                check cache -> miss (stale!)
                                start fetch, SUSPEND
resume, put(key, v1)
                                resume, put(key, v2)  // overwrites

result: two fetches, two side effects, last write wins

go deeper

for a junior

Recall that single-threaded means no simultaneous memory access, but pausing lets other code run, so multi-step updates spanning a pause can still be interleaved.

for a middle

Name the check-then-act pattern, show the duplicate-fetch example, and state the invariant rule: do not suspend while your data is halfway updated.

for a senior

Cover re-entrancy, completion ordering, cancellation landing at suspension points, and why thread-affine locks and thread-locals break when a task resumes elsewhere.

for a principal

Position it as a systems property: cooperative scheduling reduces the interleaving space to an auditable set, which is a testability and review argument — but only if single-carrier-thread execution is an enforced invariant rather than an accident of current configuration.

## Two kinds of race A **data race** is a memory-model concept: two concurrent accesses to the same location, at least one a write, without ordering between them. Consequences are ugly and low level — torn values, indefinitely stale reads, reordering surprises. A **race condition** in the broader sense is a correctness bug where the outcome depends on interleaving. An **atomicity violation** is the common shape: a sequence of operations that must appear indivisible is interrupted partway. Cooperative single-threaded scheduling eliminates the first and leaves the second fully intact. Candidates who answer "single-threaded, so no races" are answering only the first. ## What single-threaded cooperative scheduling does guarantee Between two suspension points, a task runs to completion with respect to other tasks on the same carrier thread. That is a real and useful guarantee: you can mutate a shared map across many lines with no lock, provided none of those lines suspends. It also means interleaving points are *finite and visible* — a reviewer can find them by looking for suspension markers, unlike preemptive threading where every instruction is a potential switch. ## What it does not guarantee **Check-then-act across a pause.** The classic: ``` if (!cache.contains(key)): # check value = await fetch(key) # <-- everything else runs here cache.put(key, value) # act on a stale check ``` Two tasks can both pass the check and both fetch. You get duplicate work, duplicate side effects, or a last-writer-wins overwrite. The fix is the same as with locks: record the in-flight operation itself (store a pending placeholder before awaiting so the second caller joins it) or serialise with an async-aware mutual-exclusion primitive. **Broken invariants observed at a pause.** If you take an object out of a consistent state, suspend, then finish the update, other tasks observe the inconsistent middle. Rule: never suspend while an invariant is broken; do all awaits before or after the mutation, not inside it. **Re-entrancy.** A handler that suspends can be entered again for a new event before the first completes. Any per-handler mutable state (a module-level accumulator, a "currently processing" flag) is now shared between two live activations. **Ordering assumptions.** Completions are delivered in whatever order the underlying operations finish, not in the order they were started. Code that assumes responses arrive in issue order is racing. **External state.** Time-of-check to time-of-use against a database, a file, or another service is unaffected by your scheduler. Only the external system's own atomicity (transactions, conditional writes, compare-and-set) helps there. **Cancellation.** Cancelling a task usually takes effect at its next suspension point, so a task can be cancelled in the middle of a multi-step update. Anything that must complete atomically has to be shielded or made idempotent. ## When multiple carrier threads enter Many runtimes multiplex tasks over a pool of carrier threads. Then two tasks genuinely execute at the same instant, and the memory-model hazards return in full: you need proper synchronisation and safe publication for shared mutable state, exactly as with threads. Worse, code may have been written and tested under a single-threaded assumption and only break when the pool is widened or when the deployment enables parallelism. Treating "single-threaded" as an architectural invariant means enforcing and documenting it, not assuming it. Even on one carrier thread, a task can resume on a *different* thread if the runtime allows, which invalidates thread-affine constructs: thread-local storage read after an await may belong to another logical task, and a lock acquired before a suspension may be released by a different thread than acquired it — undefined or illegal for many lock implementations. So: do not hold a thread-affine lock across a suspension point. ## The mental model to carry into an interview Say it in one line: *cooperative scheduling turns "anywhere" into "here", it does not turn "anywhere" into "nowhere"*. It shrinks the set of interleaving points to something you can enumerate and audit, which is a genuine engineering win — and the discipline it demands is to look at every suspension point and ask what invariant is currently broken and what other task could touch it.

  • Two tasks both miss a cache and both start the same expensive fetch. How do you fix that without a thread lock?
    Store the in-flight operation, not just the result: insert a pending entry (a future/promise for the value) into the cache before awaiting, so the second caller finds it and awaits the same operation. This is single-flight deduplication. Remove or replace the entry on completion, and make sure a failed fetch clears it so failures are not cached forever.
  • Why is holding a thread-affine lock across a suspension point dangerous?
    The task may resume on a different carrier thread, so the release would happen on a thread that never acquired it, which many lock implementations treat as undefined behaviour. It also blocks other tasks on that thread for the whole duration of the awaited operation, converting a short critical section into an unbounded one. Use an async-aware mutual-exclusion primitive that suspends rather than blocks.

Cooperative switching is a kitchen where cooks only swap stations when they choose. No two hands touch the same pot at once, but if you step away mid-recipe with the sauce half-salted, the next cook still finds it in a state your recipe never intended.

saying these in an interview costs you the question

  • "It's single-threaded, so there are no race conditions" — only data races are excluded
  • Assuming a value read before an await is still valid after it
  • Thinking cooperative scheduling makes multi-step updates atomic
  • Holding a thread-affine lock or relying on thread-local state across a suspension point
  • Assuming completions arrive in the order operations were started

context