skip to content

A teammate claims 'every suspend call allocates objects and is expensive.' Using your knowledge of the CPS state-machine transform, when is that true and when is it false?

level: seniorimportance: should knowfreq 35%

answer

  1. Fast path = no allocation, falls through when(label)
  2. Allocate-once lazily at first real suspension
  3. Cost is in actual suspension + dispatcher hop
  4. suspend marker itself is nearly free
  5. Avoid needless withContext hops, not suspend

basics

~10 s

It's only partly true. If a suspend call finishes without pausing, it runs almost like a normal function call with no extra allocation. Allocation and real cost happen only when it actually suspends.

solid answer

~50 s

The CPS transform has a fast path and a slow path. On the **fast path** — when a callee completes without suspending (returns a real value, not `COROUTINE_SUSPENDED`) — the call is essentially a plain method call with one extra parameter; no continuation object is allocated and execution falls straight through the `when (label)` to the next state. The continuation is allocated **lazily, once, only at the first real suspension**, and reused for the rest of that call. So a `suspend` function that never hits an actual await (e.g., a cache hit, or a `suspend` wrapper around pure work) is nearly free. The cost appears only when suspension truly occurs: heap-allocating/spilling locals, returning the sentinel up the chain, and later re-entering via `resumeWith`/`invokeSuspend`, possibly with a dispatch hop. The takeaway: don't fear marking functions `suspend`; fear unnecessary actual suspensions and dispatcher hops.

code

kotlin · 4 lines
kotlin
suspend fun lookup(id: Int): User {
    cache[id]?.let { return it }      // fast path: no real suspension, no continuation alloc
    return repository.load(id)        // slow path: real suspension allocates/spills once
}

go deeper

for a junior

Recognizes suspend isn't automatically heavy and that not pausing is cheap.

for a middle

Explains the fast-path fall-through vs slow-path allocation and the COROUTINE_SUSPENDED check.

for a senior

Articulates allocate-once-lazily, identifies dispatcher hops as the real cost, and gives correct profiling-oriented guidance.

for a principal

Reasons about allocation/GC pressure, escape analysis limits, and architectural trade-offs of granular suspend boundaries vs dispatch overhead.

## The claim, deconstructed 'Every suspend call allocates and is expensive' conflates two different paths the compiler generates. ## Fast path: completes without suspending When a suspending callee returns a **real value** rather than the `COROUTINE_SUSPENDED` sentinel: - No continuation object is allocated for the caller at that point. - The `if (r == COROUTINE_SUSPENDED) return ...` check fails, so control **falls through** to the next branch of `when (label)` within the same invocation. - The only overhead versus a normal call is: one extra `Continuation` parameter, the sentinel comparison, and the `when` dispatch — all cheap, JIT-friendly, and non-allocating. A suspend function that wraps pure computation, or returns early (cache hit), or whose callees all complete synchronously, runs at essentially normal-call speed. ## Slow path: an actual suspension Real suspension is where cost lives: - The continuation object is **allocated once** (lazily, at the first suspension) and locals that cross the suspension are spilled into its fields. - `COROUTINE_SUSPENDED` is returned up the whole call chain, unwinding the native stack. - On completion, `resumeWith` -> `invokeSuspend` re-enters the body; if a dispatcher is involved (e.g., `Dispatchers.Default`/`IO`), there may be a **thread hop** and scheduling latency. ## Practical guidance ```kotlin // Cheap: usually a fast-path return, no real suspension on cache hit suspend fun get(key: String): V = cache[key] ?: loadFromDb(key) // only loadFromDb may truly suspend // Watch out: needless dispatcher hops add real cost withContext(Dispatchers.Default) { trivialPureWork() } // hop may cost more than the work ``` - **Do** mark functions `suspend` freely for composability; the marker alone is nearly free. - **Avoid** unnecessary `withContext`/dispatcher switches around trivial work — the hop, not the suspend keyword, is the cost. - **Remember** allocate-once: a single suspend function call that suspends many times still allocates roughly one continuation, not one per point. ## Terms to cite `COROUTINE_SUSPENDED`, fast path vs slow path, lazy/allocate-once continuation, `withContext`, dispatcher hop, `invokeSuspend`.

  • If a suspend function suspends at five different points during one call, roughly how many continuation objects does it allocate?
    About one: the continuation is allocated once at the first suspension and reused for the remaining states; it does not allocate one per suspension point.
  • Where does a `withContext(Dispatchers.IO){ ... }` add cost even if the body is trivial?
    It forces a suspension and a dispatcher hop — scheduling onto another thread and back — which can dominate the cost when the wrapped work itself is tiny.

Like a tollbooth with an open gate: if traffic flows (no suspension) you barely slow down; you only pay the toll (allocate + reschedule) when you actually have to stop.

saying these in an interview costs you the question

  • Asserts every suspend call always allocates
  • Doesn't distinguish fast path from slow path
  • Thinks one continuation is allocated per suspension point
  • Blames the suspend keyword instead of dispatcher hops
  • Recommends avoiding suspend functions for performance reasons broadly

context