You need to call a flaky external API but must never have more than 5 in-flight requests at once across many coroutines. How do you implement this with coroutine sync primitives?
answer
- Semaphore(5) + withPermit { } caps concurrency at 5
- ONE shared semaphore instance — not per call
- Mutex ≈ Semaphore(1) = serialized, too strict
- extra coroutines SUSPEND, not block
- withPermit releases on throw/cancel
basics
~10 sCreate one shared Semaphore(5) and wrap each call in semaphore.withPermit { ... }. Coroutines that arrive when all 5 permits are taken suspend until a permit is freed.
solid answer
~40 sUse `kotlinx.coroutines.sync.Semaphore(permits = 5)` shared across all callers, and wrap each request in `semaphore.withPermit { callApi() }`. `withPermit` suspendingly `acquire()`s a permit, runs the block, and `release()`s in a `finally` — so it's exception- and cancellation-safe. When all 5 permits are held, additional coroutines **suspend** (not block) until one is released, capping concurrency at 5 regardless of how many you `launch`. A `Mutex` (≈ `Semaphore(1)`) would serialize to one-at-a-time, which is too strict here. The semaphore must be a single shared instance — creating one per call gives each call its own 5 permits and enforces nothing. This is the idiomatic non-blocking rate/concurrency limiter; for time-based rate limiting you'd combine it with delays.
code
kotlin · 11 linesimport kotlinx.coroutines.*
import kotlinx.coroutines.sync.Semaphore
import kotlinx.coroutines.sync.withPermit
val gate = Semaphore(permits = 5)
suspend fun limitedCall(url: String) = gate.withPermit { httpGet(url) }
suspend fun run(urls: List<String>) = coroutineScope {
urls.map { async { limitedCall(it) } }.awaitAll() // ≤5 in flight
}go deeper
Reaches for Semaphore(5) + withPermit and knows it limits concurrency.
Explains shared-instance requirement, suspend-vs-block, and Mutex-vs-Semaphore tradeoff.
Distinguishes concurrency cap from rate limiting and handles cancellation/leak correctness.
Designs the limiter as a reusable abstraction, considers fairness, backpressure, and integration with flatMapMerge/dispatcher limiting alternatives.
## Goal Bound *concurrent* access to a resource at N, no matter how many coroutines exist. This is exactly what a **counting semaphore** does. ## `Semaphore` in kotlinx.coroutines `Semaphore(permits: Int, acquiredPermits: Int = 0)` holds up to `permits` tokens: - `acquire()` — `suspend`; takes a permit, suspending if none are free. - `release()` — returns a permit, resuming a waiter. - `tryAcquire()` — non-suspending; returns `false` if none free. - `withPermit { }` — inline helper: `acquire(); try { block() } finally { release() }`. ## The implementation ```kotlin import kotlinx.coroutines.* import kotlinx.coroutines.sync.Semaphore import kotlinx.coroutines.sync.withPermit class ApiClient { private val gate = Semaphore(permits = 5) // ONE shared instance suspend fun call(url: String): Response = gate.withPermit { // ≤5 concurrent here httpGet(url) } } suspend fun fanOut(client: ApiClient, urls: List<String>) = coroutineScope { urls.map { async { client.call(it) } }.awaitAll() } ``` Even if you `async` 1000 URLs, at most 5 run inside the block at any instant; the rest **suspend** at `acquire()` and free their threads. ## Why not a `Mutex`? A `Mutex` is effectively `Semaphore(1)` — it would force strictly one request at a time. We want *five*, so we need a real semaphore. ## Common bug: per-call semaphore ```kotlin // WRONG — each call gets its own 5 permits → no global limit suspend fun call(url: String) = Semaphore(5).withPermit { httpGet(url) } ``` The semaphore must be shared (a field / singleton) to enforce a global cap. ## Safety properties - `withPermit` releases the permit on normal return, exception, and cancellation (inline `finally`). - Cancellation while waiting in `acquire()` throws `CancellationException` and no permit is leaked. ## Beyond concurrency caps For *requests-per-second* limiting, combine the semaphore with `delay()`/a token-refill scheme, or use a dedicated rate limiter — a plain semaphore caps *simultaneous* in-flight work, not throughput over time.
- Does a `Semaphore` give you requests-per-second rate limiting?No — it caps *simultaneous* in-flight work. For throughput over time you need delays or a token-bucket refill on top.
- What's the bug if you write `Semaphore(5).withPermit { ... }` inside the call function?A fresh semaphore per call means every call has its own 5 permits, so there's no global limit at all. It must be a shared field/singleton.
saying these in an interview costs you the question
- Using a `Mutex` and claiming it allows 5 concurrent
- Creating a new `Semaphore` inside each call
- Thinking the extra coroutines block threads instead of suspending
- Confusing concurrency cap with rate-per-second limiting
- Forgetting `withPermit` releases on exception/cancellation