Your service spawns one virtual thread per request, but a downstream API rate-limits you and starts failing. How do you bound concurrency correctly?
answer
- Don't shrink the thread count to throttle a dependency
- Semaphore permits = downstream limit
- acquire() before the call, release() in finally
- Decouple #tasks (per-request VT) from #resource-users (permits)
- Per-dependency semaphores; consistent lock order avoids deadlock
basics
~20 sKeep one virtual thread per request, but guard the downstream call with a Semaphore sized to what the API allows. Each thread acquires a permit before calling and releases it after, so only that many calls happen at once — without limiting how many threads exist.
solid answer
~50 sThe mistake would be to 'fix' this by switching to a fixed thread pool, which re-imposes a global concurrency cap and throws away the virtual-thread benefit. The right tool is a Semaphore sized to the downstream limit. A semaphore holds N permits; each virtual thread calls acquire() before the rate-limited call and release() in a finally block after. At most N threads hold permits and make the call concurrently; the rest park cheaply on acquire() — which, on a virtual thread, unmounts and costs almost nothing. This separates the two concerns the pool conflated: how many tasks exist (as many as requests, each its own virtual thread) versus how many may touch the constrained resource (N). You can have a different semaphore per downstream dependency, sizing each to its own limit, instead of one coarse pool throttling everything. For multiple resources, acquire permits in a consistent order to avoid deadlock.
code
java · 19 lines// One virtual thread per request; bound each downstream resource separately.
Semaphore paymentsApi = new Semaphore(50); // third-party rate limit
Semaphore database = new Semaphore(200); // DB connection capacity
try (var executor = Executors.newVirtualThreadPerTaskExecutor()) {
for (var req : incoming) {
executor.submit(() -> {
var parsed = parse(req); // full concurrency, unthrottled
paymentsApi.acquire(); // cheap park if no permit
try { charge(parsed); }
finally { paymentsApi.release(); } // always release
database.acquire();
try { persist(parsed); }
finally { database.release(); }
});
}
}go deeper
Knows a Semaphore limits how many threads do something at once and that you acquire before and release after the protected call.
Explains why a fixed pool is the wrong throttle and how acquire/release in finally bounds the downstream call while keeping one virtual thread per task.
Separates task count from resource concurrency, sizes a semaphore per dependency, handles release-in-finally and tryAcquire-based load shedding, and notes deadlock from inconsistent multi-permit ordering.
Designs end-to-end backpressure: per-dependency limits, fail-fast vs queueing policy, distinction between concurrency and rate limits, and how this composes with structured concurrency and observability across a service.
## The scenario You adopted virtual threads and now spawn **one virtual thread per request**, scaling to huge concurrency. But one of your dependencies — say a third-party API or a database with a small connection pool — can only handle, say, **50 concurrent calls**. Under load your virtual threads happily fire thousands of simultaneous calls and the dependency starts rejecting or timing out. You need **backpressure**: a way to limit how many calls hit that dependency at once. ## The wrong fix: go back to a fixed pool The tempting fix is to run requests on a `newFixedThreadPool(50)`. This *does* cap downstream calls at 50 — but it also caps **everything** at 50, including the parts of each request that *don't* touch the slow dependency (parsing, other services, CPU work). You've throttled the whole request pipeline to the bottleneck of one dependency and discarded the scalability virtual threads gave you. **Limiting threads is too coarse a knob.** ## The right fix: a Semaphore on the resource A **`Semaphore`** is a counter of **permits**. `acquire()` takes a permit, blocking if none are free; `release()` returns one. Create a semaphore with as many permits as the dependency allows, and have each task **acquire before the call and release after** (always in a `finally`, so a thrown exception still releases): ```java Semaphore apiLimit = new Semaphore(50); ... apiLimit.acquire(); try { result = callRateLimitedApi(req); } finally { apiLimit.release(); } ``` Now at most 50 virtual threads hold a permit and call the API at once. The other threads block in `acquire()` — and because they are **virtual** threads, blocking is cheap: the JVM unmounts them from their carriers, so you can have ten thousand threads waiting on the semaphore using almost no OS resources. The moment a permit frees, one waiter proceeds. ## Why this is strictly better - **Separates the two concerns the pool conflated.** *How many tasks exist* (one virtual thread per request) is decoupled from *how many may use a given resource* (the semaphore's permit count). The non-throttled parts of each request run at full concurrency. - **Per-dependency sizing.** You can hold a *different* semaphore for each downstream resource — 50 for the payments API, 200 for the DB, 10 for a fragile legacy service — instead of one pool that throttles everything to its smallest bottleneck. - **Cheap waiting.** Parked virtual threads cost almost nothing, so an overflow of waiters is harmless backpressure rather than resource exhaustion. ## Caveats - **Always release in `finally`**, or a failing call permanently leaks a permit and the limit shrinks over time. - **Multiple semaphores → deadlock risk.** If a task must hold two permits at once, always acquire them in the **same global order** across all tasks, or use `tryAcquire` with a timeout. - A semaphore bounds *concurrency*, not *rate* (calls per second). For a true rate limit (e.g. 100 req/s regardless of duration) use a rate limiter; a semaphore is the right tool when the limit is on *simultaneous* in-flight calls. - Consider `tryAcquire(timeout)` to fail fast and shed load instead of queueing unboundedly. ## Summary Don't throttle by shrinking the thread count — that re-creates the limit virtual threads removed. Keep one cheap virtual thread per task and bound each constrained downstream resource with its own Semaphore, acquiring before and releasing (in `finally`) after the call.
- A fixed thread pool of 50 also caps downstream calls at 50 — why is a Semaphore(50) better?The pool caps the entire request pipeline at 50, including the work that never touches the slow dependency, discarding virtual-thread scalability. A semaphore bounds only the guarded call, letting everything else run at full per-request concurrency, and lets you size each dependency independently.
- What goes wrong if you call semaphore.release() only on the success path?When the guarded call throws, the permit is never returned, so the available permit count permanently shrinks; over time the effective limit drops toward zero and all tasks block forever. Release must be in a finally block.
saying these in an interview costs you the question
- Reaching for a fixed thread pool to rate-limit — it throttles the entire pipeline, not just the slow dependency
- Calling release() outside finally so exceptions leak permits
- Conflating a concurrency limit (semaphore) with a requests-per-second rate limit
- Using one global semaphore/pool for all dependencies instead of sizing each independently
- Acquiring multiple semaphores in inconsistent order, risking deadlock