skip to content

Virtual Threads: Motivation & Model

Platform threads are expensive, which capped thread-per-request throughput and pushed people into async code and its colored-function complexity. Virtual threads are JVM-scheduled, mount onto carrier threads, park their stacks on the heap, and let you write blocking code again — the why behind Loom that interviewers want first.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

4

Why were virtual threads added to Java? What problem do they solve?

level: middleimportance: must knowfreq 78%

answer

  1. Platform thread = 1:1 OS thread, expensive, ~few thousand max
  2. Thread-per-request wall: blocked threads waste a scarce resource
  3. Async fixes scale but adds 'colored function' callback complexity
  4. Loom inverts it: cheap threads, write blocking code again
  5. Virtual thread runs by mounting on a carrier platform thread

basics

~20 s

Regular Java threads map one-to-one to OS threads, which are expensive and limited, so you can only have a few thousand. Virtual threads are cheap, so you can have millions and write simple blocking code that scales.

solid answer

~50 s

Before virtual threads, each Java thread was a platform thread backed one-to-one by an OS thread. OS threads are heavyweight (large stacks, kernel scheduling), so a JVM can only run a few thousand. The common server model is thread-per-request: one thread per in-flight request. With a hard cap of a few thousand threads, throughput on I/O-bound workloads is limited not by CPU but by the thread count, since most threads sit blocked on I/O. To scale further, developers switched to async/reactive frameworks, which are non-blocking but force you to break logic into callbacks/futures and lose readable, debuggable straight-line code. Virtual threads (Project Loom) make threads so cheap you can have millions, so thread-per-request scales again. You write ordinary blocking code that reads top-to-bottom, and the JVM handles the efficiency, eliminating the need for async-style 'colored' functions.

go deeper

for a junior

Knows virtual threads are cheap and you can create a great many of them, and that they let you write normal blocking code instead of callbacks.

for a middle

Explains the platform-thread 1:1-OS-thread cost, the thread-per-request scalability wall, and that Loom's goal is cheap threads so blocking code scales.

for a senior

Frames the async/reactive trade-off (the 'colored function' tax) precisely and explains mount/unmount onto carriers as the mechanism that frees the OS thread on blocking.

for a principal

Reasons about when virtual threads do and don't help (I/O-bound vs CPU-bound), migration implications for existing reactive stacks, and ecosystem effects (libraries that must not pin carriers).

## The starting point: what a thread is A **thread** is an independent path of execution within a program. Each thread has its own **stack** (the memory holding local variables and the chain of in-progress method calls) and a **program counter** (where it is in the code). Multiple threads let a program do several things at once. ## Platform threads = OS threads (the old model) Before Java 21, every Java thread was a **platform thread**: a thin wrapper over an **operating-system (OS) thread**, mapped **one-to-one**. The OS thread is the real unit the operating system's **scheduler** (the OS component that decides which thread runs on a CPU core, and when) manages. OS threads are **expensive**: - Each reserves a large fixed **stack** (commonly ~1 MB), so memory caps how many you can have. - Creating/destroying one is a **system call** (a request into the OS kernel) — relatively slow. - **Context switching** (the OS saving one thread's state and loading another's so they can share a core) costs CPU. In practice a JVM tops out around a **few thousand** platform threads. ## Thread-per-request and the scalability wall A classic server uses the **thread-per-request** model: each incoming request gets its own thread that runs the whole request top to bottom — query a database, call another service, format a response. This code is easy to write, read, and debug, because it's plain sequential code. Most server work is **I/O-bound**: the thread spends most of its life **blocked** — parked, waiting for the network or disk, using no CPU. With only a few thousand threads, you can only have a few thousand requests in flight, even though the CPU is mostly idle. The bottleneck is the **thread count**, not the hardware. That is the scalability wall. ## The async/reactive escape — and its tax To break the wall, developers moved to **asynchronous / reactive** programming (CompletableFuture, reactive streams, etc.). Here a small pool of threads is **never blocked**; instead of waiting, a task registers a **callback** to run when the I/O finishes, and the thread moves on to other work. This scales to many concurrent requests on few threads. The cost is **complexity**. Sequential logic shatters into chained callbacks/futures. You can't use ordinary loops, try/catch, or step-through debugging across an async boundary; **stack traces** become uninformative. This is the **'colored function' problem**: once a method is async, everything that calls it must also be async — the 'color' spreads through the codebase, and you can't freely mix blocking and non-blocking code. ## Loom's idea: make threads cheap instead **Project Loom** asked the inverse question: instead of avoiding threads, make them so cheap that thread-per-request scales. The result is the **virtual thread** (Java 21, JEP 444). A **virtual thread** is a thread scheduled by the **JVM**, not the OS. It is **not** permanently bound to an OS thread. To actually run, a virtual thread is **mounted** onto a **carrier thread** — a regular platform (OS) thread drawn from a pool. When the virtual thread hits a **blocking** operation (e.g. socket read), the JVM **unmounts** it: its stack is copied to the **heap** (the JVM's general-purpose memory region), and the carrier thread is freed to run other virtual threads. When the I/O completes, the virtual thread is **remounted** (possibly onto a different carrier) and resumes. Because an idle virtual thread costs only a little heap, you can have **millions**. ## The payoff You go back to writing **simple, blocking, sequential code** — readable, debuggable, with real stack traces — and still get the scalability that previously required async. The JVM does the unmount/remount dance for you, so a blocking call no longer wastes a precious OS thread.

  • What is the 'colored function' problem and how do virtual threads avoid it?
    In async code, once a function is asynchronous, every caller must also be async — the 'color' propagates and you can't freely mix blocking and non-blocking code. Virtual threads let you write ordinary blocking calls, so no color spreads; any method can block without forcing its callers to change.
  • Are virtual threads faster than platform threads for a single CPU-bound task?
    No. A single task runs at roughly the same speed; virtual threads add scheduling overhead if anything. Their benefit is allowing a huge number of concurrent, mostly-blocked tasks cheaply — they raise concurrency/throughput for I/O-bound workloads, not per-task compute speed.

saying these in an interview costs you the question

  • Claiming virtual threads make CPU-bound code faster (they help blocking/I/O-bound concurrency, not raw compute)
  • Saying a virtual thread is just a bigger thread pool of platform threads
  • Thinking each virtual thread has its own OS thread
  • Saying virtual threads remove the need for the OS scheduler entirely (carriers are still OS-scheduled)

context

open as a page

How does a virtual thread actually execute? Explain mounting, unmounting, and carrier threads.

level: seniorimportance: must knowfreq 70%

basics

~20 s

A virtual thread doesn't have its own OS thread. To run, the JVM mounts it onto a carrier (a real platform thread). When it blocks, the JVM unmounts it and stores its stack on the heap, freeing the carrier for other virtual threads.

open as a page

How do you create virtual threads, and why is creating millions of them cheap?

level: juniorimportance: should knowfreq 58%

basics

~10 s

Use Thread.ofVirtual().start(...) or Executors.newVirtualThreadPerTaskExecutor(). They're cheap because a virtual thread doesn't reserve a big OS-thread stack — its stack lives on the heap and grows as needed, so millions fit.

open as a page

When do virtual threads NOT help, and what are the main pitfalls when adopting them?

level: principalimportance: should knowfreq 52%

basics

~20 s

Virtual threads help concurrency for blocking/I/O-bound work, not raw CPU speed. They don't help CPU-bound tasks, and you can lose the benefit through pinning (e.g. blocking in synchronized) or by overloading a small downstream resource.

open as a page