skip to content

Threads & the GVL

Thread.new, join and value run work concurrently, but the GVL lets only one thread execute Ruby at a time while blocking I/O releases it. Interviewers ask when threads really speed code up.

on this pageshow

explore

questions

5

In CRuby, what is the GVL, and when do Ruby threads actually make a program run faster?

level: juniorimportance: must knowfreq 82%

answer

  1. one thread runs Ruby at a time
  2. one lock per Ractor
  3. blocking I/O and sleep release it
  4. waits overlap, computation does not
  5. RUBY_THREAD_TIMESLICE default 100 ms

basics

~20 s

The GVL (Global VM Lock) lets only one thread per Ractor execute Ruby code at a time. Threads speed up I/O-bound work, because blocking I/O and sleep release the lock, but not CPU-bound Ruby code.

solid answer

~50 s

The **GVL** (Global VM Lock, sometimes still called the GIL) is a lock inside CRuby: a thread must hold it to execute Ruby code, and each Ractor has one, so in a normal single-Ractor app only one thread runs Ruby at any instant. A thread gives the lock up when it blocks: reading a socket, waiting for a file, `sleep`, `Thread#join`, or a C extension that explicitly runs without it. A running thread is also switched out when its time slice (100 ms by default) ends. So threads help when the work is mostly **waiting** — fetching 50 exchange-rate pages with 50 threads overlaps the network waits and finishes in roughly the time of the slowest page. For **CPU-bound** Ruby code, four threads take about as long as doing the work serially; true parallelism needs processes or Ractors.

code

ruby · 10 lines
ruby
require "benchmark"

cpu = -> { 3_000_000.times.sum { |i| i * i } }
io  = -> { sleep 0.5 }

Benchmark.realtime { 4.times { cpu.call } }                    # ~ T
Benchmark.realtime { 4.times.map { Thread.new(&cpu) }.each(&:join) } # still ~ T

Benchmark.realtime { 4.times { io.call } }                     # ~ 2.0 s
Benchmark.realtime { 4.times.map { Thread.new(&io) }.each(&:join) }  # ~ 0.5 s

go deeper

for a junior

Recall that only one thread per Ractor runs Ruby code at a time, and that blocking I/O and sleep hand the lock to another thread.

for a middle

Explain the release points (blocking I/O, sleep, join, time-slice expiry, GVL-free C code) and predict wall time for I/O-bound versus CPU-bound thread pools.

for a senior

Show you can diagnose a threaded service: measure how much time is blocked versus computing, spot CPU work starving I/O threads, and move CPU work to processes.

for a principal

Frame the choice between threads, processes and Ractors as a trade of memory, isolation and CPU parallelism across the deployment, not as a language limitation.

## What the GVL is CRuby, the reference Ruby interpreter, maps every `Thread` to a native operating-system thread, but it protects its virtual machine with one lock: the **Global VM Lock (GVL)**, also called the Giant VM Lock. A thread may execute Ruby code only while it holds the GVL. Since Ruby 3.0 each **Ractor** has its own lock, but almost every application runs a single Ractor, so in practice only **one thread executes Ruby code at a time** in the whole process. The lock exists so the interpreter, the garbage collector and the core classes (`String`, `Array`, `Hash`) do not need a fine-grained lock on every object access. The price is that Ruby threads give you **concurrency** (work interleaves) but not **parallelism** (Ruby work on several cores at once). ## When a thread gives the lock up A thread releases the GVL in a few well-defined situations: - **Blocking I/O**: reading or writing a socket, pipe or file that is not ready. `Net::HTTP` waiting for a response, a database driver waiting for a query result, and `IO#read` on a slow pipe all release it. - **Sleeping and waiting**: `Kernel#sleep`, `Thread#join`, `Thread#value`, and waiting on a lock or queue. - **C extensions that opt out**: native code can call `rb_thread_call_without_gvl` around work that touches no Ruby objects, so that work can run in parallel with Ruby threads. - **The end of a time slice**: a thread that computes without blocking is asked to yield when its quantum expires, 100 ms by default, so other threads are not starved. - **`Thread.pass`**: an explicit hint to let another thread run. When the blocking call returns, the thread must **reacquire** the GVL before it can run Ruby code again. ## I/O-bound versus CPU-bound work | Workload | What the threads spend time on | Effect of more threads on CRuby | |---|---|---| | Fetching pages, querying a database, calling an API | waiting for bytes from the network | waits overlap; wall time drops sharply | | Parsing JSON, rendering templates, number crunching in Ruby | executing Ruby code | no speed-up; context switches may add a little overhead | | Mixed: fetch, then parse | both | the fetch part overlaps, the parse part serialises | The rule of thumb: threads help in proportion to the share of time the work spends **blocked**. ## The 50 exchange-rate pages Suppose a job fetches 50 exchange-rate pages, each taking about 200 ms of network wait and a millisecond of parsing: ```ruby require "net/http" urls = (1..50).map { |i| URI("https://rates.example/page/#{i}") } threads = urls.map do |url| Thread.new(url) { |u| Net::HTTP.get(u) } end bodies = threads.map(&:value) ``` Serially the job takes about 50 × 200 ms = 10 seconds. With one thread per page, each thread releases the GVL while it waits, so all 50 requests are in flight together and the job finishes in roughly the time of the slowest response plus the parsing. If the same job instead computed a checksum over each page in pure Ruby for a full second, 50 threads would still take about 50 seconds, because only one of them can execute Ruby at a time. ## What the GVL does not promise The GVL protects the interpreter's own data structures; it says nothing about **your** invariants. A thread can be switched out between any two Ruby operations, so shared mutable state still needs synchronisation. Two practical consequences follow: 1. Do not treat "we have a GVL" as a thread-safety argument for application code. 2. Mixing CPU-heavy threads with I/O threads raises latency for the I/O threads: when a response arrives, its thread must wait for the lock, and a CPU-bound thread may hold it until its time slice ends. ## Getting real parallelism For CPU-bound Ruby work, CRuby offers: - **processes**: `fork` or a preforking server runs one interpreter, and so one GVL, per worker; - **Ractors**: each Ractor has its own GVL, at the cost of strict rules on which objects may be shared; - **C extensions** that release the GVL around their native work. ## Version notes Ruby 3.3 added an M:N thread scheduler (enabled on the main Ractor only with `RUBY_MN_THREADS=1`) and Ruby 3.4 added the `RUBY_THREAD_TIMESLICE` environment variable to change the 100 ms quantum. Neither lifts the rule that one Ractor runs one thread's Ruby code at a time; this answer assumes Ruby 4.0 on CRuby.

  • Why can a CPU-heavy thread slow down the response time of I/O threads in the same process?
    When an I/O thread's data arrives it must reacquire the GVL before running Ruby code. If a CPU-bound thread holds the lock, the I/O thread waits until that thread blocks or its time slice ends, 100 ms by default (tunable with `RUBY_THREAD_TIMESLICE` since Ruby 3.4). Latency for the I/O work rises even though total throughput looks fine.
  • How can native code in a C extension run in parallel with Ruby threads on CRuby?
    The extension wraps work that touches no Ruby objects in `rb_thread_call_without_gvl`, which releases the GVL for the duration and takes it back afterwards. Other threads run Ruby code meanwhile. The native code must not allocate or read Ruby objects while unlocked.
  • If threads do not parallelise CPU-bound Ruby, what does CRuby offer instead?
    Separate processes, through `fork` or a preforking server, give each worker its own interpreter and its own GVL. Ractors give each Ractor its own GVL but restrict which objects can be shared. For numeric hot spots, a C extension that releases the GVL also runs in parallel.

The GVL is a single microphone passed between speakers: only the holder can talk, but a speaker who is waiting for a reply from the audience hands it on, so several conversations can be waiting at once while only one voice is ever heard.

saying these in an interview costs you the question

  • CRuby threads execute Ruby code in parallel on every core.
  • Threads never help in Ruby because the GVL serialises everything.
  • A thread waiting on a socket read keeps holding the GVL until data arrives.
  • Four threads make a pure-Ruby CPU loop four times faster.
  • Enabling YJIT removes the GVL for compiled methods.
open as a page

In Ruby, what do Thread.new, Thread#join and Thread#value return, and what happens to threads the main program never joins?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Thread.new starts a block in a new thread and returns the Thread. join waits and returns that thread, or nil when its limit expires; value returns the block's result or re-raises its exception. Unjoined threads die when main exits.

open as a page

In Ruby, what happens when an exception escapes a thread's block, and what do report_on_exception and abort_on_exception change?

level: middleimportance: should knowfreq 50%

basics

~20 s

The thread dies while others keep running. report_on_exception (default true) prints the error to $stderr at that moment; join or value re-raise it later. abort_on_exception (default false) re-raises it in the main thread immediately instead.

open as a page

In Ruby, how does Thread.current[:key] differ from Thread.current.thread_variable_get(:key), and when does the difference bite?

level: middleimportance: should knowfreq 32%

basics

~20 s

Thread#[] and #[]= are fiber-local: each fiber on a thread has its own store, so code running in another fiber sees nil. thread_variable_get and thread_variable_set are truly thread-local, shared by every fiber on that thread.

open as a page

Why is Ruby's Timeout.timeout risky around code that holds connections or runs ensure blocks, and what do you use instead?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Timeout.timeout makes a watcher thread raise into your thread with Thread#raise, so the exception can land on any line: mid-write, inside ensure cleanup, while a lock is held. Prefer the timeouts built into the I/O call itself.

open as a page