In Ruby, what does fork return in the parent and the child, with and without a block, and why does forking bypass the GVL?
answer
- one call, two processes
- parent gets the child's pid
- child gets nil, not 0
- block form: child exits after block
- each process has its own GVL
basics
~20 sfork copies the running Ruby process. Without a block it returns the child's pid in the parent and nil in the child; with a block, only the child runs it and then exits. Each process has its own GVL, so children run Ruby in parallel.
solid answer
~40 s`fork` (also `Process.fork`) copies the current Ruby process. Without a block the call returns in both processes: in the parent it returns the child's pid, in the child it returns `nil`, so code branches on `if pid`. With a block, only the child runs the block and then terminates, with status 0 if the block finished normally, while the parent gets the pid straight away. Only the thread that called `fork` exists in the child; other threads are not copied. Because each process has its own interpreter and its own GVL, several forked children run CPU-bound Ruby code on different cores at once, which threads in one CRuby process cannot. The parent must still collect each child with `Process.wait`, and `fork` is unavailable on some platforms, such as Windows.
code
ruby · 8 linespids = 4.times.map do |i|
fork do
total = (1..5_000_000).sum { it * i }
exit!(total.even? ? 0 : 1)
end
end
pids.each { |pid| Process.wait(pid) }go deeper
Recall the two return values of fork, pid in the parent and nil in the child, and that the block form runs only in the child.
Explain why forked workers run Ruby in parallel under separate GVLs, and list what a child inherits: memory, open descriptors, at_exit handlers, one thread.
Design a forking worker model: load code before forking, restart threads in children, end children with exit! where handlers must not rerun, and reap every child.
Weigh processes against threads and Ractors for a Ruby service: memory per worker, isolation from crashes, platform support and operational complexity.
## What fork does `Kernel#fork` (the same method as `Process.fork`) asks the operating system to duplicate the running process. The new **child** starts as a copy of the **parent** at the moment of the call: the same loaded classes, the same objects, the same open files and sockets. From then on they are separate processes with separate memory; a change in one is invisible in the other. ## Return values: with and without a block | Form | In the parent | In the child | |---|---|---| | `pid = fork` | returns the child's pid (an Integer) | returns `nil`, then continues after the call | | `pid = fork { work }` | returns the child's pid immediately | runs the block, then exits; status 0 if the block finished normally | The block form is the one to prefer: it makes it impossible for the child to fall through into code meant for the parent. ```ruby pid = fork if pid puts "parent #{Process.pid} started child #{pid}" Process.wait(pid) else puts "child #{Process.pid} doing the work" exit!(0) end ``` A detail that trips people coming from C: the underlying system call returns 0 in the child, but Ruby's `fork` returns `nil` there. ## Why fork gives real parallelism on CRuby CRuby lets only one thread per interpreter run Ruby code at a time, because of the GVL. A forked child is a **whole new interpreter** with its own GVL, so: - N children can execute CPU-bound Ruby code on N cores simultaneously. - A crash or memory leak in one child cannot corrupt the others' objects. - Nothing is shared by accident, so no `Mutex` is needed between processes; results must come back through the exit status, a pipe, a file or another explicit channel. This is why preforking servers and job runners fork workers instead of relying on threads for CPU-heavy work. ```ruby slices = (1..8_000).each_slice(1_000).to_a pids = slices.map do |ids| fork { ids.each { |id| JobRunner.perform(id) } } end pids.each { |pid| Process.wait(pid) } ``` ## What the child inherits, and what it does not 1. **Memory**: a copy of every object, shared physically with the parent until one of them writes to a page. 2. **Open file descriptors**: files, pipes and sockets stay open in both processes and point at the same underlying resource. 3. **Threads**: only the thread that called `fork` exists in the child. A background thread started in the parent (a heartbeat or metrics flusher) is simply absent there. 4. **`at_exit` handlers**: registered handlers run when the child exits too, unless the child ends with `exit!`, which skips them. ## Getting a result back A child cannot hand an object back through a variable, because its memory is separate. The exit status carries a small code (0 to 255); anything richer needs an explicit channel, and the simplest is a pipe created before the fork: ```ruby reader, writer = IO.pipe pid = fork do reader.close writer.write(JobRunner.summarise(1..1_000).to_s) writer.close end writer.close summary = reader.read Process.wait(pid) ``` Each side closes the end it does not use, so the parent's `read` sees end-of-file once the child finishes writing. ## Practical limits - **Platform**: `fork` does not exist on every platform; on Windows `Process.respond_to?(:fork)` is false, and `spawn` is the portable way to start another process. - **Cleanup**: every child must be collected with `Process.wait` or `Process.detach`, or it lingers as a zombie after it exits. - **Cost**: a fork is cheap compared with booting a new interpreter, which is exactly why you load the application first and fork afterwards.
- A parent process starts a background thread that flushes metrics, then forks workers. What happens to that thread in the workers?It does not exist there. `fork` copies only the thread that called it, so the child has one thread and the flusher is gone, even though the objects it used were copied. Restart such threads in the child after the fork, for example from a fork hook.
- Why might a forked child run the parent's `at_exit` handlers, and how do you prevent it?The child is a copy of the parent, including the list of registered `at_exit` handlers, and a normal exit runs them, so a handler that flushes a file or closes a shared connection runs twice. Ending the child with `exit!` skips `at_exit` handlers; the fork documentation recommends it for that reason.
Forking is photocopying a filled-in notebook: both copies start identical and each person can write on their own, but notes in one never appear in the other, and each copy has its own pen, the GVL.
saying these in an interview costs you the question
- Says fork returns 0 in the child, as in C
- Thinks the parent also runs the block given to fork
- Expects background threads to keep running in the child
- Believes a variable changed in the child is visible to the parent
- Thinks forked children still share one GVL