skip to content

A preforking Ruby job runner's workers grow toward the parent's full size soon after fork; why does copy-on-write sharing erode, and how does Process.warmup help?

level: seniorimportance: should knowfreq 40%

answer

  1. shared until someone writes a page
  2. new objects land on shared pages
  3. promotion writes the object header
  4. warmup: GC, compact, promote
  5. call it once, before the first fork

basics

~20 s

After fork, pages stay shared only until a process writes to them, and CRuby writes to heap pages constantly: allocating into free slots, promoting survivors to the old generation, filling string caches. Process.warmup does that work once in the parent before forking.

solid answer

~50 s

A forked worker shares the parent's memory pages until either process writes to one; then the kernel copies that page. CRuby writes to heap pages far more often than the application code suggests: new objects are allocated into free slots on pages the parent left half-empty, an object that survives enough GCs is promoted to the old generation and gets a flag set in its header, and lazily computed data such as a string's coderange is written into the object. Each write turns a shared page into a private copy, so worker memory creeps toward the parent's. `Process.warmup`, called once in the parent at the end of boot and before the first fork, runs a major GC, compacts the heap, promotes all survivors to the old generation, precomputes string coderanges, frees empty pages and calls `malloc_trim` where available. Load all code before forking, and judge sharing by proportional or private memory, not RSS.

go deeper

for a junior

Remember that forked workers share the parent's memory at first and that pages are copied once someone writes to them.

for a middle

Explain which interpreter activities write to heap pages after fork: allocation into free slots, promotion to the old generation, lazily filled caches.

for a senior

Diagnose post-fork memory growth with PSS or private memory, eager-load before forking, call Process.warmup once in the parent, and recycle workers when needed.

for a principal

Set a memory budget per host: workers per machine, recycle thresholds, and when a threaded or single-process design would cost less than preforking.

## Copy-on-write in one paragraph After `fork`, the child does not get a physical copy of the parent's memory. Both processes point at the same pages, marked read-only; the first write to a page by either process makes the kernel copy that page for the writer. A preforking job runner relies on this: boot the application once, fork eight workers, and the large, read-mostly heap of classes, methods and constants is stored once. ## Why a CRuby worker keeps writing to shared pages The job code may never modify those objects, yet the interpreter does. The main Ruby-specific causes: - **Allocation into shared pages.** A heap page that the parent left partly empty has free slots. The worker allocates its new objects into those slots, and every allocation writes to a page that was shared. - **Promotion to the old generation.** CRuby's collector is generational: an object that survives several collections is promoted, and promotion sets a flag in the object's header. Objects that happen to reach that age in the worker are written to in every worker. - **Lazily filled caches in objects.** Some data is computed on first use and stored in the object. A string's coderange (whether it is ASCII-only, valid, or broken) is one; CRuby's own source notes that computing it early avoids mutating heap pages after a fork. - **Code loaded after fork.** Classes autoloaded or required on first use inside a worker are built separately in every worker and never shared. The symptom is characteristic: each worker's private memory climbs quickly in the first minutes after fork, then flattens. ## What Process.warmup does `Process.warmup` (added in Ruby 3.3) tells the VM that boot is finished. Its documentation says a pre-forking application should call it in the original process before the first fork, and that the exact work is implementation specific. On CRuby it: 1. **Runs a major GC**, so garbage from boot is gone before it is shared. 2. **Compacts the heap**, moving live objects together so pages are full and new allocations tend to go to fresh pages rather than into shared ones. 3. **Promotes all surviving objects to the old generation**, so the header write happens once, in the parent. 4. **Precomputes the coderange of all strings**, so that cache is written before the fork. 5. **Frees empty heap pages** and calls `malloc_trim` where available, returning memory the workers would otherwise inherit for nothing. It returns `true`. ```ruby # boot.rb for the job runner require_relative "config/environment" JobRunner.eager_load! Process.warmup WORKERS.times { fork { JobRunner.work_loop } } Process.waitall ``` ## Other levers - **Load everything before forking.** Eager-load application code and warm any caches the workers will read, so that work is done once and shared. - **Measure the right number.** RSS counts shared pages in every process, so summing workers' RSS overstates usage. Proportional set size (PSS) or private memory from the operating system shows what forking actually saved. - **Recycle workers.** If private memory still creeps, restarting a worker after N jobs or above a memory threshold resets it to a fresh copy of the warmed parent. ## Common mistakes | Mistake | Why it hurts | |---|---| | Calling `Process.warmup` in each child | the child then writes to every page itself, making them private | | Calling it before the application is loaded | there is little to optimise; objects loaded later are not compacted or promoted | | Judging sharing by summed RSS | shared pages are counted once per worker | | Lazy-loading code in workers | each worker builds its own copy of those classes |

  • Why does summing each worker's RSS overstate how much memory a preforking job runner uses?
    RSS counts every page a process can touch, including pages still shared with the parent and the other workers, so a shared page is counted once per process. Proportional set size divides each shared page among the processes sharing it, and private memory shows only what each worker owns alone. Those numbers show whether copy-on-write is working.
  • Should a preforking app call `GC.start` and `GC.compact` by hand before forking now that `Process.warmup` exists?
    On Ruby 3.3 and later, `Process.warmup` already performs a major GC and a compaction on CRuby, plus promoting survivors, precomputing string coderanges and freeing empty pages, which hand-written GC calls do not cover. Call `Process.warmup` once in the parent at the end of boot instead of assembling the steps yourself.

The workers share one printed reference book until someone scribbles in it; then that page is photocopied for the scribbler. Warmup is doing all the predictable scribbling, the margin notes and bookmarks, before handing copies out.

saying these in an interview costs you the question

  • Believes a worker's memory stays shared as long as its Ruby code only reads objects
  • Calls Process.warmup in each child after forking
  • Sums workers' RSS to prove copy-on-write is not working
  • Thinks Process.warmup freezes objects so children cannot modify them
  • Lazy-loads application code inside each worker to save boot time