skip to content

A Ruby 4.0 nightly report job's memory keeps growing while it runs; how do you use GC.stat, ObjectSpace and the RUBY_GC_* variables to diagnose and contain it?

level: seniorimportance: should knowfreq 40%

answer

  1. live slots after a full GC
  2. heap growth vs retention vs malloc
  3. ObjectSpace.memsize_of_all needs objspace
  4. RUBY_GC_HEAP_GROWTH_FACTOR default 1.8
  5. RUBY_GC_MALLOC_LIMIT default 16 MB

basics

~20 s

Force a full GC and track GC.stat's heap_live_slots, old_objects and heap pages over the run: rising live slots mean the job still references its data. Size classes with ObjectSpace, fix by streaming rows, and tune RUBY_GC_HEAP_GROWTH_FACTOR or RUBY_GC_MALLOC_LIMIT only to contain the rest.

solid answer

~40 s

First classify the growth. Call `GC.start` at checkpoints and log `GC.stat` `:heap_live_slots`, `:old_objects`, `:heap_allocated_pages` and `:heap_free_slots`. Live slots that keep rising after full collections mean **retention**: the job still references its data, typically an ever-growing results Array, a memo Hash or a String report built in memory. Flat live slots with a large heap mean it grew to a **peak** and kept the pages. Flat Ruby counters but rising RSS point at **malloc'd** buffers, such as big strings or C extension memory. `ObjectSpace.count_objects` and, after `require "objspace"`, `ObjectSpace.memsize_of_all(SomeClass)` show which types dominate. The fix is structural: stream rows and write output incrementally. Then contain: a lower `RUBY_GC_HEAP_GROWTH_FACTOR` (default 1.8) grows the heap in smaller steps and a lower `RUBY_GC_MALLOC_LIMIT` (16 MB) collects sooner, both at the cost of more GC time.

code

ruby · 15 lines
ruby
require "objspace"

def checkpoint(label)
  GC.start
  s = GC.stat
  warn format("%s live=%d old=%d pages=%d free=%d strings=%dKB",
              label, s[:heap_live_slots], s[:old_objects],
              s[:heap_allocated_pages], s[:heap_free_slots],
              ObjectSpace.memsize_of_all(String) / 1024)
end

File.foreach("input.csv").each_slice(10_000).with_index do |batch, i|
  batch.each { |line| out.puts(format_row(line)) } # write out, keep nothing
  checkpoint("batch #{i}")
end

go deeper

for a junior

Recall that GC.stat reports live and free slots and that a job which keeps every row in memory will grow however the GC is tuned.

for a middle

Explain heap_live_slots versus available slots, memsize_of being shallow, and what the growth factor and malloc limit do.

for a senior

Show the method: full-GC checkpoints, classify retention, peak or off-heap, fix by streaming, then tune RUBY_GC_* with measured trade-offs.

for a principal

Decide when a job should be redesigned to stream versus given a memory budget, and who owns tuning variables across deployments.

## Three different kinds of "growing" "Memory keeps growing" can mean three different things, and each needs a different fix: 1. **Retention** - the job still references objects it no longer needs, so no collection can free them. 2. **Heap growth to a peak** - the job briefly needed many objects; CRuby added heap pages to hold them and the process keeps its size even though the objects are gone. 3. **Off-heap growth** - memory allocated with `malloc` for String and Array buffers, or by C extensions, grows while the object heap looks stable. ## Step 1: classify with GC.stat Take checkpoints during the run - say after every 10,000 rows - and at each one call `GC.start` (a full collection) and record: | `GC.stat` key | What it tells you | |---|---| | `:heap_live_slots` | objects still alive after the collection | | `:old_objects` | long-lived survivors | | `:heap_available_slots` / `:heap_allocated_pages` | how big the object heap has become | | `:heap_free_slots` / `:heap_empty_pages` | how much of it is unused | | `:malloc_increase_bytes` | malloc growth since the last collection | Read the pattern: - live slots and old objects rising checkpoint after checkpoint -> **retention**; - live slots flat, allocated pages high, many free slots -> **peak growth**; - all of the above flat while the process's RSS rises -> **off-heap**. ## Step 2: find what is taking the space - `ObjectSpace.count_objects` (core) counts live objects by internal type: `:T_STRING`, `:T_ARRAY`, `:T_HASH` and so on. - `require "objspace"` adds `ObjectSpace.memsize_of(obj)` and `ObjectSpace.memsize_of_all(klass)`. Both are **hints**: the method's own documentation says the size is incomplete, especially for C extension data, and `memsize_of` is **shallow** - an Array's size does not include the Strings it holds. - For "which line allocated this", use an allocation profiler; that is a profiling tool's job, not GC's. ## Step 3: fix the structure Report jobs grow for predictable reasons: - accumulating every row in an Array before writing any output; - building the whole report as one String or one big Hash; - a memo Hash keyed by row that is never cleared; - loading a whole table or file when a cursor or `IO#each_line` would stream it. Stream instead: read a batch, transform it, write it to the output IO, drop the references, repeat. Peak live objects then stay proportional to the batch size, not the input size. ## Step 4: contain what is left with RUBY_GC_* variables These environment variables are read when the process starts: - **`RUBY_GC_HEAP_GROWTH_FACTOR`** (default `1.8`) - the maximum factor by which the object heap grows when it needs more slots. A value such as `1.2` grows it in smaller steps, so a peak overshoots less. Values of `1.0` or below are ignored. - **`RUBY_GC_MALLOC_LIMIT`** (default 16 MB, accepts `k`/`M`/`G` suffixes) - the initial malloc growth that starts a collection. It then adapts: grown by `RUBY_GC_MALLOC_LIMIT_GROWTH_FACTOR` (1.4) up to `RUBY_GC_MALLOC_LIMIT_MAX` (32 MB). Lower values collect sooner on buffer-heavy work. - **`RUBY_GC_HEAP_OLDOBJECT_LIMIT_FACTOR`** (2.0) - how far old objects may grow before a major collection. Every one of them trades memory for more GC time, and an invalid or out-of-range value is silently ignored in a normal run. Since Ruby 3.3, initial heap size is set per size pool with `RUBY_GC_HEAP_0_INIT_SLOTS` through `RUBY_GC_HEAP_4_INIT_SLOTS`; the old single `RUBY_GC_HEAP_INIT_SLOTS` was removed. ## Summary Classify with full-GC checkpoints of `GC.stat`, size the culprits with `ObjectSpace`, fix retention by streaming, and use `RUBY_GC_HEAP_GROWTH_FACTOR` and `RUBY_GC_MALLOC_LIMIT` only to shave what the design cannot.

  • ObjectSpace.memsize_of(rows) returns a few kilobytes for an Array holding thousands of long Strings; why?
    `memsize_of` is shallow: it reports the Array object and its own element buffer, not the Strings it references. Its documentation also calls the result a hint that is incomplete. To size the Strings, use `ObjectSpace.memsize_of_all(String)` or sum `memsize_of` over the elements yourself.
  • The job's live slots return to baseline after each batch, yet RSS stays at its peak; what is happening?
    The object heap grew to hold the peak batch and CRuby kept those pages; memory freed inside them is reused by Ruby but not necessarily returned to the OS. Smaller batches lower the peak, and a smaller `RUBY_GC_HEAP_GROWTH_FACTOR` makes each growth step smaller. It is not a leak: live objects are back to baseline.
  • Why do RUBY_GC_* settings not fix a retention problem?
    They change when and how much the collector runs, but a collector can free only unreachable objects. If the job still references every row, more frequent or earlier collections find nothing to free and just cost CPU. Retention is fixed by dropping references, for example by streaming output.

saying these in an interview costs you the question

  • ObjectSpace.memsize_of includes everything an object references
  • Rising RSS always means Ruby objects are leaking
  • Lowering RUBY_GC_HEAP_GROWTH_FACTOR fixes a job that retains its rows
  • RUBY_GC_HEAP_INIT_SLOTS still sets the initial heap size in Ruby 4.0
  • Setting RUBY_GC_HEAP_GROWTH_FACTOR=1.0 stops the heap from growing
  • GC.stat counters without a GC.start show what is still referenced