skip to content

Allocation Pressure & GC

CRuby's collector marks generationally and incrementally, exposed through GC.stat, GC.compact and RUBY_GC_* variables, yet most slowness is allocation. Interviewers probe cutting objects created.

on this pageshow

explore

questions

6

In Ruby, what do GC.start and GC.disable do, and why is calling them from application code usually a mistake?

level: juniorimportance: must knowfreq 45%

answer

  1. the collector already runs itself
  2. GC.start defaults to a full collection
  3. GC.disable returns the previous state
  4. GC.start still works while disabled
  5. memory grows without bound

basics

~20 s

GC.start forces a garbage collection, a full blocking one by default; GC.disable stops automatic collection until GC.enable. Ruby already collects when it needs to, so forcing it wastes CPU and disabling it lets memory grow without limit.

solid answer

~40 s

`GC.start` runs a collection now; with its defaults (`full_mark: true, immediate_mark: true, immediate_sweep: true`) that is a **major**, stop-the-world collection that finishes before the method returns. `GC.disable` switches off automatic collection and returns whether it was *already* disabled (`false` the first time); `GC.enable` turns it back on. `GC.start` still runs while collection is disabled. In application code both are usually wrong: CRuby triggers collection itself when free slots run out or malloc'd memory crosses a limit, so a manual `GC.start` in a request adds a full-heap pause, and a `GC.disable` around a big job lets the heap grow without bound until the process is killed. They belong in benchmarks, tests and diagnosis, where you want a known heap state.

go deeper

for a junior

Recall that GC.start forces a collection, GC.disable suspends automatic ones, and that Ruby collects on its own without either.

for a middle

Explain GC.start's full_mark, immediate_mark and immediate_sweep defaults, and what GC.disable and GC.enable return.

for a senior

Show why a forced full GC in the request path hurts latency and why disabling GC postpones rather than removes the cost.

for a principal

Treat manual GC calls as a code-review smell and allow them only at deliberate points such as benchmarks or pre-fork preparation.

## What the two methods do Ruby's `GC` module is the interface to CRuby's garbage collector. Two of its methods are the ones people reach for first. **`GC.start`** starts a collection immediately. Its keyword arguments and their defaults are: - `full_mark: true` - mark every object, old and young, which is a **major** collection; `false` asks for a **minor** one that marks only young objects; - `immediate_mark: true` - finish marking before returning instead of doing it incrementally; - `immediate_sweep: true` - finish sweeping before returning instead of lazily. So a bare `GC.start` is the most expensive collection CRuby can do, and the program waits for all of it. It returns `nil`. **`GC.disable`** stops automatic collection. It returns whether collection was **already** disabled, so the first call returns `false` and a second returns `true`. `GC.enable` reverses it and returns whether collection *had been* disabled. Its documentation adds one detail that surprises people: `GC.start` "remains potent" - an explicit `GC.start` still collects while automatic collection is off. ## Why application code should rarely call them CRuby already decides when to collect. A collection starts when, for example: 1. there are no free slots left for a new object; 2. memory allocated through `malloc` since the last collection crosses the malloc limit (16 MB initially, tunable with `RUBY_GC_MALLOC_LIMIT`); 3. the number of old objects crosses the old-object limit, which forces a major collection. Against that background: | Call in app code | What actually happens | |---|---| | `GC.start` after each request or job | a full, blocking collection that the runtime would have done more cheaply, or not at all | | `GC.disable` around a heavy section | every object allocated in it stays in memory; the heap and RSS grow until `GC.enable` or an out-of-memory kill | | `GC.disable` and forgetting `GC.enable` | the process leaks memory for the rest of its life | Disabling collection does not make allocation free; it only postpones the bill, and the heap pages grown in the meantime are not necessarily handed back afterwards. ## Where they do belong - **Benchmarks** - a `GC.start` before each measured run gives every run the same starting heap; turning collection off for a short micro-benchmark removes GC noise, as long as you say so when reporting. - **Tests and diagnosis** - forcing a full collection before reading `GC.stat` or counting objects shows what is really still referenced. - **Controlled points in long-lived processes** - for example preparing a parent process before it forks workers, where a full collection (and often `GC.compact`) is done once, on purpose. ## Related methods worth knowing - `GC.count` - total number of collections so far, minor plus major. - `GC.stat` - a Hash of counters such as `:minor_gc_count`, `:major_gc_count` and `:total_allocated_objects`. - `GC.stress = true` - collect at every opportunity; its documentation says it degrades performance and is only for debugging. ## Summary `GC.start` forces a (by default full, blocking) collection; `GC.disable` suspends automatic collection and returns the previous state. The runtime already collects on its own triggers, so in production code the right number of calls to either is almost always zero.

  • In Ruby, what do two consecutive calls to GC.disable return, and why?
    `false`, then `true`. `GC.disable` returns whether collection was already disabled before the call: the first call finds it enabled, the second finds it disabled. `GC.enable` mirrors this, returning `true` when it re-enables a disabled collector and `false` when collection was already on.
  • How do you ask CRuby for a minor collection instead of a full one?
    Pass `full_mark: false`: `GC.start(full_mark: false)` marks only young objects. The keyword arguments are implementation- and version-dependent by the method's own documentation, so use them in diagnosis and tests rather than in application logic.

saying these in an interview costs you the question

  • Calling GC.start after each request keeps memory lower at no cost
  • GC.disable makes allocation faster, so disable it in hot paths
  • GC.start does nothing while GC.disable is in effect
  • GC.disable returns true when it successfully disables collection
  • A bare GC.start only collects young objects
open as a page

In Ruby, how do you reduce the objects a hot loop allocates, and how do you prove the change with GC.stat?

level: middleimportance: must knowfreq 50%

basics

~20 s

Stop building intermediate collections and per-iteration temporaries: aggregate with block forms such as sum, count and filter_map, append to strings with << rather than +=, and mutate in place where safe. Prove it with the GC.stat(:total_allocated_objects) delta before and after.

open as a page

In CRuby, how do you tell from GC.stat and GC.latest_gc_info whether a process runs mostly minor or major collections, and what triggers the majors?

level: middleimportance: should knowfreq 35%

basics

~20 s

GC.stat reports :minor_gc_count and :major_gc_count, and GC.latest_gc_info(:major_by) names why the last major ran: :nofree, :oldgen, :shady, :force or :oldmalloc, or nil when it was minor. Rising majors usually mean too many objects are being promoted to old.

open as a page

In a Ruby 4.0 preforking server, why would you run GC.compact in the parent before forking workers, and what are its limits?

level: seniorimportance: should knowfreq 25%

basics

~20 s

GC.compact moves live objects together so the parent's heap pages are full and dense. Forked workers share those pages copy-on-write; fewer half-empty pages means fewer pages copied when workers allocate. It does not move pinned objects and is not supported on every platform.

open as a page

A Ruby 4.0 nightly report job's memory keeps growing while it runs; how do you use GC.stat, ObjectSpace and the RUBY_GC_* variables to diagnose and contain it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Force a full GC and track GC.stat's heap_live_slots, old_objects and heap pages over the run: rising live slots mean the job still references its data. Size classes with ObjectSpace, fix by streaming rows, and tune RUBY_GC_HEAP_GROWTH_FACTOR or RUBY_GC_MALLOC_LIMIT only to contain the rest.

open as a page

In Ruby, how does ObjectSpace::WeakMap differ from a Hash and from ObjectSpace::WeakKeyMap when you use it as a cache?

level: middleimportance: nice to knowfreq 15%

basics

~20 s

A Hash keeps keys and values alive. ObjectSpace::WeakMap holds both weakly and compares keys by identity, so entries vanish once either is collected. ObjectSpace::WeakKeyMap, since Ruby 3.3, holds only keys weakly, keeps values strongly and compares keys with eql?.

open as a page