In CRuby, how do you tell from GC.stat and GC.latest_gc_info whether a process runs mostly minor or major collections, and what triggers the majors?
answer
- minor_gc_count vs major_gc_count
- old after surviving 3 collections
- old_objects crossing old_objects_limit
- latest_gc_info(:major_by)
- :nofree, :oldgen, :shady, :force, :oldmalloc
basics
~20 sGC.stat reports :minor_gc_count and :major_gc_count, and GC.latest_gc_info(:major_by) names why the last major ran: :nofree, :oldgen, :shady, :force or :oldmalloc, or nil when it was minor. Rising majors usually mean too many objects are being promoted to old.
solid answer
~40 sCRuby's collector is generational: a **minor** GC marks only young objects (plus old ones remembered as pointing at young), and a **major** marks everything. `GC.stat(:minor_gc_count)` and `GC.stat(:major_gc_count)` give the split. `GC.latest_gc_info(:major_by)` says why the last major happened: `:oldgen` when `:old_objects` passed `:old_objects_limit` (twice the old count after the previous major, via `RUBY_GC_HEAP_OLDOBJECT_LIMIT_FACTOR` 2.0), `:oldmalloc` when malloc growth attributed to old objects passed its limit, `:shady` for too many write-barrier-unprotected objects, `:nofree` when a minor could not free enough slots, `:force` for `GC.start`. `:gc_by` shows the trigger, such as `:newobj` or `:malloc`. Objects become old after surviving three collections, so long-lived-but-temporary data is what drives majors up.
code
ruby · 8 linesbefore = GC.stat.slice(:minor_gc_count, :major_gc_count)
run_one_batch # the workload you are inspecting
after = GC.stat.slice(:minor_gc_count, :major_gc_count)
p after.to_h { |k, v| [k, v - before[k]] }
# => {minor_gc_count: 41, major_gc_count: 3} (illustrative)
p GC.latest_gc_info.slice(:major_by, :gc_by)
# => {major_by: nil, gc_by: :newobj} when the last one was minorgo deeper
Recall that CRuby has minor collections of young objects and major collections of everything, and that GC.stat counts both.
Explain promotion after three collections, the old_objects_limit trigger, and the major_by and gc_by reasons in GC.latest_gc_info.
Show you diagnose from deltas and major_by which data is being promoted, and fix lifetimes before touching tuning variables.
Frame major-GC frequency as a consequence of data lifetimes the design chooses, and set expectations for it per workload type.
## Minor and major in CRuby terms CRuby's default collector marks and sweeps, and it is **generational**. An object starts young; once it has survived **three** collections it is promoted to **old** (`GC.stat`'s documentation describes `:old_objects` as objects that "survived at least 3 garbage collections"). - A **minor** collection marks only young objects, plus the old objects recorded as pointing at young ones. It is frequent and cheap. - A **major** collection marks every live object. It is rarer and costs time proportional to the whole live heap. Why this split pays off is general collector theory; what matters in a Ruby interview is how to read it from a running process. ## Reading the counters `GC.stat` returns a Hash, and `GC.stat(:key)` returns one value without building the Hash: | Key | Meaning | |---|---| | `:count` | all collections since start, minor plus major | | `:minor_gc_count` | minor collections | | `:major_gc_count` | major collections | | `:old_objects` | live old objects | | `:old_objects_limit` | when `:old_objects` crosses it, a major is triggered | | `:oldmalloc_increase_bytes_limit` | malloc growth by old objects that triggers a major | | `:time` | milliseconds spent in GC (measured while `GC.measure_total_time` is true, the default) | Sample these twice, a minute apart, and compare deltas; absolute values since boot hide what is happening now. ## Why did that major happen? `GC.latest_gc_info` describes the most recent collection. Its `:major_by` key is `nil` for a minor collection and otherwise one of: 1. `:nofree` - after a minor there were still not enough free slots; 2. `:oldgen` - old objects outgrew `:old_objects_limit`, which is set to `RUBY_GC_HEAP_OLDOBJECT_LIMIT_FACTOR` (default 2.0) times the old-object count after the previous major; 3. `:shady` - too many old objects without write barriers were remembered (`:remembered_wb_unprotected_objects` passed its limit); 4. `:force` - an explicit request such as `GC.start` or a C API call; 5. `:oldmalloc` - memory allocated with `malloc` attributed to old objects passed its limit (`RUBY_GC_OLDMALLOC_LIMIT`, 16 MB initially). Its `:gc_by` key tells you what started the collection: `:newobj` (no free slot for a new object), `:malloc` (the young malloc limit was reached), `:method` (`GC.start`), `:capi` or `:stress`. `:need_major_by` shows whether the next collection is already due to be major. ## What to do with the answer - **Mostly `:oldgen`** - objects that live a while and then die are being promoted: caches that churn, big per-job data structures, memoised values kept too long. Shorter lifetimes, or not creating them, lowers the major rate. - **Mostly `:oldmalloc` or `gc_by: :malloc`** - large strings and arrays whose buffers live outside the object slots dominate; reducing that buffer churn helps more than slot tuning. - **`:force` you did not expect** - something is calling `GC.start`; find it. - **`:shady`** - often C extension objects without write barriers; mostly outside your control. Ruby 3.4 added `GC.config`, and its `rgengc_allow_full_mark: false` setting stops the default collector from running majors at all, marking only young objects. Its documentation suggests calling `Process.warmup` first. It is a specialised lever for latency-critical processes, not a default. ## Summary Read `:minor_gc_count` against `:major_gc_count` as deltas, ask `GC.latest_gc_info(:major_by)` why majors happen, and reduce whatever keeps promoting objects to old.
- What does GC.latest_gc_info(:major_by) return after a minor collection?`nil`. `:major_by` is set only when the latest collection was major, and then names the reason: `:nofree`, `:oldgen`, `:shady`, `:force` or `:oldmalloc`. `:gc_by` is filled either way and names the trigger, such as `:newobj` or `:malloc`.
- Why does a large, short-lived per-job data structure raise the major GC rate?If it survives three collections while the job runs, its objects are promoted to old. When the job ends they die as old objects, which only a major collection can reclaim, and meanwhile `:old_objects` climbs toward `:old_objects_limit`, triggering majors with `major_by: :oldgen`. Building less, or processing in smaller batches, keeps them young.
saying these in an interview costs you the question
- GC.stat(:count) counts only major collections
- An object becomes old the first time it survives a collection
- A major GC is triggered only by calling GC.start
- GC.latest_gc_info(:major_by) is :minor after a minor collection
- Lowering RUBY_GC_HEAP_OLDOBJECT_LIMIT_FACTOR reduces the number of majors