skip to content

How does gc.get_objects() with collections.Counter give a census of live objects?

level: middleimportance: should knowfreq 35%

answer

  1. First ask what is piling up
  2. The collector already keeps a list
  3. Count by type, then subtract
  4. Diff two snapshots, not absolutes
  5. Untracked atoms never appear

basics

~20 s

gc.get_objects() returns a list of every container object the cyclic collector currently tracks. Feeding the type name of each one into collections.Counter yields counts per class; taking two censuses around a workload and diffing them shows which type is accumulating.

solid answer

~40 s

The idiom is `Counter(type(obj).__name__ for obj in gc.get_objects())` after a `gc.collect()`, run once at a quiet baseline and again after the workload, then subtracted. It answers the first question of any leak hunt -- *what* is piling up -- before you spend time on *who* holds it. Two caveats matter. First, `gc.get_objects()` returns only objects the collector tracks, which means containers and ordinary class instances; atomic values such as `int`, `str`, `float` and `bytes` are untracked, so a pile of strings only shows up as growth in whatever list or dict holds them. Second, the returned list is the size of the heap and briefly holds a reference to everything in it, so the call is slow, allocation-heavy, and something you delete immediately rather than keep.

code

python · 12 lines
python
import gc
from collections import Counter

class LabResult:
    pass

kept = [LabResult() for _ in range(3)]

gc.collect()
census = Counter(type(obj).__name__ for obj in gc.get_objects())
print(census["LabResult"])
print(census.most_common(3))

go deeper

for a junior

Know that the collector can hand you the live tracked objects and that counting them by type is a real technique. Recall the one-liner shape: build a Counter over type names from gc.get_objects().

for a middle

Explain the mechanics an interviewer is after: collect first, diff two snapshots, and know that only gc-tracked objects appear, so ints and strings are absent by design. Be able to say why the returned list itself is expensive.

for a senior

Show operational judgement -- when you would take this in a live process, how you keep the diagnostic from becoming the leak, and how you separate warm-up growth from real accumulation by repeating the workload and checking the delta scales.

for a principal

Own where this sits in a diagnosis strategy: a zero-setup triage step that narrows the search before anyone pays for heavier instrumentation, plus a view on what a service should expose by default so leak hunts do not require attaching to production by hand.

### What the call actually returns `gc.get_objects()` asks the cyclic garbage collector for its tracking lists and returns them as one new list. Those lists are how the collector does its job: to find unreachable cycles it must know every object that *could* participate in one, which means every object capable of holding a reference to another object. So the result is a snapshot of the tracked heap -- lists, dicts, sets, tuples that contain non-atomic items, instances of ordinary classes, functions, frames, modules, cells, and so on. Since Python 3.8 the call also accepts an optional generation argument, so you can ask for only the youngest tracked objects; the census technique below is unchanged through 3.14 either way. ### The census idiom ```python import gc from collections import Counter def census(): gc.collect() return Counter(type(obj).__name__ for obj in gc.get_objects()) before = census() run_one_batch() after = census() print((after - before).most_common(15)) ``` Three details make this useful rather than noise: - **`gc.collect()` first.** Without it the snapshot includes cyclic garbage that is dead but not yet reclaimed, and every diff is polluted by collection timing rather than by your workload. - **Diff, never absolutes.** An idle interpreter already carries thousands of functions, dicts and wrapper descriptors. Only the delta across a workload means anything. - **Repeat the workload.** Run the batch several times and check the delta scales with iterations. A one-off delta is usually warm-up: interned constants, lazily imported modules, a filled cache that has now reached its ceiling. Using `type(obj).__name__` keeps the counter keys hashable and printable; using `type(obj)` itself works too and disambiguates same-named classes from different modules, at the cost of holding a reference to each class. ### What the census cannot see This is the part interviewers probe, because getting it wrong sends people down a blind alley. The collector only tracks objects that can refer to other objects. Atomic values do not: `int`, `float`, `str`, `bytes` and `bool` instances are never tracked, and CPython also untracks tuples and dicts once it can prove they contain only untracked items. `gc.is_tracked()` will confirm this on any object. The practical consequence is that a leak made of strings shows in the census as growing `list` or `dict` counts, not as growing `str` counts -- so you learn the shape of the container rather than the payload, and you size the payload some other way. Memory held outside the Python heap is invisible too: buffers allocated by a native extension, memory the allocator has freed to its own pools but not returned to the OS, and fragmentation all move RSS without moving any count in the census. That is why the census is a triage step and not a memory accounting tool -- it tells you which *kind* of Python object is multiplying, and nothing about bytes. ### Cost and safety in production The call walks the whole tracked heap and materialises a list as long as it. On a large process that is tens of millions of pointers: hundreds of milliseconds, plus a transient allocation, plus a reference to every tracked object that keeps anything otherwise dying alive until the list is released. Assign it, use it, then `del` it -- and never hold a census list across a request or stash it in a global, or the diagnostic becomes the leak. Because it stops nothing and takes no lock of its own, it is safe to run in a live process from a debug endpoint or a signal handler, but treat it as an occasional operation rather than a periodic metric. In a multi-threaded process the snapshot is inherently fuzzy: other threads keep allocating while you iterate. ### Where it sits in a leak hunt A disciplined sequence is: confirm growth is real and monotonic, take a type census diff to learn what is accumulating, then take one instance of the accumulating type and walk backwards to find the object still referring to it. The census is deliberately cheap in thinking effort -- it needs no instrumentation, no restart, and no prior hypothesis, which is exactly why it goes first. Line-level attribution of allocations is a different tool with a different setup cost, and it becomes far more useful once the census has told you which type to care about.

  • A census shows str counts flat while memory climbs. What does that tell you?
    Nothing on its own, because `str` objects are untracked and never appear in `gc.get_objects()` at all. Flat is the expected reading whether or not strings are the payload. Look instead at which container types grew -- a rising `list` or `dict` count with steady instance counts usually means one big container is being appended to -- and size the payload separately rather than by counting objects.
  • Why call gc.collect() before taking each snapshot?
    Because otherwise the snapshot includes cyclic garbage that is unreachable but not yet reclaimed, and whether it appears depends on when the collector last happened to run. That makes two censuses non-comparable for reasons unrelated to your workload. Collecting first puts both snapshots in the same, deterministic state: everything the collector can prove dead is gone, so the delta reflects retention rather than collection timing.
  • Is it safe to run this census on a live production process?
    Occasionally, yes, and it needs no restart or instrumentation, which is its main virtue. But it walks the entire tracked heap and builds a list as long as it, so expect a pause proportional to heap size and a transient spike, and it briefly pins every tracked object. Trigger it from a debug endpoint, delete the list immediately, and never sample it on a timer.

saying these in an interview costs you the question

  • Expects str or int counts to reveal a string leak
  • Reads absolute counts instead of diffing two snapshots
  • Keeps the returned list around in a global
  • Thinks gc.get_objects() reports memory in bytes
  • Samples the census on a timer in production
  • Skips gc.collect() and blames noisy results on the tool

context