skip to content

What does resource.getrusage(resource.RUSAGE_SELF) report about a process?

level: middleimportance: nice to knowfreq 18%

answer

  1. Accounting, not limiting
  2. The process asks about itself
  3. CPU time plus a memory high-water mark
  4. ru_maxrss units differ between platforms
  5. Children count only once reaped

basics

~20 s

It returns a resource.struct_rusage of kernel counters for the calling process: user and system CPU seconds, the peak resident set size reached so far, page faults, block I/O and context switches. It measures consumption, not limits.

solid answer

~40 s

`resource.getrusage(resource.RUSAGE_SELF)` returns a `resource.struct_rusage` for the calling process only. The fields worth knowing are `ru_utime` and `ru_stime` (CPU seconds in user and kernel code, summed across all threads), `ru_maxrss` (the **peak** resident set size — a high-water mark that never falls), the minor and major page-fault counters, and the voluntary/involuntary context-switch counters. The classic trap is that `ru_maxrss` units differ by platform: kilobytes on Linux, bytes on macOS. `resource.RUSAGE_CHILDREN` covers only children that have terminated **and been waited for**, and its `ru_maxrss` is the maximum over those children rather than a sum. For CPU timing alone `time.process_time` is cleaner; what `getrusage` adds is the memory watermark and the fault and switch counters.

code

python · 6 lines
python
import resource

usage = resource.getrusage(resource.RUSAGE_SELF)
print("user cpu seconds:", round(usage.ru_utime, 3))
print("system cpu seconds:", round(usage.ru_stime, 3))
print("peak resident size:", usage.ru_maxrss)

go deeper

for a junior

Recall that the call reports what the process has consumed, not what it is allowed to consume, and that CPU time and a peak-memory number are the headline fields.

for a middle

Explain the field semantics: CPU time summed across threads, ru_maxrss as a never-falling high-water mark, and the kilobytes-versus-bytes unit difference between platforms.

for a senior

Show how you use it in production: differencing CPU fields per job, emitting the watermark as a gauge next to the limits from resource.getrlimit, and knowing when time.process_time or tracemalloc is the better instrument.

for a principal

Own what the service publishes about its own resource consumption, so capacity decisions rest on measured headroom against configured ceilings rather than on guesswork.

### Accounting, not limiting `resource.getrlimit` tells you what your process is *allowed* to consume; `resource.getrusage` tells you what it *has* consumed. `resource.getrusage(resource.RUSAGE_SELF)` returns a `resource.struct_rusage` — a named tuple-like object of counters the kernel has been maintaining for this process since it started. It is a single, cheap syscall, always available on Unix, and needs no third-party profiler. ### The fields that carry real information * `ru_utime` and `ru_stime` — CPU seconds spent in user code and in the kernel on this process's behalf, as floats. They are summed across all threads, so a process with four busy threads can report four seconds of `ru_utime` per wall-clock second. * `ru_maxrss` — the **peak** resident set size reached so far: a high-water mark that never goes down, even after the memory is freed. It answers "how big did this ever get", not "how big is it now". * `ru_minflt` and `ru_majflt` — page faults served without and with disk I/O. A climbing major-fault count is the fingerprint of a process being swapped. * `ru_inblock` and `ru_oublock` — block input and output operations. * `ru_nvcsw` and `ru_nivcsw` — voluntary context switches (the process blocked on something) and involuntary ones (the scheduler preempted it). A high involuntary count means CPU contention. Several of these fields are simply zero on some platforms; the kernel fills in what it tracks, and Python passes the structure through unchanged. Check a value is non-zero on your target OS before building an alert on it. ### The portability trap in `ru_maxrss` The units of `ru_maxrss` are **not** the same everywhere. On Linux the value is in kilobytes; on macOS it is in bytes. The same code therefore reports a number a thousand times larger on one platform than on the other, and the classic bug is a memory dashboard that looks fine in CI on one OS and absurd in production on the other. Normalise explicitly at the point of measurement, keyed on `sys.platform`, and label the unit in whatever you emit. ### `RUSAGE_SELF` versus `RUSAGE_CHILDREN` `resource.RUSAGE_SELF` covers the calling process only — not its children, however many it has spawned. `resource.RUSAGE_CHILDREN` covers children that have **terminated and been waited for**; a child still running contributes nothing, and one that was never reaped contributes nothing either. For the children entry, `ru_maxrss` is the maximum over those children, not their sum, so it will never tell you the aggregate footprint of a worker pool. A parent that wants a true total has to collect a measurement inside each worker before it exits. ### How it compares with the alternatives `time.process_time` returns the same CPU-time concept with a cleaner interface and higher resolution, so for pure CPU timing it is the better call; `time.monotonic` gives wall clock, and the gap between wall clock and CPU time is exactly how much of a slow operation was spent waiting rather than computing. What `getrusage` adds over both is the memory high-water mark and the fault and context-switch counters, which no `time` function reports. For live memory rather than the peak, `tracemalloc` measures Python-level allocations, and the operating system's own per-process reporting measures the whole process. ### Using it well The useful pattern is a difference, not an absolute: capture the structure before and after a unit of work and subtract the CPU-time fields, because the raw values include everything since process start. `ru_maxrss` cannot be differenced meaningfully — it only ever rises — so treat it as a lifetime watermark, sampled at the end of a job or emitted periodically as a gauge. Logging `ru_utime`, `ru_stime` and `ru_maxrss` once per completed job, next to the soft limits read from `resource.getrlimit`, gives a cheap, dependency-free answer to "how close is this workload to the ceiling it was given" — which is the question the limits themselves cannot answer.

  • Why can the ru_maxrss value not be differenced across a unit of work?
    Because it is a high-water mark, not a current reading: it only ever rises, and it does not fall when memory is freed. Subtracting two samples measures whether a new peak was set, not how much the work used. Difference the CPU-time fields, which do accumulate, and treat `ru_maxrss` as a lifetime watermark sampled at the end of a job.
  • What does resource.getrusage(resource.RUSAGE_CHILDREN) include?
    Only children that have terminated and been waited for; a running child, or one that was never reaped, contributes nothing. The CPU-time fields sum across those children, but `ru_maxrss` is the maximum over them rather than a total, so it never reports a worker pool's aggregate footprint. Measuring inside each worker before it exits is the way to get that.

saying these in an interview costs you the question

  • Reading ru_maxrss as current rather than peak memory
  • Assuming ru_maxrss units are the same on every platform
  • Expecting RUSAGE_CHILDREN to include still-running children
  • Treating the children's ru_maxrss as a sum
  • Confusing getrusage with getrlimit — usage versus ceiling

context