What does resource.getrusage(resource.RUSAGE_SELF) report about a process?
answer
- Accounting, not limiting
- The process asks about itself
- CPU time plus a memory high-water mark
- ru_maxrss units differ between platforms
- Children count only once reaped
basics
~20 sIt returns a resource.struct_rusage of kernel counters for the calling process: user and system CPU seconds, the peak resident set size reached so far, page faults, block I/O and context switches. It measures consumption, not limits.
solid answer
~40 s`resource.getrusage(resource.RUSAGE_SELF)` returns a `resource.struct_rusage` for the calling process only. The fields worth knowing are `ru_utime` and `ru_stime` (CPU seconds in user and kernel code, summed across all threads), `ru_maxrss` (the **peak** resident set size — a high-water mark that never falls), the minor and major page-fault counters, and the voluntary/involuntary context-switch counters. The classic trap is that `ru_maxrss` units differ by platform: kilobytes on Linux, bytes on macOS. `resource.RUSAGE_CHILDREN` covers only children that have terminated **and been waited for**, and its `ru_maxrss` is the maximum over those children rather than a sum. For CPU timing alone `time.process_time` is cleaner; what `getrusage` adds is the memory watermark and the fault and switch counters.
code
python · 6 linesimport resource
usage = resource.getrusage(resource.RUSAGE_SELF)
print("user cpu seconds:", round(usage.ru_utime, 3))
print("system cpu seconds:", round(usage.ru_stime, 3))
print("peak resident size:", usage.ru_maxrss)go deeper
Recall that the call reports what the process has consumed, not what it is allowed to consume, and that CPU time and a peak-memory number are the headline fields.
Explain the field semantics: CPU time summed across threads, ru_maxrss as a never-falling high-water mark, and the kilobytes-versus-bytes unit difference between platforms.
Show how you use it in production: differencing CPU fields per job, emitting the watermark as a gauge next to the limits from resource.getrlimit, and knowing when time.process_time or tracemalloc is the better instrument.
Own what the service publishes about its own resource consumption, so capacity decisions rest on measured headroom against configured ceilings rather than on guesswork.
### Accounting, not limiting `resource.getrlimit` tells you what your process is *allowed* to consume; `resource.getrusage` tells you what it *has* consumed. `resource.getrusage(resource.RUSAGE_SELF)` returns a `resource.struct_rusage` — a named tuple-like object of counters the kernel has been maintaining for this process since it started. It is a single, cheap syscall, always available on Unix, and needs no third-party profiler. ### The fields that carry real information * `ru_utime` and `ru_stime` — CPU seconds spent in user code and in the kernel on this process's behalf, as floats. They are summed across all threads, so a process with four busy threads can report four seconds of `ru_utime` per wall-clock second. * `ru_maxrss` — the **peak** resident set size reached so far: a high-water mark that never goes down, even after the memory is freed. It answers "how big did this ever get", not "how big is it now". * `ru_minflt` and `ru_majflt` — page faults served without and with disk I/O. A climbing major-fault count is the fingerprint of a process being swapped. * `ru_inblock` and `ru_oublock` — block input and output operations. * `ru_nvcsw` and `ru_nivcsw` — voluntary context switches (the process blocked on something) and involuntary ones (the scheduler preempted it). A high involuntary count means CPU contention. Several of these fields are simply zero on some platforms; the kernel fills in what it tracks, and Python passes the structure through unchanged. Check a value is non-zero on your target OS before building an alert on it. ### The portability trap in `ru_maxrss` The units of `ru_maxrss` are **not** the same everywhere. On Linux the value is in kilobytes; on macOS it is in bytes. The same code therefore reports a number a thousand times larger on one platform than on the other, and the classic bug is a memory dashboard that looks fine in CI on one OS and absurd in production on the other. Normalise explicitly at the point of measurement, keyed on `sys.platform`, and label the unit in whatever you emit. ### `RUSAGE_SELF` versus `RUSAGE_CHILDREN` `resource.RUSAGE_SELF` covers the calling process only — not its children, however many it has spawned. `resource.RUSAGE_CHILDREN` covers children that have **terminated and been waited for**; a child still running contributes nothing, and one that was never reaped contributes nothing either. For the children entry, `ru_maxrss` is the maximum over those children, not their sum, so it will never tell you the aggregate footprint of a worker pool. A parent that wants a true total has to collect a measurement inside each worker before it exits. ### How it compares with the alternatives `time.process_time` returns the same CPU-time concept with a cleaner interface and higher resolution, so for pure CPU timing it is the better call; `time.monotonic` gives wall clock, and the gap between wall clock and CPU time is exactly how much of a slow operation was spent waiting rather than computing. What `getrusage` adds over both is the memory high-water mark and the fault and context-switch counters, which no `time` function reports. For live memory rather than the peak, `tracemalloc` measures Python-level allocations, and the operating system's own per-process reporting measures the whole process. ### Using it well The useful pattern is a difference, not an absolute: capture the structure before and after a unit of work and subtract the CPU-time fields, because the raw values include everything since process start. `ru_maxrss` cannot be differenced meaningfully — it only ever rises — so treat it as a lifetime watermark, sampled at the end of a job or emitted periodically as a gauge. Logging `ru_utime`, `ru_stime` and `ru_maxrss` once per completed job, next to the soft limits read from `resource.getrlimit`, gives a cheap, dependency-free answer to "how close is this workload to the ceiling it was given" — which is the question the limits themselves cannot answer.
- Why can the ru_maxrss value not be differenced across a unit of work?Because it is a high-water mark, not a current reading: it only ever rises, and it does not fall when memory is freed. Subtracting two samples measures whether a new peak was set, not how much the work used. Difference the CPU-time fields, which do accumulate, and treat `ru_maxrss` as a lifetime watermark sampled at the end of a job.
- What does resource.getrusage(resource.RUSAGE_CHILDREN) include?Only children that have terminated and been waited for; a running child, or one that was never reaped, contributes nothing. The CPU-time fields sum across those children, but `ru_maxrss` is the maximum over them rather than a total, so it never reports a worker pool's aggregate footprint. Measuring inside each worker before it exits is the way to get that.
saying these in an interview costs you the question
- Reading ru_maxrss as current rather than peak memory
- Assuming ru_maxrss units are the same on every platform
- Expecting RUSAGE_CHILDREN to include still-running children
- Treating the children's ru_maxrss as a sum
- Confusing getrusage with getrlimit — usage versus ceiling