skip to content

You are writing a metrics collector that reads Linux /proc files directly instead of shelling out to tools. Which properties of procfs, as opposed to an ordinary file on disk, does the code have to handle?

level: seniorimportance: should knowfreq 34%

answer

  1. each read runs kernel code
  2. totals since boot, not rates
  3. the PID can be gone, or be somebody else
  4. page-table walks are not free
  5. no inotify on a generated file

basics

~20 s

Procfs files are generated per read, so each open costs kernel work, values are instantaneous samples, most counters are cumulative since boot and need two samples to become rates, process directories vanish mid-scrape, PIDs are reused, and some entries are far more expensive to read than they look.

solid answer

~50 s

Treat every read as a kernel function call rather than a disk fetch. Open, read to EOF and close each time — caching a descriptor and re-reading assumes stability that procfs does not offer, and seeking is unreliable. The numbers are snapshots: most system counters such as those in /proc/stat are monotonic totals since boot, so rates require two samples and a divide, and you must handle counter resets when the machine reboots. Process enumeration is racy — a PID directory can disappear between the listing and the open, so ENOENT is a normal outcome to ignore rather than an error to log, and because PIDs are reused you may sample a different process than the one you enumerated. Finally, cost varies enormously: /proc/<pid>/stat is cheap, while entries that must walk a process's page tables are expensive enough to show up in your own CPU profile when scraping thousands of processes.

go deeper

for a junior

Know that /proc files are produced by the kernel on each read, so you open and read them fresh rather than watching them for changes like a normal file.

for a middle

Explain the difference between cumulative counters and level gauges in procfs, and why turning /proc/stat into a rate needs two samples plus handling for a counter that resets.

for a senior

Demonstrate that you have shipped this: expected ENOENT on vanished PIDs, start-time verification against PID reuse, cheap-versus-expensive file scheduling, and the awareness that a high-frequency collector perturbs the host it measures.

for a principal

Own the observability tradeoff — what scrape interval and cardinality the fleet can afford, when a text-parsing collector should give way to a purpose-built kernel interface, and how you keep parsers resilient to fields appended by newer kernels.

## Reading is executing An ordinary read pulls bytes that already exist, possibly from page cache. A procfs read runs kernel code that formats current state into a buffer. Everything unusual about writing a collector follows from this one fact. The first practical consequence is the access pattern. Open the file, read until EOF, close, and do it again next cycle. Keeping a descriptor open and re-reading is fragile: many procfs files are implemented so that content is produced during an iteration starting at open, and lseek-and-reread behaviour is not something to rely on. The cost of the open is small; the cost of being subtly wrong is not. ## Snapshots, not events Everything procfs reports is instantaneous. There is no history, no ring buffer of past values, and no notification when something changes — you cannot watch a procfs file for changes the way inotify works on a real file, because there is no write event to observe. A related trap is the *type* of number you are reading. Most system-wide statistics are cumulative counters since boot: the fields in /proc/stat are jiffies accumulated per CPU state, /proc/net/dev counts bytes and packets since the interface came up, and /proc/diskstats accumulates I/O totals. A collector must store the previous sample and emit a delta over elapsed time to produce a rate, and it must detect the counter going backwards — a reboot, an interface being recreated, or a counter wrapping — and drop that interval rather than emitting an enormous spike. By contrast, /proc/meminfo and /proc/loadavg report levels, which are used as-is. ## Consistency and races There is no snapshot isolation across files, and only weak guarantees within one. Two reads of two different files are two moments in time, so a ratio computed from them is approximate by construction. Within a single large file — a long process list, a big mapping list — the content is produced incrementally as your reads consume it, so a very long listing can be internally inconsistent: an entry added while you were reading may appear, and one removed may vanish. Process enumeration is the sharpest edge. The sequence "list /proc, filter numeric directories, open /proc/<pid>/stat for each" has a window in which the process can exit. The `open()` then fails with ENOENT, and reads on an already-open descriptor for a dead process can fail too. Correct code treats this as an expected, silent outcome. Logging it at error level produces a log line for every short-lived process on the box, which on a build machine is thousands per minute. Worse than a missing PID is a *reused* one. The kernel allocates PIDs increasing until it wraps at the value of `kernel.pid_max`, and on a busy host that wrap happens often. If you enumerate PID 4711, take a sample, and come back a minute later, the 4711 you find may be a different process entirely. Robust collectors capture a distinguishing field alongside the sample — the process start time from /proc/<pid>/stat is the conventional choice — and treat a change in it as "old process gone, new process started" rather than as a wild jump in that process's counters. ## Cost is not uniform Some files are nearly free; others are not. Reading /proc/<pid>/stat or /proc/<pid>/status is a formatted dump of fields the kernel already tracks. Reading /proc/<pid>/smaps requires walking the process's page tables mapping by mapping, and for a large-heap process with thousands of mappings this is measurably expensive — enough that a naive per-second scrape of every process on a dense host burns real CPU and perturbs the very system it is measuring. Where you need only totals, /proc/<pid>/smaps_rollup gives the aggregate at a fraction of the cost (it requires Linux 4.14 or newer). The general discipline is to scrape cheap files often and expensive files rarely, and to be honest that an observer at high frequency is part of the workload. ## Permissions and visibility Your collector will not see everything unless it is privileged. Another user's /proc/<pid>/environ, fd/ and cwd are restricted to that user and root. Beyond that, procfs accepts a `hidepid` mount option that hides other users' process directories outright, so on a hardened host an unprivileged collector may enumerate almost nothing and report a near-empty machine rather than an error. Detect that case explicitly instead of publishing a confidently wrong zero. ## Stability of the format Procfs is text, unversioned, and parsed by an enormous amount of software, so fields are appended rather than reordered — but they *are* appended, and new lines appear in new kernel versions. Parse by field name where the format is `key: value`, and by index only where the layout is genuinely fixed and documented, tolerating extra trailing fields rather than rejecting the line.

  • How do you avoid attributing one process's CPU sample to another after a PID is reused?
    Record a stable identifier alongside the PID and verify it on every sample. The conventional one is the process start time reported in /proc/<pid>/stat, which is expressed in clock ticks since boot and does not change for the life of the process. If the start time differs from what you stored, treat it as a new process rather than continuing the old series — otherwise the counter appears to jump backwards or forwards wildly.
  • Why does emitting a rate from /proc/net/dev require handling counters that go backwards?
    Those fields are cumulative totals, and the total resets when the machine reboots or when the interface is destroyed and recreated. A naive delta then produces a huge negative number, or a huge positive one if you take the absolute value. The correct behaviour is to detect the decrease, discard that interval as unmeasurable, and start a fresh baseline from the new sample.
  • Your collector's own CPU use grows with the number of processes on the host. What is the likely cause and fix?
    Per-process scraping cost, dominated by whichever files require the kernel to walk real data structures rather than print cached fields — page-table walks for detailed memory breakdowns are the usual offender. Scrape the cheap per-process files at the normal interval, use the rolled-up aggregate for memory, sample expensive detail on a slower cycle or only for processes above a threshold, and accept a cap on how many processes you report in full detail.

saying these in an interview costs you the question

  • Caches an open descriptor and re-reads it for fresh values
  • Reports /proc/stat jiffy totals as if they were rates
  • Logs a vanished PID directory as an error condition
  • Assumes a PID identifies the same process across samples
  • Thinks every /proc read costs the same tiny amount

context