skip to content

An incident happened at 03:00 on a Linux server and nobody was logged in at the time. If sysstat is installed, how do you go back and look at CPU and disk behaviour for that window, and what determines whether the data exists at all?

level: seniorimportance: should knowfreq 44%

answer

  1. someone has to have been writing it down
  2. the reader is not the collector
  3. a binary file per day under /var/log
  4. default cadence smears short spikes
  5. it knows the machine, never the process

basics

~20 s

Read the sysstat archive with sar -f /var/log/sa/saDD -s 03:00 -e 04:00, adding -u for CPU, -d -p for devices and -q for run-queue history. Whether it exists depends on collection being enabled, the collection interval, and the retention setting — Debian and Ubuntu ship it disabled.

solid answer

~50 s

`sar` is a reader; the writer is `sadc`, invoked periodically by `sa1` from a cron entry or from `sysstat-collect.timer` on systemd distributions, and it appends binary samples to a daily file under `/var/log/sa`. So the first question is whether that timer or cron job has been running: on Debian and Ubuntu the package ships with collection disabled and you have to set `ENABLED` in `/etc/default/sysstat`. The second is the interval — the default is every ten minutes, so a ninety-second spike is averaged into a ten-minute window and may be invisible. The third is retention, set by `HISTORY` in the sysstat configuration file, which decides how many days of files are kept. When it is all in place, `sar -f /var/log/sa/sa15 -s 02:30 -e 04:00 -u -P ALL` and the same with `-d -p` will show me CPU and per-device behaviour across the incident. What it will never show me is which process was responsible — sysstat's periodic collection records system-wide counters, not per-process history.

code

bash · 8 lines
bash
# CPU, per-CPU, across the incident window from this month's day-15 file
sar -f /var/log/sa/sa15 -s 02:30 -e 04:00 -u -P ALL

# per-device I/O for the same window, with readable device names
sar -f /var/log/sa/sa15 -s 02:30 -e 04:00 -d -p

# is anything being collected at all on a systemd host?
systemctl status sysstat-collect.timer

go deeper

for a junior

Know that sar replays history that was collected earlier into files under /var/log/sa, and that -f, -s and -e select the file and the time window.

for a middle

Explain the split between the sadc collector driven on a schedule and the sar reader, name the report flags for CPU, devices and queues, and say that every value is an average over the collection interval.

for a senior

Show that you check whether the data can exist before hunting in it — collection enabled, interval, retention, which counter groups were recorded — and state plainly that there is no per-process history, so attribution has to come from elsewhere.

for a principal

Treat host-local history as a deliberate baseline rather than an accident of packaging. Decide which fleets get collection enabled, at what cadence and retention, and be clear about what that cheap always-on recorder is for versus what belongs to a real telemetry system.

## The two halves of sysstat People say "sar" and mean the whole thing, but there are two pieces: - **`sadc`** — the collector. It reads kernel counters and appends a binary record to a daily data file, conventionally `/var/log/sa/saDD` where `DD` is the day of the month. - **`sar`** — the reader. With no arguments it reports today's file; `-f <file>` reads a specific one. `sadc` does not run itself. A wrapper script, `sa1`, is invoked on a schedule — historically from a cron entry installed by the package, and on current systemd distributions from `sysstat-collect.timer`, with a companion `sysstat-summary.timer` that produces a daily summary. If neither is active, there is no history, and no amount of `sar` invocation will conjure it. ## Reading the window ``` # CPU across the incident, all CPUs broken out sar -f /var/log/sa/sa15 -s 02:30 -e 04:00 -u -P ALL # per-device I/O with friendly device names sar -f /var/log/sa/sa15 -s 02:30 -e 04:00 -d -p # run-queue and load history sar -f /var/log/sa/sa15 -s 02:30 -e 04:00 -q # network interface counters sar -f /var/log/sa/sa15 -s 02:30 -e 04:00 -n DEV ``` The report flags are the same ones the live tools use, which is the point: `-u` CPU, `-d` block devices, `-b` I/O rates, `-q` queue statistics, `-n` network, `-P ALL` per-CPU rather than aggregate. `-s` and `-e` bound the time range. Because the data files are binary and versioned, a file written by a much older sysstat may not be readable by a much newer `sar` — worth knowing before you copy archives off a decommissioned host. ## What determines whether the data exists This is the part interviewers are actually probing, because it is where people discover at 03:05 that they have nothing. 1. **Is collection enabled?** On Debian and Ubuntu the package installs with data collection switched off; `ENABLED` in `/etc/default/sysstat` must be set and the service restarted. On RHEL-family distributions the collection schedule is installed active. Check whether `sysstat-collect.timer` is active, or that the cron entry exists. 2. **What is the interval?** Ten minutes is the conventional default. Every value you read back is an average over that window, so a two-minute CPU saturation or a short I/O storm is smeared across ten minutes and may barely register. Shortening the interval costs disk and produces more samples; it is a deliberate trade, and worth making on hosts that matter. 3. **How long is it kept?** The `HISTORY` setting in the sysstat configuration file (`/etc/sysstat/sysstat` on Debian-family, `/etc/sysconfig/sysstat` on RHEL-family) sets the number of days retained. With the traditional day-of-month file naming, files also wrap after a month — a `sa15` written last month is overwritten this month — so retention beyond a month needs the alternative dated file naming. 4. **Which statistics were collected?** `sadc` can be told which groups of counters to record. On some distributions disk statistics are not collected by default and need the disk group added to the collector's options; if `sar -d` reports nothing for a window in which the machine was clearly busy, that is usually why rather than the disks being idle. ## What sar cannot tell you Be explicit about the limits, because overstating them is a red flag: - **No per-process history.** The periodic collection records system-wide and per-device counters. It will show you that the box was CPU-saturated at 03:10; it will not name the process. `pidstat` exists and can run live, but it is not part of the default periodic collection. - **Averages, not peaks.** Every value is a mean over the collection interval. A sub-interval spike is gone. - **Only what was configured.** Counters that were not in the collection set were never written. ## Where it sits in an investigation sysstat is the host's own flight recorder, and its value is that it is already there on a box that was never wired into anything else — the last resort when the incident is over and nobody was watching. Treat it as a first stop for shape and timing (when did it start, was it CPU or I/O, was it one CPU or all of them), then move to whatever live evidence you can still gather, and to the application's own logs for attribution. Confirming that collection is enabled with a sensible interval on hosts you care about is a five-minute job that pays for itself the first time you are woken at 04:00.

  • A 90-second CPU spike is definitely known to have happened, but nothing shows in the sar history. Why?
    Almost certainly the collection interval. With the conventional ten-minute cadence, ninety seconds at 100% averages to roughly 15% across the window, which looks unremarkable next to normal load. If those spikes matter on that host, shorten the interval and accept more samples and more disk, or rely on something sampling continuously instead.
  • Can sar tell you which process caused the CPU saturation at 03:00?
    No. The periodic collection records system-wide and per-device counters, not per-process ones, so it can show CPU saturated and which CPUs, but never who. Attribution has to come from the application's own logs, from a scheduled job you can correlate by time, or from something that was recording per-process data live.
  • You copy `/var/log/sa` files off a decommissioned host and `sar -f` refuses to read them. What is the likely cause?
    The data files are a versioned binary format, and a file written by a substantially different sysstat release may not be readable by the `sar` on your machine. Read them with a matching sysstat version, or convert them on the original host — `sadf` can render a data file into a text format such as CSV or JSON that survives the move.

saying these in an interview costs you the question

  • Assuming sysstat collection is on by default everywhere
  • Expecting sar to name the responsible process
  • Reading interval averages as peak values
  • Forgetting that day-of-month files wrap after a month
  • Believing sar samples live when you run it

context