How do you run tcpdump for days on a production host to catch an intermittent fault without filling the disk, using -C, -G and -W?
answer
- many files, bounded total
- size first, then a count
- time needs a strftime name
- -G with -W stops
- stop it before it wraps
basics
~20 sRotate with -C (a new file each time one passes N million bytes) and cap the set with -W for a ring that overwrites the oldest file. -G rotates by time into strftime-named files; with -W it exits after N files.
solid answer
~40 sWrite with `-w` and rotate. `-C 500` opens a new file whenever the current one passes 500 million bytes, appending a number to the whole name (`edge.pcap`, `edge.pcap1`, ...). Adding `-W 40` turns that into a ring of 40 zero-padded files that overwrites from the start, so disk use stays near 20 GB. `-G 3600` rotates every hour instead, and the `-w` name needs a strftime pattern such as `edge-%Y%m%d-%H%M.pcap` or each file overwrites the last; with `-W`, `-G` stops after that many files and exits with status 0 rather than looping. Size the ring from the filtered rate and the time until someone notices the fault, add a capture filter and a headers-only snaplen, and stop or copy the ring as soon as the fault fires.
code
bash · 3 linessudo mkdir -p /var/capture
sudo tcpdump -i eth0 -s 256 -C 500 -W 40 -w /var/capture/edge.pcap 'host 198.51.100.20 and tcp port 5432'
sudo tcpdump -i eth0 -s 256 -G 3600 -z gzip -w '/var/capture/edge-%Y%m%d-%H%M.pcap' 'tcp port 5432'go deeper
Remember that -w can be split into many files and that -C, -G and -W control the split: size, time and count.
Explain the exact behaviours: -C in millions of bytes with numbers appended, -C with -W as a ring, -G needing a strftime name, and -G with -W exiting after n files.
Size a ring from the filtered rate and the time to notice, shrink the volume with a filter and snaplen, keep it off the service's disk, and plan how the ring is stopped or copied when the fault fires.
Weigh ad-hoc rings on hosts against a standing capture appliance or flow telemetry: retention, who may hold packet data, and what evidence an incident review will expect.
## The problem rotation solves An intermittent fault, such as a burst of resets at 03:00 or a database connection that stalls once a day, means starting a capture before you know when the evidence will appear. A single `-w` file grows until the disk fills, which can take the production service down with it. tcpdump has three options that turn one file into a managed set of files: `-C`, `-G` and `-W`. ## Size-based rotation with `-C` - Before writing each packet, tcpdump checks whether the current file is **already larger than** the `-C` value; if so, it closes it and opens a new one. Files therefore end slightly above the limit. - The unit is **millions of bytes** (1,000,000, not 1,048,576). A suffix changes it: `k`/`K` for KiB, `m`/`M` for MiB, `g`/`G` for GiB. - The first file has the `-w` name; later ones get a number **appended to the whole name**: `edge.pcap`, `edge.pcap1`, `edge.pcap2`. The `.pcap` extension is no longer last, which matters to tools that decide by extension. ## Turning it into a ring with `-W` - `-C` with `-W n` limits the set to n files and **starts overwriting from the beginning**: a rotating buffer. Names get enough leading zeros to sort, so `-W 40` produces `edge.pcap00` to `edge.pcap39`. - Disk use is bounded at roughly n times the size: `-C 500 -W 40` is about 20 GB. - Once the ring has wrapped, the file number no longer tells you which file is newest; sort by modification time. ## Time-based rotation with `-G` - `-G seconds` rotates every that many seconds. The `-w` name **should contain a strftime format**, such as `%Y%m%d-%H%M%S`, expanded in local time. - Without a time format, each new file **overwrites the previous one**. A format coarser than the period, such as `%H` with `-G 600`, also overwrites, because consecutive files get the same name. - `-G` with `-W n` is **not a ring**: tcpdump stops after creating n files and **exits with status 0**. - `-G` together with `-C`: files are named `file<count>`, and `-W` is ignored except for its effect on the name. | Combination | Behaviour in tcpdump 4.99 | |---|---| | `-C size` | a new numbered file each time the size is passed; unbounded | | `-C size -W n` | a ring of n files, oldest overwritten | | `-G secs` with a strftime name | a new timestamped file every period; unbounded | | `-G secs -W n` | n files, then exit with status 0 | | `-C` and `-G` together | files named `file<count>`; `-W` affects only the name | ## Sizing the window 1. Estimate the filtered rate. A 20 Mbit/s flow is about 2.5 MB/s, roughly 9 GB an hour. 2. Decide how long after the fault someone will notice and act; that is the window the ring must cover. 3. Choose n and size to cover it with a margin, and cut the volume first: a tight capture filter and a headers-only snapshot length often shrink it by an order of magnitude. 4. Arrange to **stop the capture or copy the ring** promptly when the fault fires; otherwise the evidence is overwritten. ## Rehearse before leaving it running 1. Run the exact command for a minute with tiny values, such as `-C 1 -W 3`, and watch the files appear: you see the real names, the zero padding and the wrap-around before trusting it for days. 2. Check that the directory is writable by the account tcpdump will run as after any privilege drop; the first rotation is where a permission problem shows up. 3. Check the free space against n times the size, plus the slight overshoot of each file. 4. Note how the capture will be stopped (a service manager, a kill by PID) and who is allowed to do it. ## Operational details - `-z command` runs a command on each file as it is closed, for example `-z gzip`; tcpdump runs it in parallel with the capture at the lowest priority. - Keep the files off the filesystem the service depends on. - If tcpdump drops privileges (`-Z`, or a build with a default user), every rotated file is opened as that user, so the directory must be writable by it. - Run the capture under a service manager or `nohup` so that closing your SSH session does not end it, and record the exact command with the files.
- Your ring holds `edge.pcap00` to `edge.pcap39`; which file holds the newest packets?The names do not tell you once the ring has wrapped: tcpdump overwrites from `00` again after reaching `39`. Sort by modification time (`ls -lt`) or read the first packet's timestamp in each file. The newest file is the one tcpdump is still writing, so stop the capture before copying.
- Why did `tcpdump -G 3600 -W 24 -w 'edge-%H.pcap'` stop after a day instead of rotating forever?With `-G`, `-W` limits how many files are created, and tcpdump exits with status 0 when it reaches the limit; it is not a ring. For a rolling window by time, leave `-W` out and delete old files with a scheduled job, or use the `-C` with `-W` ring, which does overwrite.
- What goes wrong with `-G 600 -w 'edge-%Y%m%d-%H.pcap'`?The name changes only once an hour, but tcpdump rotates every ten minutes, so each rotation within the hour opens the same name and overwrites the data already there. The man page warns against a time format coarser than the rotation period; add `%M` (and `%S` if needed).
A ring of capture files works like a dashcam that keeps overwriting its memory card: it always holds the last stretch of road, so after a crash you must pull the card before the loop records over the evidence.
saying these in an interview costs you the question
- tcpdump's -C 100 means 100 MiB per file.
- -G alone always creates new files, even when the -w name has no time format.
- -W always turns tcpdump's output into a ring, whatever it is combined with.
- Rotated tcpdump files keep the .pcap extension at the end of their names.
- A ring buffer keeps the evidence until someone gets round to looking next week.