skip to content

A busy service's log lines go missing from the systemd journal during traffic spikes, and the journal itself contains a line saying messages from that unit were suppressed. What is systemd-journald doing, and which settings control it?

level: seniorimportance: nice to knowfreq 30%

answer

  1. the log system says it dropped them
  2. token bucket, per service
  3. default window is thirty seconds
  4. fix the unit, not the host
  5. volume itself is the real defect

basics

~20 s

journald rate-limits each service separately: past roughly 10000 messages in 30 seconds it drops the remainder and logs a suppression notice. Tune RateLimitIntervalSec= and RateLimitBurst= in journald.conf, or LogRateLimitIntervalSec= and LogRateLimitBurst= on the unit itself.

solid answer

~40 s

That suppression line is journald's own rate limiter announcing itself. journald applies a token-bucket limit **per service**: within each `RateLimitIntervalSec=` window (default 30 seconds) it accepts up to `RateLimitBurst=` messages (default 10000) from that unit, and discards the rest until the window rolls over, emitting a `Suppressed N messages…` entry so the loss is visible rather than silent. The defaults exist to stop one runaway service from filling the disk and evicting everyone else's history. If a service legitimately needs to log through a spike, override it for that unit alone with `LogRateLimitIntervalSec=` and `LogRateLimitBurst=` in its `[Service]` section — setting either to 0 disables the limit for it — rather than raising the global values and removing the protection for the whole host. Then ask whether the volume itself is the real defect.

go deeper

for a junior

Recognise a Suppressed N messages entry as journald reporting its own rate limiting, not as evidence that logging is broken, and know the defaults are roughly 10000 messages per 30 seconds.

for a middle

Explain the token-bucket model and that it is applied per service via the cgroup, then name both pairs of settings: RateLimitIntervalSec=/RateLimitBurst= globally and LogRateLimitIntervalSec=/LogRateLimitBurst= on the unit.

for a senior

Scope the fix to the offending unit with a drop-in, verify it with systemctl show, and treat the message volume itself as the defect — debug level left on, a per-request line on a hot path, or a crash loop replaying its banner.

for a principal

Own the shared-resource framing: the journal is one bounded budget per host, and per-service limits are what stop the noisiest workload from deciding how much history everyone else keeps.

## The symptom An application's own instrumentation says it logged a burst; the journal has a gap in the middle of it and a line like: ``` systemd-journald[412]: Suppressed 4312 messages from /system.slice/api.service ``` Nothing is broken. journald deliberately dropped those entries and told you so. Recognising this line is worth an interview on its own, because the alternative diagnosis — "our logging library is losing messages" — sends a team down a long, wrong road. ## What the limiter does journald maintains a rate limit **per service**, keyed on the sending process's cgroup, not one global limit for the machine. Two settings define it, in `/etc/systemd/journald.conf`: - `RateLimitIntervalSec=` — the length of the window. Default 30 seconds. - `RateLimitBurst=` — how many messages from one service are accepted per window. Default 10000. Within a window, the first `RateLimitBurst=` messages from a unit are stored; the rest are discarded. When the window rolls over, journald logs how many it dropped and starts accepting again. Setting **either** value to 0 disables rate limiting. One refinement worth knowing: journald scales the effective burst with the available space on the journal's filesystem, allowing more when there is plenty of room and tightening as it fills. So the observed threshold is not always exactly the configured number, and a host that starts dropping messages it never used to drop may be telling you the disk is filling. ## Why the limit exists Without it, a single service in a tight error loop can write the entire journal budget in seconds. Because retention is enforced by deleting the oldest archived files, that does not merely waste disk — it **evicts every other service's history**, including the entries you would need to explain the loop. The rate limiter converts "one bad service destroys the journal" into "one bad service loses some of its own messages, and says so". It also protects journald's own CPU and the write bandwidth on the log device. ## Tuning it correctly The wrong reflex is to raise `RateLimitBurst=` globally, which removes the protection for every unit on the box. Override for the one service instead — these are unit settings, not journald settings: ```ini # /etc/systemd/system/api.service.d/logging.conf [Service] LogRateLimitIntervalSec=30s LogRateLimitBurst=50000 ``` Then `systemctl daemon-reload` and restart the unit. Setting either directive to 0 turns the limit off for that unit alone. This per-unit override is the modern answer, and it is the one an interviewer is listening for: scoped, reversible, and it leaves the rest of the host protected. Before reaching for it, though, interrogate the volume. Tens of thousands of messages in thirty seconds from one service is usually one of: - a debug or trace level left enabled after an investigation; - a per-request log line on a hot path that should be sampled or aggregated; - a crash-restart loop, in which case each restart replays the same startup banner and the interesting entry is the exit status, not the flood; - a stack trace printed per failed request during a downstream outage. Raising the limit makes all four louder without making any of them better, and it hands the journal budget to the noisiest service permanently. ## Confirming and observing ```bash # find the suppression notices themselves journalctl -b _COMM=systemd-journald | grep -i suppressed # see the effective settings, drop-ins included systemd-analyze cat-config systemd/journald.conf systemctl show api.service -p LogRateLimitIntervalUSec -p LogRateLimitBurst ``` `systemctl show` is the reliable check for the per-unit values, because it reports what the manager actually loaded rather than what a file appears to say. Note that the property name it prints uses the internal microsecond form, `LogRateLimitIntervalUSec`, while the directive you write in the unit file is `LogRateLimitIntervalSec=`. ## Version note The `…IntervalSec=` spellings and the per-unit `LogRateLimit*` directives arrived in systemd 240; older releases used `RateLimitInterval=` in `journald.conf` and had no per-unit override, so on a genuinely ancient host the only lever is the global one. Every currently supported distribution is well past that boundary. ## The judgement being tested This question separates "logs are missing, therefore something is broken" from "logs are missing, and the log system told me exactly why". The strong answer names the mechanism, scopes the fix to the offending unit, and then treats the message volume itself as the defect worth investigating.

  • How do you lift the limit for one service without weakening it for the whole host?
    Put `LogRateLimitIntervalSec=` and `LogRateLimitBurst=` in that unit's `[Service]` section — ideally as a drop-in under `/etc/systemd/system/<unit>.d/` — then `daemon-reload` and restart it. Setting either to 0 disables the limit for that unit alone. Raising `RateLimitBurst=` in journald.conf instead removes the protection everywhere, which is what the limit exists to prevent.
  • Why does journald rate-limit per service rather than globally?
    Because the damage it is preventing is one service evicting everyone else's history. Retention deletes the oldest archived files, so an unbounded flood from one unit silently removes the entries you need to explain that unit's own failure. A per-cgroup bucket confines the loss to the noisy service and leaves the rest of the journal intact.
  • You raised the burst and messages are still dropped below the configured number. What else is involved?
    journald scales the effective burst with free space on the journal's filesystem, tightening as it fills. A host that suddenly starts suppressing at a lower threshold than configured is often reporting a disk problem rather than a logging one — check `journalctl --disk-usage` and the free space on `/var` before tuning further.

saying these in an interview costs you the question

  • Blames the application's logging library for the gap
  • Raises RateLimitBurst globally as the first fix
  • Thinks the limit is one global bucket for the host
  • Assumes dropped messages are silently lost
  • Treats the flood as normal and never finds its cause

context