skip to content

What do the Linux I/O scheduling classes set by ionice (realtime, best-effort, idle) do, and why can running a backup under `ionice -c 3` have no effect at all?

level: seniorimportance: nice to knowfreq 22%

answer

  1. three classes, one of them polite
  2. the default level comes from somewhere else
  3. a hint needs someone to honour it
  4. check the queue's active scheduler
  5. the scheduler it was written for is gone

basics

~20 s

ionice tags a process's block-layer requests with an I/O class and level, so a lower-priority job yields disk service to others. It only works if the active I/O scheduler honours those tags — with the none scheduler common on NVMe, the tag is simply ignored.

solid answer

~50 s

`ionice` sets a per-thread I/O priority through `ioprio_set(2)`: class **realtime** (1) gets device service first, **best-effort** (2) is the default with levels 0-7, and **idle** (3) only gets served when nothing else wants the device. Unset best-effort priority is derived from the CPU nice value, which is why renicing a job can quietly move its I/O priority too. The reason `ionice -c 3` on a backup can change nothing is that the priority is only a hint attached to requests — some I/O scheduler in the block layer has to act on it. `bfq` implements the classes properly; the `none` scheduler, the usual default for NVMe devices, does no reordering at all and ignores them entirely, and `mq-deadline` only gained limited priority awareness in recent kernels. Check `/sys/block/<dev>/queue/scheduler` before trusting ionice. It also cannot help with writes already absorbed into the page cache, since those are flushed later by kernel writeback threads rather than by your process.

code

bash · 8 lines
bash
# ask for the idle I/O class
ionice -c 3 tar -cf /backup/home.tar /home

# confirm what the process ended up with
ionice -p $!

# the check that decides whether any of it matters
cat /sys/block/sda/queue/scheduler

go deeper

for a junior

Know that ionice sets an I/O priority class — realtime, best-effort or idle — and that the idle class is the one used to keep background jobs off a busy disk.

for a middle

Explain that the class is metadata on block-layer requests, that an unset best-effort level is derived from the nice value, and that some scheduler must actually implement the classes.

for a senior

Show the verification habit: read the device queue's active scheduler first, know that bfq implements the classes while none ignores them, and account for buffered writeback escaping the process's priority entirely.

for a principal

Own the difference between advisory hints and enforced limits, and set the platform default — when background jobs get an enforced I/O ceiling or their own device rather than a class tag nobody may honour.

## The three classes `ionice` manipulates a per-thread I/O priority, set through `ioprio_set(2)` and readable with `ionice -p PID`. It has three classes: - **Realtime (class 1)** — requests are served ahead of everything else. Levels 0-7 within the class, 0 highest. Dangerous: a realtime-class writer can starve the rest of the system's I/O, so it is privileged. - **Best-effort (class 2)** — the normal class, also with levels 0-7. This is where everything lives by default. - **Idle (class 3)** — served only when no other process has requested the device for a period. This is the one people reach for when running backups, `rsync` jobs, `updatedb`-style indexers or checksum sweeps. ```bash ionice -c 3 tar -cf /backup/home.tar /home # run a backup in the idle class ionice -c 2 -n 7 -p 4242 # best-effort, lowest level, on a running PID ionice -p 4242 # report a process's current class and level ``` One detail worth knowing because it produces surprising behaviour: if a process has not set an I/O priority explicitly, its best-effort level is **derived from its CPU nice value**. So `nice -n 19` also de-prioritises the process's I/O relative to a nice 0 process, without anyone having run `ionice` at all — the two knobs are coupled at the default. ## Why the flag can be a complete no-op The priority is metadata attached to block-layer requests. It changes nothing on its own. Something has to *reorder* requests according to it, and that something is the **I/O scheduler** attached to the device queue: ```bash cat /sys/block/sda/queue/scheduler # e.g. [none] mq-deadline kyber bfq ``` - **`none`** — no reordering whatsoever; requests are passed straight to the device. This is the common default for NVMe drives, where the device's own parallelism makes host-side scheduling mostly counterproductive. With `none`, I/O priorities are ignored and `ionice` does nothing. - **`bfq`** — the Budget Fair Queueing scheduler, which implements the ionice classes properly, including genuine idle-class deferral. This is the scheduler ionice was designed around. - **`mq-deadline`** — a deadline scheduler that gained limited I/O-priority awareness in recent kernels, but it is not a full implementation of the class model. - **`kyber`** — latency-target based, not priority-class based. The historical context explains most of the confusion online: nearly all `ionice` documentation and blog advice was written for **CFQ**, the scheduler that implemented these classes for years. CFQ was removed along with the legacy single-queue block layer in **Linux 5.0**. Advice that was correct in 2015 quietly became a no-op on modern hardware, and nothing warns you — `ionice` returns success either way. ## The other reasons it may not help Even with `bfq` active, ionice has real blind spots: - **Buffered writes.** A write that lands in the page cache returns immediately, and the actual device I/O is issued later by kernel writeback threads, not by your process. The priority you set on the process does not straightforwardly follow that deferred writeback, so a write-heavy backup can still bury the device. - **Read-ahead and shared cache pressure.** A big sequential read evicts other processes' cached pages regardless of I/O class, so the victims see more cache misses and more I/O even when your requests are deferred politely. - **Virtualised or network storage.** On a SAN, a network filesystem, or a cloud block device, the reordering that matters happens on the far side of a boundary that never sees your ioprio. The guest's scheduler can only order what it holds. - **Device-side queueing.** Modern drives keep deep internal queues. Once requests are handed to the device, host priorities no longer apply to them. ## What to do instead When the idle class does not deliver, the alternatives are structural rather than advisory: enforce a throughput or IOPS ceiling on the job through the block-layer I/O controller, rate-limit inside the tool itself (many backup and sync tools have a bandwidth flag), schedule the work into a window where contention is acceptable, or move the noisy job to a different device. These are limits the storage stack enforces, rather than hints a scheduler may or may not consult. ## What an interviewer is checking Mostly one thing: whether you know that a command returning success does not mean it did anything. Knowing the three classes is table stakes; the senior signal is "check which I/O scheduler is active before you trust it, and remember buffered writeback is not your process's I/O any more." That habit — verifying that the mechanism you invoked is actually implemented in the path you are using — is what the question is really probing.

  • Why does a write-heavy job in the idle I/O class still hammer a shared disk?
    Buffered writes complete into the page cache and return immediately; the device I/O happens later, issued by kernel writeback threads rather than by the job's own threads. The priority attached to the process does not cleanly follow that deferred work, so the disk sees a flood of writeback with no useful class attached. Throughput limits or in-tool rate limiting work better here.
  • How would you enforce, rather than suggest, a cap on a backup job's disk usage?
    Use a mechanism the storage stack enforces: the block-layer I/O controller can impose IOPS and bandwidth ceilings on a group of processes, and most backup or sync tools offer their own bandwidth flag. Both produce a limit the job cannot exceed, whereas an I/O class is only a hint that the active scheduler may ignore entirely.
  • A process was never touched by ionice. What is its I/O class and level?
    Best-effort, with the level derived from its CPU nice value rather than a fixed default. That coupling means renicing a process also shifts its I/O priority relative to others, and it explains why two processes at the same explicit class can still be served unequally when their nice values differ.
  • Why is the none I/O scheduler the default on many NVMe devices?
    NVMe devices have deep internal queues and enough parallelism that host-side reordering adds latency without improving throughput — the device schedules better than the host can. The tradeoff is that with no host scheduler there is nothing to interpret I/O priorities, so class-based tools become inert on exactly the hardware most new systems ship with.

saying these in an interview costs you the question

  • Assumes ionice works regardless of the active I/O scheduler
  • Says the idle class guarantees the job never affects others
  • Thinks ionice throttles buffered writes to the page cache
  • Believes ionice priorities reach a SAN or cloud volume
  • Treats a successful exit status as proof it took effect

context