skip to content

What is the Linux kernel's dentry cache, what is a negative dentry, and why can that cache grow to occupy gigabytes of memory without being a leak?

level: seniorimportance: nice to knowfreq 26%

answer

  1. path lookup is per component
  2. cache the misses as well
  3. reclaimable slab, not consumption
  4. shrinkers free it under pressure
  5. one sysctl tunes the aggressiveness

basics

~20 s

The dentry cache stores the results of name-to-inode lookups so path resolution does not re-read directories. A negative dentry caches the fact that a name does not exist. It lives in reclaimable slab memory, so the kernel shrinks it automatically under pressure.

solid answer

~50 s

Resolving a path means looking up each component in its parent directory, one at a time. That would mean reading directory data from disk for every component of every `open()`, so the kernel caches the results as **dentry** objects — in-memory (name, parent, inode) associations. A **negative dentry** caches the opposite result: this name does not exist in this directory. That sounds wasteful but is valuable, because failed lookups are extremely common — dynamic-library search paths, `PATH` scanning, and applications probing for optional config files all miss repeatedly, and the cached miss short-circuits the whole search. Dentries live in kernel slab memory and are counted as **reclaimable** in `/proc/meminfo` under `SReclaimable`. The kernel's shrinkers evict them when memory is needed, so a multi-gigabyte dentry cache on a metadata-heavy host is normal rather than a leak. `vm.vfs_cache_pressure` tunes how aggressively that reclaim happens relative to the page cache.

go deeper

for a junior

Know that the kernel caches path lookups so opening a file does not re-read every directory on the way, and that this cache is memory the kernel can take back at will.

for a middle

Explain what a dentry object associates, why the kernel also caches lookups that failed, and how reclaimable slab differs from memory the kernel cannot give back.

for a senior

Read the numbers correctly in an incident: distinguish a large reclaimable cache from genuine pressure, resist scheduled cache dropping, and explain what the shrinkers do when a cgroup or the host hits its limit.

for a principal

Own the guidance — what your monitoring should count as used memory, when tuning reclaim pressure is justified for a metadata-heavy fleet, and why cache-dropping cron jobs should be removed rather than tuned.

## Why path lookup needs a cache Opening `/usr/lib/x86_64-linux-gnu/libssl.so.3` is not one operation. The kernel resolves it component by component: find `usr` in `/`, then `lib` in that directory, and so on, checking execute (search) permission on each directory along the way. Without caching, each component would require reading directory data and its inode from the filesystem. Since a busy server performs an enormous number of path lookups, that would dominate its I/O. The VFS therefore keeps a **dentry cache** (dcache): in-memory objects that record a (parent directory, name) pair and the inode it resolves to. Dentries are hashed, so a lookup is a hash probe rather than a directory read. They also form the tree structure the kernel uses to walk paths and to produce paths back from an open file. Alongside them the kernel caches the inode objects themselves, which is why the two are usually discussed and reclaimed together. ## Negative dentries A dentry may also record that a name **does not** exist in a directory. That is a negative dentry, and it exists because misses are extremely common: - The dynamic linker searches a list of directories for a shared library; every directory but the last produces a miss. - Shell and `execvp()` lookups walk `PATH` in order, missing at each entry until they hit. - Applications probe for optional files — a config override, a `.env`, a per-user override — on every request in some frameworks. Without negative caching, each of those misses would force a real directory search. With it, the second and later attempts are answered from memory. The cost is one small object per remembered miss, and a workload that probes many distinct nonexistent names — a web application stat-ing a unique path per request, for instance — can accumulate a very large number of them. ## Why a huge dentry cache is not a leak Dentries and cached inodes are allocated from kernel **slab** caches. Slab memory splits into two classes, and the kernel reports them separately in `/proc/meminfo`: `SUnreclaim` for allocations it cannot give back, and `SReclaimable` for caches it can drop on demand. Dentry and inode caches are in the reclaimable class. The kernel registers *shrinkers* for them. When memory is needed — for a new allocation, for the page cache, or because a cgroup is at its limit — the shrinkers are called and evict unused dentries, freeing that memory. So a host that has walked millions of paths can legitimately show gigabytes of `SReclaimable` and still be nowhere near out of memory: that figure is cache, not consumption, and it is counted toward the memory considered available. The misreading of this is a classic incident: someone sees kernel slab in the gigabytes, concludes the kernel is leaking, and starts dropping caches on a schedule. Dropping caches does not fix anything; it discards useful work and produces a latency spike while the caches refill. ## The knobs, and when they are appropriate `/proc/sys/vm/drop_caches` exists for benchmarking and diagnosis. Writing `1` drops the page cache, `2` drops dentries and inodes, `3` drops both. It is non-destructive — dirty data is written back, not discarded — but it is a debugging aid, not a maintenance task, and putting it in a cron job is a well-known anti-pattern. `vm.vfs_cache_pressure` (default 100) controls how aggressively the kernel reclaims dentry and inode caches relative to page cache and swap. Raising it above 100 makes the kernel discard metadata caches sooner, which suits a host whose working set is file *contents*. Lowering it keeps metadata cached longer, which helps a metadata-heavy workload such as a file server walking huge trees, at the expense of page cache. Setting it to 0 tells the kernel never to reclaim them, which is a reliable way to run the machine out of memory — the safe range is a tuning adjustment, not an off switch. ## What it is good to notice Slab growth that tracks a workload creating and probing many paths is expected. Slab growth that never recedes *under memory pressure* is the thing worth investigating, since that means the objects are pinned rather than cached — for example dentries held by open descriptors or by mount points. The distinction between "large" and "unreclaimable" is the whole judgment here, and it is what separates an engineer who reads the number from one who understands it.

  • What does the vm.vfs_cache_pressure sysctl control, and when would you change it?
    It sets how aggressively the kernel reclaims dentry and inode caches relative to page cache and swap; 100 is the default and neutral point. Raise it when the working set is file contents and metadata caching is crowding out the page cache; lower it for metadata-heavy workloads such as file servers walking large trees. Setting it to 0 tells the kernel never to reclaim those caches, which reliably ends in an out-of-memory situation.
  • Is scheduling a periodic write to drop_caches a reasonable way to keep memory free?
    No. Dentry and inode caches are already reclaimable — the kernel drops them automatically the moment memory is needed. Dropping them on a schedule discards work that was correct, causes a latency spike while lookups go back to disk, and hides whatever real memory question prompted it. The facility is for benchmarking and diagnosis, where you want a cold cache deliberately.
  • Why does caching failed lookups pay off rather than waste memory?
    Because misses are structurally common: the dynamic linker searches a list of library directories, shells and execvp() walk PATH in order, and many applications probe for optional config files on every request. Each of those misses on all but one candidate. A negative dentry turns the repeat of that whole search into a hash probe, and the entry is a small object the kernel can evict at any time.

saying these in an interview costs you the question

  • Calls a large dentry cache a kernel memory leak
  • Says a negative dentry records a deleted file
  • Believes cached dentries must be dropped manually to free memory
  • Treats drop_caches as routine maintenance
  • Assumes all kernel slab memory is unreclaimable

context