skip to content

Why is filepath.WalkDir cheaper than filepath.Walk over a large directory tree?

level: middleimportance: should knowfreq 52%

answer

  1. the callbacks differ by one parameter
  2. one promises more than a directory read knows
  3. size and mode are not free
  4. the cheap one lets you opt back in
  5. a CPU profile fills with stat calls

basics

~10 s

filepath.WalkDir hands its callback an fs.DirEntry that the directory read already produced, so it needs no per-entry metadata call. filepath.Walk hands an fs.FileInfo, which forces a stat of every file and directory it visits.

solid answer

~50 s

The two functions differ in one parameter, and that parameter decides the cost. `filepath.Walk` takes a `filepath.WalkFunc`, whose second argument is an `fs.FileInfo` — to produce it, Walk must stat every single entry before calling you, whether or not you look at the metadata. `filepath.WalkDir` takes an `fs.WalkDirFunc`, whose second argument is an `fs.DirEntry`, and that is essentially free: the directory read already knew the name and the type bits. So WalkDir does roughly one directory read per directory, while Walk adds one metadata syscall per entry on top. You opt back into the cost only where you need it, by calling `d.Info()`. On a repository with tens of thousands of small files this is the difference between a walk that is dominated by stat calls in a CPU profile and one that is not. Walk is kept for compatibility; new code should use WalkDir.

code

go · 9 lines
go
err := filepath.Walk(root, func(path string, info fs.FileInfo, err error) error {
	if err != nil {
		return err
	}
	if !info.IsDir() {
		n++
	}
	return nil
})

go deeper

for a junior

Know that Go has two tree-walking functions and that filepath.WalkDir is the one to reach for in new code; be able to write its callback signature with the path, entry and error parameters.

for a middle

Explain the mechanism precisely: fs.FileInfo cannot be produced from a directory read, fs.DirEntry can, so one function stats every entry and the other does not. Name d.Info() as the opt-in cost.

for a senior

Tie it to evidence. Describe recognising the pattern in a CPU profile of a slow indexing job and confirming the fix by counting metadata calls, not just by eyeballing wall-clock time.

for a principal

Own the porting decision: when a tool arrives from a language whose walk API hands back full metadata, the naive translation carries a cost model that only shows up at production tree sizes. Decide where that gets caught.

## The two signatures ```go func filepath.Walk(root string, fn filepath.WalkFunc) error type filepath.WalkFunc func(path string, info fs.FileInfo, err error) error func filepath.WalkDir(root string, fn fs.WalkDirFunc) error type fs.WalkDirFunc func(path string, d fs.DirEntry, err error) error ``` Everything else about them matches: both start at `root` and call the callback for `root` itself first, both descend in lexical order within each directory, both call the callback sequentially from one goroutine, and neither follows symbolic links — a symlink is reported as an entry and not descended into. The only difference is the second parameter, `fs.FileInfo` versus `fs.DirEntry`. ## Why that parameter is the whole story Reading a directory yields, per entry, a name and — on the platforms that support it — a type hint saying whether it is a directory, a regular file, a symlink and so on. That is exactly what an `fs.DirEntry` exposes: `Name()`, `IsDir()`, `Type()`. Nothing more had to be asked of the filesystem to answer them. An `fs.FileInfo` is a bigger promise: `Size()`, `Mode()` with permission bits, `ModTime()`, `Sys()`. None of that comes out of the directory read. To hand your callback one, `filepath.Walk` must issue a per-entry metadata call before it can call you — for every file and every directory in the tree, unconditionally. So on a tree with `D` directories and `N` entries: * `filepath.WalkDir` ≈ `D` directory reads (plus one stat of the root). * `filepath.Walk` ≈ `D` directory reads **plus N metadata calls**. For an asset indexer over a checked-out repository — a deep tree of very many small entries, most of which the indexer discards on the strength of the name or the extension alone — `N` dwarfs `D`, and almost every one of those metadata calls is thrown away. The standard library's own documentation for `Walk` says as much: it is less efficient than `WalkDir`, which avoids the per-entry call. ## What it looks like when you are wrong about it The symptom is not a crash. It is a walk whose wall-clock time scales with file count far worse than expected, and a CPU profile in which the hot path under your walk function is filesystem metadata work rather than your own filtering. This bites hardest on the engineer porting a tool from a language whose directory-walk API hands back full metadata for every entry as a matter of course: the natural translation reaches for `filepath.Walk`, because its callback signature is the familiar one, and the cost model quietly comes along. The same mistake survives a switch to `WalkDir` if the first line of the callback is `info, err := d.Info()`. That reintroduces exactly the per-entry call you just removed. The rule is to filter on `d.Name()`, `d.IsDir()` and `d.Type()` first, and reach for `d.Info()` only for the entries that survive the filter. ## When Walk is still fine If your callback genuinely needs `Size()` or `ModTime()` for **every** entry — a `du`-style disk-usage tool, say — then Walk's eager stat costs the same as WalkDir plus `d.Info()` everywhere, and there is no win. Even then, WalkDir is the better default: it keeps the choice in your hands, and it is the shape that generalises to `fs.WalkDir` over an abstract filesystem. ## Things that are not the difference * Neither function is concurrent. Both are a single sequential traversal, so your callback does not need to be safe for concurrent use, and neither will saturate a fast disk on its own. * Neither skips hidden files or any other class of entry for you; both report everything they find. * Neither follows symbolic links, so neither can be sent into a cycle by one. * `WalkDir` is not lazy in the sense of streaming results back — it is a push traversal in both cases, driven by the callback's return value. ## Practical shape ```go err := filepath.WalkDir(root, func(path string, d fs.DirEntry, err error) error { if err != nil { return err } if d.IsDir() || filepath.Ext(d.Name()) != ".png" { return nil } info, err := d.Info() // paid only for the entries that matter if err != nil { return err } return record(path, info.Size()) }) ``` The error parameter is not decoration — handling it is a separate discipline, and returning it unconditionally stops the whole traversal.

  • Under filepath.WalkDir, when do you still pay for a metadata system call?
    Whenever you call d.Info() on an entry. That is the point of the design: the traversal itself costs one directory read per directory, and you opt into per-entry metadata only for the entries that survive your filter. Calling d.Info() on the first line of the callback throws the advantage away and makes WalkDir cost the same as filepath.Walk.
  • In what order does filepath.WalkDir visit entries, and can the callback be invoked concurrently?
    Entries are visited in lexical order within each directory, and the traversal is depth-first from the root. The callback is called sequentially from the calling goroutine, never concurrently, so it needs no internal synchronisation. If you want parallelism you build it yourself by handing paths off to workers, and then your own handoff needs to be safe.
  • Does filepath.WalkDir follow symbolic links into other parts of the filesystem?
    No. A symbolic link is reported to the callback as an entry whose type bits say symlink, and the walk does not descend through it. That is what keeps a walk from looping forever on a link that points at an ancestor directory. If you want to follow one, you resolve and walk the target deliberately, with your own cycle guard.

saying these in an interview costs you the question

  • Says filepath.WalkDir is faster because it walks concurrently
  • Thinks the difference is caching rather than the callback's parameter
  • Calls d.Info() on every entry and keeps claiming a win
  • Believes filepath.Walk follows symlinks and WalkDir does not
  • Assumes the callback may be invoked from several goroutines