Compare Files.walk() and Files.walkFileTree() for directory traversal. When would you choose each, and what are the pitfalls?
answer
- walk = lazy Stream<Path>, close it
- walkFileTree = FileVisitor callbacks
- postVisitDirectory needed for recursive delete/copy
- visitFileFailed tolerates unreadable entries
- FileVisitResult: CONTINUE/SKIP/TERMINATE
- No symlink follow unless FOLLOW_LINKS (cycle risk)
basics
~20 sFiles.walk returns a lazy Stream of all paths under a directory, great for simple filtering — but you must close it. Files.walkFileTree uses a visitor callback giving fine control over each file, directory entry/exit, and errors, which is better for complex traversals like recursive delete.
solid answer
~40 sBoth traverse a directory tree depth-first. Files.walk(start, maxDepth, options) returns a lazy Stream<Path> of every entry; it is concise for filter/map pipelines (find all .java files), but the Stream is resource-backed so it must be closed in try-with-resources, and an I/O error mid-walk surfaces as UncheckedIOException. Files.walkFileTree(start, visitor) is the callback-based API: you implement a FileVisitor (usually extending SimpleFileVisitor) with preVisitDirectory, visitFile, visitFileFailed, and postVisitDirectory, each returning a FileVisitResult (CONTINUE/SKIP_SUBTREE/SKIP_SIBLINGS/TERMINATE). That gives you per-entry control, graceful handling of unreadable entries via visitFileFailed, and the postVisitDirectory hook needed to delete directories after their contents — making it the right tool for recursive delete or copy. Both default to NOT following symlinks (you opt in with FOLLOW_LINKS, which then guards against cycles). Choose walk for simple queries, walkFileTree for stateful or order-sensitive operations.
go deeper
Knows both walk a directory tree and that walk returns a stream of paths.
Explains walk (stream) vs walkFileTree (visitor) and that walk must be closed.
Chooses correctly per use case, uses postVisitDirectory for recursive delete/copy, handles visitFileFailed and FileVisitResult control, and knows symlink defaults.
Designs safe, resilient traversal for large/untrusted trees: cycle and symlink defenses, error tolerance, depth limits, and resource lifecycle conventions.
## The task: visiting every file under a directory Neither `Path` nor a single `Files` call recurses by default. NIO.2 gives two recursive traversal tools, both **depth-first**. ## Files.walk — the Stream way ```java try (Stream<Path> paths = Files.walk(start)) { List<Path> javaFiles = paths .filter(p -> p.toString().endsWith(".java")) .toList(); } ``` `Files.walk(start)` returns a **lazy `Stream<Path>`** of `start` and everything beneath it. Overloads accept a `maxDepth` and `FileVisitOption`s. It is wonderfully concise for *queries*: filter by extension, count, map to sizes, collect. **Pitfalls of walk:** - **Must be closed.** Like `Files.lines`, the stream holds open directory resources — wrap it in **try-with-resources** or leak descriptors. - **Lazy errors.** An I/O failure encountered while walking (e.g. an unreadable subdirectory) is thrown as an **`UncheckedIOException`** during stream consumption, and by default it **aborts the whole walk**. You cannot easily "skip and continue" past a bad entry with `walk`. - **Directories appear before their contents** (pre-order), so you can't directly use it to delete a tree without sorting in reverse order first (`sorted(Comparator.reverseOrder())`) so children are deleted before parents. ## Files.walkFileTree — the Visitor way ```java Files.walkFileTree(start, new SimpleFileVisitor<Path>() { @Override public FileVisitResult visitFile(Path f, BasicFileAttributes a) throws IOException { Files.delete(f); return FileVisitResult.CONTINUE; } @Override public FileVisitResult postVisitDirectory(Path d, IOException e) throws IOException { Files.delete(d); // delete dir AFTER its contents return FileVisitResult.CONTINUE; } }); ``` `walkFileTree` drives a **`FileVisitor`** with four callbacks: - **preVisitDirectory** — before entering a directory (chance to `SKIP_SUBTREE`). - **visitFile** — for each regular file. - **visitFileFailed** — when an entry can't be read (you decide whether to continue) — this is how you *gracefully tolerate* unreadable files. - **postVisitDirectory** — *after* all children are visited — the hook that makes recursive **delete** and **deep copy** correct, because you act on a directory only once its contents are handled. Each callback returns a **`FileVisitResult`**: `CONTINUE`, `SKIP_SUBTREE`, `SKIP_SIBLINGS`, or `TERMINATE` — fine-grained control over the walk. Extending `SimpleFileVisitor` lets you override only the callbacks you need. ## Symbolic links and cycles Both methods **do not follow symbolic links by default**. You opt in with `FileVisitOption.FOLLOW_LINKS`. Once following links, a directory could link back into an ancestor, creating an **infinite cycle**; `walkFileTree` detects already-visited directories and reports a cycle via `visitFileFailed` (with `FileSystemLoopException`), whereas with `walk` you must be cautious. Following links is therefore opt-in precisely because of cycle and escape risks. ## Choosing between them | Need | Prefer | |---|---| | Find/filter/count entries (read-only query) | `Files.walk` (in try-with-resources) | | Recursive delete or deep copy | `walkFileTree` (uses postVisitDirectory) | | Skip subtrees / stop early with control | `walkFileTree` (FileVisitResult) | | Tolerate unreadable entries and keep going | `walkFileTree` (visitFileFailed) | | One-liner in a Stream pipeline | `Files.walk` | There is also `Files.find(start, depth, matcher)` — a Stream variant that takes a `BiPredicate<Path,BasicFileAttributes>`, useful when your filter wants attributes without a separate stat call. ## Summary `walk` = concise, lazy, stream-friendly, but close it and it aborts on errors. `walkFileTree` = verbose but powerful: per-entry control, error tolerance, and the post-order hook required for tree mutations.
- Why is walkFileTree better suited to recursively deleting a directory than Files.walk?walkFileTree exposes postVisitDirectory, which fires after all of a directory's contents are visited, so you can delete files first and the directory last — the correct order. With Files.walk you'd have to materialize and reverse-sort the stream so children precede parents.
- What happens with symbolic links during traversal, and what risk does following them introduce?By default neither walk nor walkFileTree follows symlinks. If you opt in with FOLLOW_LINKS, a link pointing back to an ancestor can create an infinite cycle; walkFileTree detects revisited directories and reports a FileSystemLoopException via visitFileFailed.
saying these in an interview costs you the question
- Not closing the Files.walk stream
- Trying to delete a tree with walk in default (pre-order) order
- Assuming traversal follows symlinks by default
- Believing walk lets you skip a single bad entry and continue easily
- Forgetting cycles are possible once FOLLOW_LINKS is set