skip to content

How do you recursively traverse a directory tree in Kotlin using only kotlin.io? Compare walkTopDown and walkBottomUp, explain what a FileTreeWalk is, and how to filter, limit depth, and handle errors during the walk.

level: seniorimportance: should knowfreq 40%

answer

  1. walkTopDown / walkBottomUp -> FileTreeWalk : Sequence<File>
  2. Top-down emits dir before children; bottom-up after
  3. Bottom-up for deletes (dir must be empty)
  4. maxDepth, onEnter(false=skip), onLeave, onFail
  5. Lazy Sequence -> filter/map/first short-circuits

basics

~20 s

Call directory.walkTopDown() or walkBottomUp() to get a lazy sequence of every file and folder under it. Top-down visits a folder before its contents; bottom-up visits contents first. You can filter, limit depth, and handle errors with builder methods.

solid answer

~40 s

kotlin.io provides File.walk(direction), File.walkTopDown(), and File.walkBottomUp(), which return a FileTreeWalk — a lazy Sequence<File> over the directory tree. walkTopDown emits a directory before descending into its children (good for filtering whole subtrees early); walkBottomUp emits children before their parent directory (required when deleting, since a dir must be empty before removal). FileTreeWalk is configurable via maxDepth(n), onEnter { dir -> Boolean } (return false to skip a subtree), onLeave { dir }, and onFail { dir, IOException -> } to handle directories you can't list. Because it's a Sequence it's lazy and composes with filter/map/count, etc. The starting file itself is included. It does not follow symlinks into cycles by default in a way that loops infinitely on a normal tree, but be cautious with symlinked directories.

code

kotlin · 11 lines
kotlin
import java.io.File

// Find the first large log, pruning .git, depth-limited, fault-tolerant:
val hit = File("/srv").walkTopDown()
    .maxDepth(4)
    .onEnter { it.name != ".git" }
    .onFail { dir, e -> System.err.println("skip $dir: ${e.message}") }
    .firstOrNull { it.isFile && it.extension == "log" && it.length() > 1_000_000 }

// Bottom-up so directories are visited after their contents (delete order):
File("/tmp/build").walkBottomUp().forEach { it.delete() }

go deeper

for a junior

Knows walkTopDown gives all files under a directory and you can forEach over them.

for a middle

Distinguishes top-down vs bottom-up and uses maxDepth and a basic filter.

for a senior

Explains FileTreeWalk is a lazy Sequence, uses onEnter/onFail/onLeave, picks direction by use case (delete=bottom-up), and notes symlink cycles.

for a principal

Reasons about traversal performance/laziness on huge trees, fault tolerance via onFail, security of untrusted trees, and how deleteRecursively/copyRecursively build on the walk.

## The API `kotlin.io` adds three traversal entry points on `java.io.File`: - `File.walk(direction: FileWalkDirection)` - `File.walkTopDown()` — shorthand for `walk(TOP_DOWN)` - `File.walkBottomUp()` — shorthand for `walk(BOTTOM_UP)` Each returns a **`FileTreeWalk`**, which **is a `Sequence<File>`** (it implements `Sequence`). "Sequence" means it is **lazy**: nodes are produced on demand as you iterate, so you can stop early or chain `filter`/`map`/`first`/`count` without materializing the whole tree into a list. ```kotlin import java.io.File File("/project").walkTopDown() .filter { it.isFile && it.extension == "kt" } .forEach { println(it) } ``` The **starting `File` is included** as the first element. ## Top-down vs bottom-up - **`walkTopDown`** — a directory is emitted **before** its children, then it descends. Use when you want to **prune** whole subtrees early (skip `node_modules`) or process parents first. - **`walkBottomUp`** — children are emitted **before** their parent directory. Use when an operation requires the directory to be **empty first** — classically **deleting** a tree (you must remove files before the folder). ## Configuring the walk (FileTreeWalk builder methods) `FileTreeWalk` returns a new configured walk from these (call before iterating): - **`maxDepth(n: Int)`** — limit how deep to descend; `1` means only the immediate children. - **`onEnter { dir: File -> Boolean }`** — called before entering a directory; **return `false` to skip that whole subtree** (only meaningful top-down). - **`onLeave { dir: File -> }`** — called after a directory's contents are processed. - **`onFail { dir: File, e: IOException -> }`** — called when a directory **cannot be listed** (e.g., permission denied). Without it, a failure surfaces differently; with it you can log and continue. ```kotlin val walk = File("/data").walkTopDown() .maxDepth(3) .onEnter { dir -> dir.name != ".git" } // prune .git .onFail { dir, e -> log.warn("skip $dir: ${e.message}") } val bytes = walk.filter { it.isFile }.sumOf { it.length() } ``` ## Symlinks and cycles The walk descends into directories via `listFiles()`. A **symlink that points back up the tree** can create a cycle; the standard library does not guarantee cycle protection, so on untrusted trees consider checking with `java.nio` real-path or tracking visited canonical paths. ## Why it's a Sequence (laziness payoff) Because `FileTreeWalk : Sequence<File>`, terminal operations like `first { ... }` short-circuit and never visit the rest of the tree — important for large directories. Building a `List` (`.toList()`) forces full traversal. ## Relationship to other helpers `deleteRecursively()` is implemented on top of a **bottom-up** walk; `copyRecursively()` uses a top-down walk. Understanding the walk explains those higher-level helpers.

  • Why must a manual recursive delete use walkBottomUp rather than walkTopDown?
    A directory can only be removed once empty; bottom-up visits files before their parent folder, so deletes succeed in order.
  • How do you skip an entire subtree (like node_modules) without visiting its contents?
    Use onEnter { dir -> dir.name != "node_modules" } and return false; the walk won't descend into it.
  • What does it mean that FileTreeWalk is a Sequence?
    It is lazy: elements are produced on demand, so first/take short-circuit and the whole tree need not be loaded into memory.

Top-down is reading a table of contents before each chapter's pages; bottom-up is stacking the pages and only then filing the chapter folder.

saying these in an interview costs you the question

  • Saying walk returns a List (it is a lazy Sequence)
  • Using top-down to delete a tree and being surprised non-empty dirs fail
  • Not knowing onEnter returning false prunes a subtree
  • Forgetting the start File is included in the walk
  • Ignoring symlink cycles on untrusted trees

context