Explain Path.resolve, Path.relativize, and Path.normalize — what does each do and when would you use them?
answer
- resolve = append child (absolute child wins)
- relativize = inverse of resolve, same type required
- normalize = collapse . and .. lexically
- resolve+normalize+startsWith = anti-traversal
- normalize is lexical; toRealPath follows symlinks
basics
~20 sresolve joins two paths (base + child). relativize finds the route from one path to another. normalize cleans up redundant . and .. segments. They are pure string-style operations that do not touch the disk.
solid answer
~40 sAll three are pure, disk-free Path transformations. resolve(other) appends other to this path; if other is absolute it just returns other, so it is the standard way to build a child path under a base directory (base.resolve("sub/file.txt")). relativize(other) computes the relative path that, when resolved against this, reaches other — both paths must be the same type (both absolute or both relative). It is the inverse of resolve. normalize() removes redundant elements: . (current dir) and .. (parent), collapsing "a/b/../c" to "a/c" lexically, without consulting the file system. A common pattern for safely placing user input under a base is base.resolve(userPart).normalize(), then checking the result still startsWith(base) to prevent path-traversal escapes. Note normalize is purely lexical, so it does not resolve symbolic links — use toRealPath() for that.
go deeper
Can describe resolve as joining paths and normalize as cleaning up dots, even if fuzzy on relativize.
Correctly explains all three, knows relativize is resolve's inverse with the same-type constraint, and that they are lexical.
Applies resolve+normalize+startsWith for traversal defense and knows the lexical limitation requires toRealPath for symlinks.
Designs safe path-handling utilities, codifies the traversal-defense idiom across the codebase, and reasons about symlink and cross-filesystem edge cases.
These three methods are the core *path arithmetic* of NIO.2. Crucially, **all are lexical** — they manipulate the name elements as data and never read the disk. ## Vocabulary first - A path is **absolute** if it has a root (`/home/x` or `C:\x`); otherwise **relative** (`docs/a.txt`). - `.` means "this directory"; `..` means "the parent directory". These are *name elements*, not magic. ## resolve — joining paths `a.resolve(b)` builds a new path by appending `b` to `a`: ```java Path base = Path.of("/var/data"); base.resolve("reports/jan.csv"); // /var/data/reports/jan.csv ``` Three rules: (1) if `b` is **absolute**, resolve returns `b` unchanged (you can't append an absolute path); (2) if `b` is **empty**, it returns `a`; (3) otherwise it concatenates. There is also `resolveSibling(b)`, which resolves against the *parent* — handy for renaming a file in place: `p.resolveSibling("new-name.txt")`. ## relativize — the inverse of resolve `a.relativize(b)` answers: *what relative path takes me from `a` to `b`?* It is constructed so that `a.resolve(a.relativize(b))` equals `b` (lexically). ```java Path x = Path.of("/a/b"); Path y = Path.of("/a/c/d"); x.relativize(y); // ../c/d ``` **Constraint:** both paths must be the *same kind* — both absolute or both relative — otherwise it throws `IllegalArgumentException`, because you cannot lexically bridge between a rooted and an unrooted path. ## normalize — cleaning the path string `normalize()` removes redundant `.` and collapses `..` against the preceding element, **lexically**: ```java Path.of("a/./b/../c").normalize(); // a/c ``` Because it is purely textual, `normalize()` does **not** follow symbolic links. If `a/b` is a symlink, lexically dropping `b/..` may not match where the link actually pointed. When you need the *true*, link-resolved, absolute path that must exist on disk, use `Files`'... `toRealPath()` instead. ## The security pattern: stopping path traversal A classic vulnerability is letting a user supply `../../etc/passwd` to escape an intended directory. The safe idiom: ```java Path base = Path.of("/srv/uploads").toAbsolutePath().normalize(); Path target = base.resolve(userSuppliedName).normalize(); if (!target.startsWith(base)) { throw new SecurityException("path traversal attempt"); } ``` `resolve` builds the candidate, `normalize` collapses the `..`, and `startsWith` proves the result is still inside `base`. (For full robustness against symlinks, also compare `toRealPath()`.) ## When to use which - **resolve:** build a child/sibling path under a known base. - **relativize:** compute a portable relative reference (e.g. to store a path independent of the install root, or to mirror a directory tree). - **normalize:** clean up user/derived paths before comparing or using them, especially for security checks.
- Why isn't normalize() enough to fully prevent a path-traversal attack?normalize is purely lexical, so it cannot account for symbolic links. A symlink inside the base could still point outside it. For full safety, resolve the real path with toRealPath() and verify it still startsWith the base's real path.
- What happens if you call /a/b.relativize(c/d) with one absolute and one relative path?It throws IllegalArgumentException. relativize requires both paths to be the same kind (both absolute or both relative) because it cannot lexically bridge a rooted and an unrooted path.
saying these in an interview costs you the question
- Thinking these methods access the disk or resolve symlinks
- Calling relativize across an absolute and a relative path
- Trusting normalize alone to defeat symlink-based traversal
- Forgetting that resolve returns the argument unchanged when it is absolute