skip to content

What is the difference between Path.glob and Path.rglob in pathlib, and what do they return?

level: middleimportance: should knowfreq 55%

answer

  1. One of the two is recursive
  2. Nothing is a list here
  3. Two stars cross directory levels
  4. rglob prepends a recursive prefix
  5. Sort it yourself if order matters

basics

~10 s

Both return a lazy generator of Path objects, not a list. Path.glob(pattern) matches the pattern against entries under that directory; Path.rglob(pattern) is shorthand for glob("**/" + pattern), so it searches every subdirectory recursively.

solid answer

~40 s

`Path.glob(pattern)` walks the receiver directory applying a shell-style pattern — `*`, `?`, `[seq]` and the recursive `**` — and yields `Path` objects prefixed with the receiver, so `Path("exports").glob("*.json")` yields `exports/schedule.json`. `Path.rglob(pattern)` is defined as `glob("**/" + pattern)`: the only difference is the implicit recursive prefix. Both are **generators**, so nothing is materialised until you iterate, they can only be consumed once, and `len()` needs `list()` or `sorted()` around them. Ordering is whatever the filesystem hands back — never assume alphabetical. Neither follows symlinked directories when recursing unless you pass `recurse_symlinks=True` (3.13+), and `case_sensitive=` (3.12+) overrides the platform default. For a single level with no pattern use `Path.iterdir()`; for a full walk with pruning and an error hook use `Path.walk()` (3.12+).

code

python · 12 lines
python
from pathlib import Path
import tempfile, os

os.chdir(tempfile.mkdtemp())
root = Path("exports")
(root / "2026-03-01").mkdir(parents=True, exist_ok=True)
(root / "2026-03-01" / "schedule.json").write_text("{}", encoding="utf-8")
(root / "manifest.json").write_text("{}", encoding="utf-8")

print(sorted(p.name for p in root.glob("*.json")))
print(sorted(p.name for p in root.rglob("*.json")))
print(sorted(str(p) for p in root.glob("**/*.json")))

go deeper

for a junior

Be able to list files by extension in a directory and, separately, throughout a tree, and to say which of the two calls is the recursive one. Remember to wrap the result in list() or sorted() before you index it.

for a middle

Explain the mechanics: rglob is glob with a '**/' prefix, both are generators, * never crosses a separator, and ordering is whatever the filesystem returns. Know the case_sensitive and recurse_symlinks keywords and when they matter.

for a senior

Show that you treat traversal as a cost centre: prune with a narrower pattern or Path.walk() instead of filtering a firehose, avoid mutating a tree while iterating it, and know that glob offers no hook for unreadable directories.

for a principal

Own the policy for scanning large trees: sorted, deterministic listings where output feeds a diff or a manifest, bounded traversal for untrusted directory structures, and a clear rule for when a filesystem scan should be replaced by an index.

Globbing is pattern-matched directory traversal, and the two `pathlib` entry points differ in exactly one thing: whether the pattern is implicitly recursive. ### What the two calls mean `Path.glob(pattern)` interprets the pattern relative to the path it is called on. `Path.rglob(pattern)` is documented as equivalent to calling `glob()` with `"**/"` prepended to the pattern. So `Path("exports").rglob("*.json")` and `Path("exports").glob("**/*.json")` do the same work. There is no third mode; `rglob` is a convenience, not a different algorithm. The pattern language is shell-style, not regular expressions: `*` matches any run of characters **within one component** and never crosses a separator, `?` matches one character, `[abc]` and `[a-z]` match a character class, and `**` matches any number of directory levels. A pattern is matched against components, so `glob("*/*.json")` is exactly one directory deep — a common off-by-one boundary when someone means "one level or deeper" and gets "exactly one level" instead. Anything deeper needs `**`. ### What comes back Both calls return a **generator of `Path` objects**, and each yielded path is prefixed with the receiver, so the results are usable directly rather than being bare filenames. The consequences are the ones that catch people out: - Nothing happens until you iterate. A `glob` call that appears to be instant has done no work yet. - It is single-use. Iterating twice gives you an empty second pass; wrap it in `list()` or `sorted()` if you need it more than once, or a count. - `len(p.glob("*"))` is a `TypeError`; `sum(1 for _ in ...)` or `len(list(...))` is the count. - **Order is unspecified.** It comes from the underlying directory scan and varies by filesystem. If output order matters — a diff, a manifest, a checksum over filenames — sort explicitly. - Because it is lazy, mutating the tree while iterating (renaming or deleting matches as you go) gives you undefined results. Materialise first, then mutate. ### Symlinks, case and errors Recursion does not descend into symlinked directories by default. Python 3.13 made that explicit by adding a `recurse_symlinks` keyword, defaulting to `False`; passing `True` opts into following them, and then a symlink cycle is your problem to bound. Python 3.13 also changed what a pattern *ending* in `**` matches: it now yields files as well as directories, where earlier versions yielded directories only. Python 3.12 added `case_sensitive=`, which overrides the platform default (case-sensitive on POSIX, insensitive on Windows) — useful when a test must behave the same on both. Traversal errors are not surfaced through `glob`: there is no error callback, so if a subdirectory cannot be read the walk simply produces fewer results. When you need to know, use `Path.walk()` (3.12+), which yields `(dirpath, dirnames, filenames)` with `dirpath` as a `Path`, lets you prune the walk by mutating `dirnames` in place, and accepts `on_error=` and `top_down=`. ### Choosing the right tool - One directory, no pattern: `Path.iterdir()`. - One directory, a pattern: `Path.glob("*.json")`. - Whole tree, a pattern: `Path.rglob("*.json")`. - Whole tree with pruning, error handling or per-directory logic: `Path.walk()`. - Testing a path you already have against a pattern, with no filesystem access at all: `PurePath.match()`, which anchors from the right, or `PurePath.full_match()` (3.13+), which matches the whole path and understands `**`. ### A note on cost Consider a flight-schedule differ that compares daily export files across an archive tree. Its parse cache hits about 83% of the time, so parsing is no longer where the wall-clock goes — the traversal is. Two things dominate: how much of the tree the pattern forces you to visit, and how many stat calls the matching needs. `rglob("*.json")` visits every directory under the root; `glob("2026-*/*.json")` prunes at the first level and can be an order of magnitude cheaper on a deep archive. Narrowing the pattern, or switching to `Path.walk()` and pruning `dirnames`, beats filtering a firehose of results in Python.

  • Why does iterating the result of Path.rglob() a second time yield nothing?
    Because it is a generator, not a sequence. The first pass consumes it and leaves it exhausted; there is no rewind. If you need the results twice — say to count them and then process them — bind `list(...)` or `sorted(...)` once and iterate that. The same laziness is why `len()` fails on the result and why deleting or renaming files while iterating gives undefined behaviour.
  • Does Path.rglob() follow symlinked directories while recursing?
    Not by default. Python 3.13 added an explicit `recurse_symlinks` keyword defaulting to `False`, which preserves the older behaviour of not descending through symlinked directories. Pass `True` only when you genuinely want to traverse them, and be aware that you then own the risk of a symlink cycle turning the walk into an unbounded one.
  • When would you reach for Path.walk() instead of rglob()?
    When you need control the pattern cannot express: pruning whole subtrees by mutating `dirnames` in place, handling per-directory state, or reacting to traversal errors through `on_error=` — `glob` has no error hook and silently yields fewer results. `Path.walk()` arrived in 3.12 and yields `(dirpath, dirnames, filenames)` with `dirpath` as a `Path`, so it is the `os.walk` shape with pathlib objects.

saying these in an interview costs you the question

  • Says glob returns a list of filename strings
  • Assumes results come back alphabetically sorted
  • Thinks a single * crosses directory separators
  • Passes recursive=True, which belongs to glob.glob
  • Calls len() directly on the glob result
  • Deletes matched files while still iterating the generator

context