What is the difference between shutil.copy, shutil.copy2 and shutil.move?
answer
- A ladder of how much comes along
- Bytes, then permission bits, then metadata
- One of them is a rename first
- Crossing a filesystem changes everything
- Ownership survives none of them
basics
~10 sshutil.copy duplicates contents plus permission bits. shutil.copy2 additionally preserves metadata such as modification and access times. shutil.move relocates a path, renaming it when possible and otherwise copying with copy2 and deleting the source.
solid answer
~40 s`shutil.copyfile` copies bytes only. `shutil.copy` copies the bytes and the permission bits, and accepts a directory as the destination. `shutil.copy2` is `copy` plus `shutil.copystat`, so modification and access times and other metadata survive — it is the right default when something downstream compares timestamps, and it is what `shutil.copytree` and `shutil.move` use internally. `shutil.move` first tries a rename, which is instantaneous and atomic within one filesystem; if the source and destination are on different devices it falls back to copying with `copy2` and unlinking the source, so it is neither cheap nor atomic in that case. If the destination is an existing directory the source is moved *inside* it, and `move` returns the final path. None of these preserve ownership, which needs root and a separate `os.chown`.
code
python · 13 linesimport os
import shutil
import tempfile
with tempfile.TemporaryDirectory() as d:
src = os.path.join(d, "ledger.csv")
with open(src, "w", encoding="utf-8") as f:
f.write("id,amount\n")
os.utime(src, (0, 0))
plain = shutil.copy(src, os.path.join(d, "plain.csv"))
kept = shutil.copy2(src, os.path.join(d, "kept.csv"))
print("copy kept mtime:", os.stat(plain).st_mtime == 0)
print("copy2 kept mtime:", os.stat(kept).st_mtime == 0)go deeper
Recall which function to reach for: shutil.copy2 when you want a faithful duplicate, shutil.move to relocate, shutil.copytree for a directory. Knowing they take paths and that copy accepts a destination directory is enough here.
Explain the ladder precisely — copyfile bytes, copy adds permission bits, copy2 adds timestamps and other metadata via copystat — and that move tries os.rename first and only copies when the paths straddle filesystems.
Show the operational consequences: a cross-device move is a full copy and is not atomic, an observer can see a partial destination, and the fix is placing the staging file on the destination's filesystem rather than retrying the move.
Own the layout decision — which volumes staging, temp and published data live on, so that hand-offs are renames rather than byte copies — and set the expectation that publication is atomic by construction rather than by hoping a move stays cheap.
## The family, from thinnest to thickest `shutil` sits on top of `os` and the `open` builtin to provide the bulk file operations that are tedious to hand-roll. The copy family is a ladder, and the difference between the rungs is entirely about *how much of the file besides its bytes* comes along. **`shutil.copyfile(src, dst)`** copies contents and nothing else. The destination must be a full file path, not a directory, and the function refuses to copy a file onto itself, raising `shutil.SameFileError`. **`shutil.copy(src, dst)`** copies the contents and then the permission bits, via the same logic as `shutil.copymode`. It accepts a directory as `dst` and reuses the source's basename inside it. It returns the destination path. **`shutil.copy2(src, dst)`** is `copy` plus `shutil.copystat`, which carries over the access and modification times, the permission bits, and where the platform supports them the file flags and extended attributes. It is the closest the standard library gets to a faithful duplicate, and it is the default `copy_function` for both `shutil.copytree` and `shutil.move`. What *none* of them copy is ownership. The user and group of the destination are whatever your process would normally create; changing that requires privilege and an explicit `os.chown`. Candidates routinely claim `copy2` preserves the owner, and it does not. For directories there is **`shutil.copytree(src, dst)`**, which walks recursively and applies `copy_function` per file. Its useful knobs are `dirs_exist_ok=True` (since Python 3.8) to merge into an existing destination instead of raising, `ignore=` to skip patterns, and `symlinks=` to choose between copying links as links or dereferencing them. Errors from individual files are collected and raised together as a `shutil.Error` at the end rather than aborting on the first one. ## `move` is a rename with a fallback `shutil.move(src, dst)` is not a copy at all in the good case. It attempts `os.rename` first: within one filesystem that is a metadata-only operation, instantaneous regardless of file size, and atomic — an observer sees the old path or the new path, never a half-file. When the two paths are on different filesystems, rename cannot work, and `move` falls back to copying with `copy2` and then deleting the source (`os.unlink` for a file, `shutil.rmtree` for a directory tree). Three consequences follow, and they are the reason this question gets asked: 1. **Cost.** A "move" across a mount point reads and writes every byte. Moving a large staging file from a RAM-backed temp filesystem to a data volume is a full copy, not the free operation the name suggests. 2. **Non-atomicity.** During the fallback the destination exists partially written. Anything watching that directory can see a truncated file. If you need atomic publication, keep the temp file on the *same* filesystem as the destination, which is what the `dir=` argument to `tempfile.mkstemp` is for. 3. **Failure in the middle.** A crash between copy and delete leaves both copies. Two more behaviours worth knowing. If `dst` is an existing directory, the source is moved *into* it, keeping its basename — so `move("batch.csv", "archive")` produces `archive/batch.csv`, and `move` returns that final path, which is why you should use the return value rather than reconstructing it. And on a same-filesystem rename, an existing destination file is silently replaced on Unix, so `move` is not a safe way to avoid clobbering; check first or pick a name that cannot collide. ## Speed Since Python 3.8 the copy functions use platform-specific zero-copy calls where they can — `os.sendfile` on Linux and a native copy call on macOS — so the bytes move inside the kernel instead of through a Python-level read/write loop. That makes `shutil.copyfile` substantially faster than the obvious hand-written loop, which is a good reason to reach for the module rather than reimplementing it. `shutil.copyfileobj(fsrc, fdst)` remains the tool when you have open file objects rather than paths, for instance streaming a response body to disk. ## The rest of the bulk toolkit `shutil.make_archive(base_name, format, root_dir=...)` writes a `.zip` or a tar variant of a directory and returns the archive's path; `shutil.get_archive_formats()` lists what the build supports, which on Python 3.14 includes a Zstandard tar format alongside the gzip, bzip2 and lzma ones. `shutil.disk_usage(path)` returns a named tuple of `total`, `used` and `free` bytes — the cheap pre-flight before writing something large into a temp directory you do not control the size of. `shutil.which(name)` resolves an executable the way the shell would. ## Choosing in practice Use `copy2` by default; the metadata almost never hurts and its absence eventually confuses something that compares timestamps. Use `copyfile` when you explicitly want a clean destination with your own permissions. Use `move` when the paths are on one filesystem and you want the rename; when they are not, be honest that you are doing a copy and size the operation accordingly.
- Your job moves a two-gigabyte staging file out of the temp directory and it takes minutes. Why?The temp directory is almost certainly on a different filesystem from the destination, so `shutil.move` could not use `os.rename` and fell back to copying every byte with `copy2` and then unlinking the source. The fix is to create the staging file on the destination's filesystem in the first place by passing `dir=` to `tempfile.mkstemp` or `TemporaryDirectory`, which turns the move back into a metadata-only rename.
- Does shutil.copy2 preserve the file's owner and group?No. It copies contents, permission bits, access and modification times, and where supported file flags and extended attributes — but the destination is owned by whoever ran the process. Reproducing ownership needs `os.chown`, which normally requires root. This matters when copying trees that a service account must later write to; the permission bits arriving unchanged while the owner changes can leave a directory nobody can modify.
- How do you copy a directory tree into a destination that already exists?`shutil.copytree(src, dst, dirs_exist_ok=True)`, added in Python 3.8; without it the call raises `FileExistsError`. Individual file failures during the walk are accumulated and raised together as a `shutil.Error` at the end rather than stopping at the first one, so you get the whole picture. `ignore=` skips patterns and `symlinks=True` copies links as links instead of dereferencing them.
saying these in an interview costs you the question
- Claims shutil.copy2 preserves the file's owner and group
- Thinks shutil.move is always a cheap atomic rename
- Cannot say what copy adds over copyfile
- Passes a directory to shutil.copyfile as the destination
- Rebuilds the destination path instead of using move's return value
- Hand-writes a read/write loop where copyfile is faster