Why is git filter-branch discouraged, and what does git filter-repo do better?
answer
- one shell process per commit
- the defaults leave content reachable
- check what refs/original still holds
- tags are not rewritten for free
- Git's docs now point elsewhere
basics
~20 sgit filter-branch forks a shell per commit, so it is extremely slow, and its defaults are unsafe: it leaves refs/original backups, keeps empty commits, and ignores tags and other refs unless told otherwise. Git's own docs now point at git filter-repo.
solid answer
~40 s`git filter-branch` is a builtin that replays history by running your filter as a shell command **once per commit** — `--tree-filter` even checks out the whole tree each time — so a large repository takes hours. Worse are its defaults: it rewrites only the refs you name, so tags and other branches silently keep the old objects unless you add `-- --all` and `--tag-name-filter cat`; it leaves now-empty commits behind unless you pass `--prune-empty`; and it stashes the originals under `refs/original/`, which keeps every object you meant to delete alive until you remove those refs and gc. Modern Git prints a warning telling you to use `git filter-repo` instead (silenced only by `FILTER_BRANCH_SQUELCH_WARNING`). `git filter-repo` streams the object graph once, rewrites all refs, prunes empty commits, and cleans up reflogs and objects itself.
code
bash · 5 linesgit filter-branch --prune-empty --tag-name-filter cat \
--index-filter 'git rm --cached --ignore-unmatch config/prod.env' \
-- --all
git for-each-ref --format='%(refname)' refs/original | xargs -n1 git update-ref -d
git reflog expire --expire=now --all && git gc --prune=nowgo deeper
Recall that filter-branch is the old, discouraged tool and filter-repo is the current recommendation. You are not expected to have run either.
Explain the per-commit shell execution cost and name at least two unsafe defaults — refs/original backups, unrewritten tags, unpruned empty commits — and the flags that fix them.
Demonstrate that you verify a rewrite rather than trust it: check remaining refs, confirm the path is absent across all refs, and confirm objects were actually pruned before declaring done.
Own the tooling standard: pick one rewriting tool, document the runbook including the coordination step, and treat ad-hoc shell filters as an audit risk rather than a clever trick.
## Two tools, one job Both rewrite an entire repository's history. `git filter-branch` is an old builtin implemented as a shell script; `git filter-repo` is a separately installed program written for the same job with two decades of hindsight. Git's own documentation recommends against `filter-branch`, and running it prints a warning to that effect unless the `FILTER_BRANCH_SQUELCH_WARNING` environment variable is set. ## Why filter-branch is slow It walks commits one at a time and invokes your filter as a shell command for each one. With `--tree-filter` it materialises the full working tree for every commit before running your command, then re-hashes the result. With `--index-filter` it skips the checkout and manipulates the index directly — the classic incantation being `git filter-branch --index-filter 'git rm --cached --ignore-unmatch <path>' -- --all` — which is far faster but still one process spawn per commit. On a history with tens of thousands of commits that is hours versus a single streaming pass measured in minutes. ## Why its defaults are the real problem Speed is annoying; the defaults are dangerous, because each one leaves the unwanted content reachable: - **Only the named refs are rewritten.** Run it on one branch and every other branch and tag still points at the original commits, so the blob you wanted gone is still fetched by every clone. You need `-- --all` to cover all refs. - **Tags are not rewritten** unless you pass `--tag-name-filter cat`, which makes it re-create tags pointing at the new commits. - **Empty commits survive** unless you pass `--prune-empty`, littering history with commits whose entire content was the removed path. - **Backups are kept in `refs/original/`.** Those are real refs, so everything you "deleted" is still reachable and still cloned. Until you delete them, expire the reflog and run `git gc --prune=now`, nothing has actually been removed. Miss any of these and you can honestly believe the secret is gone while every fresh clone still downloads it. That failure mode is the reason the tool is deprecated in practice. ## What filter-repo changes `git filter-repo` inverts the defaults so the safe thing happens when you type nothing extra: - it rewrites **all** refs — branches, tags, notes — in one run; - it prunes commits that became empty; - it expires reflogs and garbage-collects at the end, so objects are genuinely gone from that copy; - it refuses to run outside a fresh clone unless `--force` is given, so an intact original exists; - it writes `.git/filter-repo/commit-map` so you can translate old SHAs to new ones; - it removes the `origin` remote, forcing a deliberate re-add before publishing. It also expresses common intents directly — `--invert-paths`, `--replace-text`, `--strip-blobs-bigger-than`, `--mailmap`, `--path-rename` — instead of requiring shell one-liners whose correctness you cannot easily verify. ## Where BFG fits BFG Repo-Cleaner is a third-party JVM tool from the same era that targets the common cases — delete files matching a pattern, replace text — and is far faster than `filter-branch`. It is not a Git builtin and it is not general-purpose: it deliberately protects the commit at the tip, so the current state of your project is untouched, which is fine for purging old blobs and wrong if the offending content is in the latest commit. If you can install either, `git filter-repo` is the one to reach for; the reason to know BFG is that older runbooks still reference it. ## What to say in an interview Name the two axes: performance (per-commit shell forks versus one streaming pass) and safety-by-default (leftover `refs/original/` backups, unrewritten tags, unpruned empty commits versus a tool that handles all refs and cleans up after itself). Then add the operational sentence: whichever tool you use, the rewrite only fixes the copy you ran it on.
- After a filter-branch run, why might the removed blob still be in every clone?Because `refs/original/` still points at the pre-rewrite commits, and any branch or tag you did not include still points at them too. Those are live refs, so the objects stay reachable and are transferred on clone until the refs are deleted and gc prunes them.
- What does --index-filter buy over --tree-filter?It manipulates the index directly instead of checking out the full working tree for every commit, which removes the dominant cost. It is the only filter-branch mode tolerable on a large history, and it is still far slower than a single streaming pass.
- Is BFG Repo-Cleaner a reasonable substitute?It is a third-party tool that handles the common delete-files and replace-text cases quickly, but it deliberately leaves the tip commit untouched and is not general-purpose. Where you can install either, git filter-repo is the better default.
saying these in an interview costs you the question
- Says filter-branch is fine, just slower
- Forgets refs/original keeps the old objects alive
- Rewrites one branch and calls the repo clean
- Assumes tags follow the rewrite automatically
- Thinks BFG is a Git builtin