skip to content

questions

5

In Git, why does deleting a file and committing not remove it from history?

level: juniorimportance: must knowfreq 55%

answer

  1. snapshots, not a change log
  2. older commits are untouched
  3. objects are named by their content
  4. reachability decides what survives
  5. every clone already has it

basics

~20 s

A Git commit only adds a new snapshot. Every earlier commit still references the blob holding that file, so the content stays reachable and clonable until history itself is rewritten and the old objects are pruned.

solid answer

~40 s

Git history is append-only in practice. A commit that removes a file records a **new** tree in which the file is absent, but it does not touch the commits before it. Those commits still point at trees that point at the blob containing the file's bytes, so `git show <old-commit>:<path>` still prints the content and every clone still carries it. Deleting the file, adding it to `.gitignore`, or deleting the branch changes none of that. To actually remove the content you must rewrite every commit that ever contained it — in practice with `git filter-repo` — which produces brand-new SHAs for the whole downstream history, and then expire reflogs and garbage-collect so the now-unreferenced objects are dropped. That is why a committed credential is a rotation problem first and a history problem second.

go deeper

for a junior

Recall that a commit is a snapshot and that earlier snapshots are untouched by a later delete. Be able to say the file is still retrievable from any older commit.

for a middle

Explain reachability: refs point to commits, commits to trees, trees to blobs, and Git keeps everything reachable. Name the rewrite plus reflog-expiry plus gc sequence that actually drops an object.

for a senior

Show the operational consequence: every clone, fork, mirror and CI cache holds the same objects, so a rewrite is a coordinated event and never an undo. Lead with invalidating whatever leaked.

for a principal

Frame it as blast radius and prevention: once content is pushed you own an exposure timeline, not a cleanup task. Argue for controls that stop the commit rather than processes that rewrite history.

## The mental model people get wrong Newcomers picture a repository as "the current files, plus a log of changes". It is closer to the opposite: a repository is a content-addressed store of immutable objects, and history is a chain of full snapshots over that store. Removing a file and committing produces one more snapshot in which the file is absent. It does not, and cannot, edit snapshots that already exist. ## The objects involved - A **blob** holds the raw bytes of one version of one file, named by the hash of its content. - A **tree** is a directory listing: names mapped to blobs and to other trees. - A **commit** points to one root tree plus its parent commit(s). Because every object is named by its own hash, an object cannot be modified in place — changing the bytes would change the name. So the only way to make history not contain something is to build a *different* history. ## Why the delete commit is not enough Say `config.yml` was added in commit A and deleted in commit Z. Commit Z's tree simply lacks that entry. Commit A's tree still has it, and commit A is still an ancestor of Z, so the blob is still **reachable** — reachable meaning: walk from any ref (branch, tag, HEAD, reflog entry) through commits, trees and blobs, and you arrive at it. Reachable objects are exactly the objects Git keeps, and exactly the objects a clone or fetch transfers. You can prove it in seconds: - `git log --all --oneline -- config.yml` still lists the commits that touched it. - `git show A:config.yml` still prints the file. - `git rev-list --objects --all` still enumerates the blob. ## What .gitignore and branch deletion do not do `.gitignore` only stops *untracked* files from being staged casually; it has no effect on content already committed and no retroactive effect at all. Deleting a branch removes one ref, but the commits may still be reachable from other refs, and even if they are not, they linger in the reflog and in packfiles until `git gc` prunes them — which by default only prunes objects past a grace period. ## What actually removes it Three things must all happen: 1. **Rewrite every commit that contained the content.** Git cannot edit a commit; it can only create a replacement with different content and a different SHA, plus replacement descendants for everything downstream. `git filter-repo` does this across the whole ref graph. 2. **Drop the old references.** Reflogs, `refs/original/` backups left by older tooling, and stale tags all keep the old commits alive. `git reflog expire --expire=now --all` removes those handles. 3. **Garbage-collect.** `git gc --prune=now` deletes objects that are now unreachable. `git filter-repo` performs steps 2 and 3 itself at the end of a run. ## The part that is not technical Even a perfect rewrite only fixes the copy you rewrote. Every other copy — teammates' clones, mirrors, forks, backup snapshots, CI caches, a laptop that has been offline for a month — still has the original objects, and each one is a complete repository. So the honest answer in an interview is: rewriting history reduces future exposure, it does not undo past exposure. If the content was a credential, it is compromised from the moment it was pushed, and the first action is to invalidate that credential; the rewrite is cleanup that follows. ## The one exception worth naming If the content is only in the commit you just made and nothing has been pushed or fetched by anyone, you are not rewriting shared history at all — you are fixing a local tip, which is cheap. The moment it is published, the cost becomes a coordinated rewrite for everyone.

  • How would you prove to a reviewer that the file really is still in history?
    Run `git log --all --oneline -- <path>` to list the commits that touched it, then `git show <commit>:<path>` to print the bytes from one of them. `git rev-list --objects --all` enumerates every reachable object and will still show the blob and its path.
  • Does adding the path to .gitignore help at all after the fact?
    No. `.gitignore` only affects untracked files going forward — it prevents re-adding the path, but it has zero retroactive effect and does not stop Git from serving the already-committed blob to any clone.
  • If the bad commit is the most recent one and nothing was pushed, is a full rewrite needed?
    No. Nothing downstream depends on it, so replacing the tip commit locally is enough, followed by expiring the reflog and garbage-collecting. The heavy tooling exists for when many commits and many clones already reference the content.

saying these in an interview costs you the question

  • Says git rm removes the file from history
  • Thinks .gitignore retroactively hides committed content
  • Believes a newer commit overwrites older blobs
  • Assumes deleting the branch deletes the content
  • Thinks garbage collection alone purges reachable objects

context

open as a page

How do you use git filter-repo to strip a path from every commit in a repo?

level: middleimportance: must knowfreq 50%

basics

~20 s

Run git filter-repo on a fresh clone with a path filter: git filter-repo --path secrets/ --invert-paths keeps everything except that path. It rewrites every commit, so all downstream SHAs change and the result must be force-pushed.

open as a page

A credential was committed to your Git repository months ago — what do you do?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Treat the credential as compromised and invalidate it first — a history rewrite cannot un-leak it. Then purge it with git filter-repo --replace-text on a fresh clone, force-push every rewritten ref, and have everyone re-clone.

open as a page

Why is git filter-branch discouraged, and what does git filter-repo do better?

level: middleimportance: should knowfreq 38%

basics

~20 s

git filter-branch forks a shell per commit, so it is extremely slow, and its defaults are unsafe: it leaves refs/original backups, keeps empty commits, and ignores tags and other refs unless told otherwise. Git's own docs now point at git filter-repo.

open as a page

How do you find and purge large blobs bloating a Git repository's clone size?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Measure with git count-objects -vH, then list the biggest objects by piping git rev-list --objects --all through git cat-file --batch-check. Purge them with git filter-repo --strip-blobs-bigger-than or a path filter, then re-measure; every downstream SHA changes.

open as a page