How do you find and purge large blobs bloating a Git repository's clone size?
answer
- clone size is history, not checkout
- measure before you rewrite
- rev-list piped into cat-file batch-check
- a size threshold or a path glob
- re-measure to prove it worked
basics
~20 sMeasure with git count-objects -vH, then list the biggest objects by piping git rev-list --objects --all through git cat-file --batch-check. Purge them with git filter-repo --strip-blobs-bigger-than or a path filter, then re-measure; every downstream SHA changes.
solid answer
~40 sA repository is slow to clone because history is heavy, not because the checkout is. Start by measuring: `git count-objects -vH` reports the packed size, and `git rev-list --objects --all | git cat-file --batch-check='%(objecttype) %(objectname) %(objectsize) %(rest)'` gives every reachable object with its size and path, which you sort to find the offenders. `git filter-repo --analyze` produces the same picture as reports under `.git/filter-repo/analysis/`. Then purge: `git filter-repo --strip-blobs-bigger-than 10M` drops oversized blobs wherever they appear, or use `--path-glob '*.psd' --invert-paths` when a file type is the problem. Both rewrite every commit, so all downstream SHAs change and every clone must be replaced. Afterwards prevent recurrence — otherwise the same binaries land again next quarter.
code
bash · 4 linesgit count-objects -vH
git rev-list --objects --all |
git cat-file --batch-check='%(objecttype) %(objectname) %(objectsize) %(rest)' |
awk '$1=="blob"' | sort -k3 -n | tail -20go deeper
Recall that clone size comes from all of history, so a large file stays costly after it is deleted, and that removing it requires rewriting history.
Explain the measurement path — count-objects for totals, rev-list piped into cat-file for per-object sizes — and the size and path filters that purge the offenders.
Show judgment about whether the rewrite is worth it, plan the coordination and the external SHA references, verify the size actually dropped, and add prevention so it does not recur.
Own the tradeoff across teams: repository splits, artefact storage strategy, clone-time budgets for CI, and whether a once-off coordinated rewrite beats living with the weight.
## Why the checkout size misleads you People judge repository weight by the working tree, but a clone transfers **history**: every version of every file that is still reachable. A 40 MB binary replaced twenty times is roughly 800 MB of objects even though the checkout shows one 40 MB file, because each version is a distinct blob and compressed binaries do not delta-compress well against each other. That is why deleting the file today does nothing for clone times. ## Measuring first Never rewrite before you know what you are rewriting. - `git count-objects -vH` reports loose and packed object counts and total pack size — the number that correlates with clone time. - `git rev-list --objects --all` lists every reachable object and, for blobs, the path it was seen at. - Piping that into `git cat-file --batch-check='%(objecttype) %(objectname) %(objectsize) %(rest)'` yields type, SHA, size and path per line; sorting numerically on size and taking the tail shows the worst offenders and where they came from. - `git filter-repo --analyze` does this for you and writes size-by-path and size-by-blob reports under `.git/filter-repo/analysis/`, usually the faster route to a decision. What you are looking for is the shape of the problem: one accidental commit of a build artefact, a whole directory of fixtures, or a file type that keeps coming back. ## Purging On a fresh clone, pick the filter that matches the shape: - **By size**, when the offenders are unrelated one-offs: `git filter-repo --strip-blobs-bigger-than 10M` removes every blob above the threshold from all of history, regardless of path. - **By path**, when a directory or file type is the cause: `git filter-repo --path assets/renders/ --invert-paths` or `--path-glob '*.iso' --invert-paths`. Both rewrite every commit that contained the content and everything downstream, rewrite all branches and tags, prune commits that became empty, then expire reflogs and garbage-collect. Re-measure with `git count-objects -vH` to confirm the drop before publishing — if the size barely moved, your filter missed the real offenders and you should return to the analysis. ## The publishing cost This is where a size cleanup stops being a technical exercise. Every commit SHA from the earliest rewritten commit onward is new, so: - every clone in the organisation must be replaced, not pulled; - SHAs recorded outside the repository — deployment records, release notes, issue references, submodule pointers — now name commits that exist only in old copies; - signatures on rewritten commits no longer verify; - the force-push must cover every rewritten ref including tags. Budget for the coordination, pick a quiet window, and tell people before rather than after. ## Alternatives to rewriting Sometimes the rewrite is not worth it. Options to weigh: - **Leave history alone and stop the growth.** If clone time is tolerable, preventing new large blobs may be enough. - **Shallow or partial clones for consumers** who do not need full history — a per-consumer fix that requires no rewrite and breaks nobody, though it does not shrink the repository itself. - **Move binary assets out of the repository** entirely, into an artefact store or Git LFS, and rewrite once as part of that migration so the coordination cost is paid a single time. - **Archive and restart** for a repository whose history has little ongoing value: keep the old repository read-only and begin a fresh one. ## Preventing recurrence After a rewrite, add ignore rules for build outputs and a size check that runs before content can be published; otherwise the same binaries return and you will run the same painful rewrite next year. An interviewer listens for exactly that closing thought — the rewrite is the remediation, the prevention is the fix.
- Why does one 40 MB binary committed twenty times cost far more than a 40 MB source tree?Each committed version is a separate blob, and compressed binaries delta-compress poorly against each other, so the pack holds close to twenty full copies. Text files with small edits delta well, so their history is a fraction of size times revisions.
- Can you reduce clone pain without rewriting history at all?Yes, for consumers: shallow and partial clones fetch less history, and moving future binaries out of the repository stops the growth. Neither shrinks the repository itself, but both avoid breaking every existing clone and stored SHA.
- How do you confirm the purge actually worked before force-pushing?Re-run `git count-objects -vH` and compare the packed size, and re-run the rev-list plus cat-file listing to confirm the offending blobs are gone. If the size did not move, the filter missed the real content and the analysis should be redone.
saying these in an interview costs you the question
- Judges repo weight by the working tree size
- Rewrites before measuring what is actually large
- Thinks a delete commit reclaims the space
- Skips repacking and re-measuring afterwards
- Ignores that stored commit SHAs become invalid