skip to content

What does the --squash flag change about a git subtree add or pull import?

level: middleimportance: should knowfreq 35%

answer

  1. Content without the history
  2. One commit per import
  3. Message records the imported upstream SHA
  4. Whatever you choose, stay consistent
  5. Blame inside the prefix pays the price

basics

~10 s

With --squash, Git collapses the upstream commits being imported into one synthetic commit before merging, so your history gains a single snapshot entry instead of the other project's entire commit history.

solid answer

~40 s

Without `--squash`, `git subtree add` and `git subtree pull` merge the upstream branch itself, which makes every one of its commits an ancestor of yours — `git log` then interleaves two projects. With `--squash`, Git first synthesises one commit representing the upstream tree, with a message like `Squashed 'vendor/lib/' content from commit <sha>`, and merges that instead. Your log gains one entry per import, and the upstream's individual commits are not reachable from your branch. The synthetic commit still carries `git-subtree-dir` and `git-subtree-split` lines, which is how a later pull knows what was last imported. The practical rule is consistency: if you imported with `--squash`, pass `--squash` on every subsequent pull for that prefix, because mixing modes produces confusing merges and can reintroduce the history you were avoiding.

code

bash · 5 lines
bash
git subtree add --prefix=vendor/lib liblog v1.4.0 --squash
# ... later ...
git subtree pull --prefix=vendor/lib liblog v1.5.0 --squash

git log --oneline --grep='git-subtree-dir: vendor/lib'

go deeper

for a junior

Recall the headline: --squash imports the code as one commit instead of dragging the other project's whole commit history into your log.

for a middle

Explain the mechanism — a synthetic commit carrying the git-subtree-split line, merged in place of the upstream branch — and name what you give up: per-commit blame and bisect inside the prefix.

for a senior

Demonstrate the operating discipline: one mode per prefix for its whole life, documented, and a recovery story for investigating regressions by diffing the recorded upstream SHAs in the upstream repository.

for a principal

Own the convention across repositories: decide whether the organisation keeps third-party history at all, what that costs every clone, and how provenance or licence attribution is preserved when the upstream authors' commits are deliberately excluded.

## The problem --squash solves When you vendor another repository with `git subtree add --prefix=<dir> <repo> <ref>`, the default behaviour is a genuine merge: the fetched upstream commit becomes the second parent of the import commit. Because merge parents are ancestors, every commit that project ever made is now part of *your* repository's history. For a library with ten years of development that means thousands of foreign commits in `git log`, foreign author names in `git shortlog`, and foreign commits traversed by `git bisect`. The objects for that whole history are also in your object database, so every clone downloads them. `--squash` exists to import the *content* without the *history*. ## What --squash actually does With `--squash`, Git builds one new commit whose tree is the upstream tree, and whose message looks like: `Squashed '<dir>/' content from commit <short-sha>` followed by the machine-readable lines `git-subtree-dir: <dir>` and `git-subtree-split: <upstream-sha>`. That synthetic commit is then merged into your branch. The upstream's real commits are never made ancestors of your branch, so they do not appear in `git log`, and only the objects needed for the current snapshot are pulled in rather than the full history. On a later `git subtree pull --prefix=<dir> <repo> <ref> --squash`, Git reads the `git-subtree-split` line of the previous squash commit to learn which upstream commit you are currently at, synthesises a new squashed commit representing the delta from there to the new tip, and merges that. So each update is one commit in your history rather than the batch of upstream commits it contains. ## The trade-offs What you gain: - A readable log: your project's commits, plus one clearly-labelled import line per update. - A smaller repository: you are not carrying a decade of someone else's objects. - Cleaner `git bisect` and `git shortlog` over your own work. What you lose: - Granularity. You cannot `git log` your way to which upstream commit introduced a bug in the vendored code — you only have snapshots. The upstream repository is still the place to investigate. - `git blame` inside the prefix attributes lines to the squash commits, not to the original authors. - Attribution generally: the upstream authors' commits are not in your history at all, which some projects care about for provenance or licensing bookkeeping. ## The consistency rule The most common practical failure is mixing modes on one prefix. If the initial add used `--squash` and a later pull omits it, that pull merges the real upstream branch — which pulls the full upstream history in after all, defeating the earlier squash and producing a history that is half one shape and half the other. The guidance is simple: choose a mode per prefix at import time, write it down (a line in the prefix's README or in your contributing docs), and always pass the same flag. Wrapping the command in a small script or a Git alias is a reasonable way to make that mechanical. ## Interaction with splitting changes back out `git subtree split` and `git subtree push` reconstruct a prefix-only history to send upstream. They work with squashed imports because the squash commits record `git-subtree-split: <sha>`, giving the tool the upstream commit your local changes are based on. What they cannot do is recover the upstream detail you chose not to import — the reconstructed history contains *your* commits touching the prefix, not the upstream ones. ## Choosing Use `--squash` for third-party dependencies you consume but do not co-develop: you want the code, you do not want their history, and the upstream repository remains the authoritative place to study it. Skip `--squash` when the vendored project is genuinely part of your work — for example code you also maintain and regularly push back — and you want a single unified history in which upstream and local commits are directly comparable and bisectable. A short way to say it in an interview: `--squash` trades per-commit forensics inside the vendored directory for a clean, small history in your own repository, and whichever you pick you must stay consistent for the life of that prefix.

  • What breaks if you use --squash on the initial add but omit it on a later pull?
    That pull merges the real upstream branch, so all of the upstream history you deliberately avoided becomes an ancestor of your branch after all. You end up with a hybrid history — one squash commit followed by the full upstream log — which defeats the point and confuses later readers. Pick a mode per prefix and always pass the same flag.
  • With squashed imports, how do you investigate which upstream change broke the vendored code?
    Not from your repository — the individual upstream commits are not there. Read the `git-subtree-split` lines of the two squash commits that bracket the regression to get the old and new upstream SHAs, then bisect or diff that range in the upstream repository itself.
  • Does --squash reduce the size of the vendored files themselves?
    No. The current content is stored in full either way; only the historical versions are avoided. If the dependency is large because its current tree is large, squashing does not help — that is a different problem, addressed by vendoring less of it or not vendoring at all.

saying these in an interview costs you the question

  • Thinks --squash compresses or shrinks the vendored files
  • Believes --squash is only meaningful on the initial add
  • Says you can still bisect upstream commits after a squashed import
  • Assumes mixing squashed and non-squashed pulls is harmless
  • Confuses it with squashing your own commits before merging

context