skip to content

What does git subtree add --prefix do to your repository's files and history?

level: juniorimportance: should knowfreq 30%

answer

  1. Copy, not pointer
  2. Real files land in your tree
  3. History joins as a merge commit
  4. git-subtree-dir and git-subtree-split lines

basics

~20 s

It fetches another repository and copies its content into the given directory as ordinary tracked files, then records the import as a commit in your history, so a plain clone of your repo already contains that code.

solid answer

~40 s

`git subtree add --prefix=vendor/lib <repo-or-remote> <ref>` runs from the top level of your working tree. Git fetches the named ref from that repository, places its tree under `vendor/lib`, and commits the result. Without `--squash` the import is a merge commit whose second parent is the fetched upstream commit, so the upstream's individual commits become part of your history; with `--squash` you get one synthetic commit representing that snapshot instead. Either way the files afterwards are normal blobs in your tree — no pointer, no extra init step. Anyone who clones or exports an archive of your repo gets the vendored code automatically. The prefix directory must not already exist, and the working tree must be clean.

code

bash · 6 lines
bash
git remote add liblog https://git.internal/liblog.git
git subtree add --prefix=vendor/liblog liblog main

# the code is now ordinary tracked content
git ls-files vendor/liblog | head
git log --oneline -1

go deeper

for a junior

Recall the shape of the command and its headline effect: real files under a directory you choose, and a plain clone just works with no extra step.

for a middle

Be ready to describe the commit it creates — a merge whose second parent is the fetched upstream tip, or one synthetic commit with --squash — and where the git-subtree-dir and git-subtree-split metadata lives.

for a senior

Explain the operational consequences you would flag in review: repository growth, interleaved foreign commits in the log, and the fact that the vendoring relationship is discoverable only from commit messages.

for a principal

Own the policy angle for Git specifically: whether importing a third party's full history into your object database is acceptable, and what convention (squash mode, one prefix per dependency, a README marker) you standardise on so future maintainers can follow the imports.

## What the command is for `git subtree` is a way to *vendor* another repository — to keep a copy of its code inside your own repository, under a directory you choose, while still being able to pull newer upstream versions later. The command ships with Git in its `contrib` area and is present in most distributions' Git packages. The entry point is: `git subtree add --prefix=<dir> <repository> <ref>` `--prefix` (short form `-P`) names the directory the foreign code will live in. `<repository>` is a URL or the name of a remote you have already configured. `<ref>` is the branch or tag in that repository you want to import. The command insists on being run from the top level of the working tree, and it refuses to run if `<dir>` already exists or if you have uncommitted changes. ## What happens to your files Git fetches the requested ref, takes the tree of that commit, and grafts it into your tree at `<dir>`. Afterwards those files are ordinary tracked files: `git status` sees them, `git grep` finds them, your build reads them straight off disk. There is no special object type involved and nothing that a later `git clone` has to be told to resolve. That is the entire point of the technique — a consumer who does nothing but `git clone` gets a working checkout. ## What happens to your history This is where people are surprised. Two shapes are possible. **Without `--squash`**, the import is a **merge commit with two parents**: your previous `HEAD` and the fetched upstream commit. Because the upstream commit is now an ancestor of yours, every commit from that other project becomes part of your repository's history. `git log` will show them interleaved with your own work, and their file paths are the paths that project used — usually *not* prefixed by `<dir>` — because those old commits are not rewritten. Only the merge places the content at the prefix. **With `--squash`**, Git collapses the upstream history into a single synthetic commit whose message reads like `Squashed '<dir>/' content from commit <short-sha>`, and merges that. Your history gains one commit per import instead of hundreds, and the foreign project's commits are not reachable from your branch. ## The bookkeeping Git leaves behind Subtree commits carry machine-readable lines at the end of the commit message: `git-subtree-dir: <dir>` and `git-subtree-split: <upstream-sha>` These are how later `git subtree pull` and `git subtree split` invocations discover which prefix an import belongs to and which upstream commit was last imported. If someone rewrites or drops those commit messages, subtree loses its memory of the relationship and later operations get confused. There is no separate config file recording the link — unlike `.gitmodules` for submodules, the relationship lives only in commit messages, so you cannot see at a glance from a working tree which directories are vendored. ## Everyday consequences - Updates are deliberate: `git subtree pull --prefix=<dir> <repo> <ref>` merges newer upstream code in and can conflict like any merge, which you resolve normally. - Local edits inside the prefix are just edits — nothing stops you making them, and nothing sends them upstream automatically. - Your repository grows by the size of the vendored content, and (without `--squash`) by the whole imported history. - Because file paths inside the prefix did not exist in the upstream commits, `git log <dir>` and `git blame` across the import boundary can look discontinuous. ## Common mistakes Running the command from a subdirectory fails outright. Forgetting `--prefix` and passing a bare path fails as well — the option is required. Adding into an existing directory fails; if you already copied the code in by hand you must remove it and commit that first. And choosing `--squash` on the initial add but omitting it on later pulls (or vice versa) produces awkward history, so pick one mode per prefix and keep to it.

  • How do you bring in newer upstream code after the initial import?
    Run `git subtree pull --prefix=<dir> <repo> <ref>` from the top level. Git fetches the ref and merges the new upstream content into the prefix; conflicts inside that directory are resolved exactly like any merge conflict. If the original import used `--squash`, keep passing `--squash` on every pull so the history stays consistent.
  • How can a reviewer tell that a directory is a vendored subtree rather than hand-written code?
    Only from history. Look for commits whose messages end with `git-subtree-dir: <dir>` and `git-subtree-split: <sha>`, or for a merge commit that introduced the whole directory at once. There is no tracked manifest equivalent to `.gitmodules`, which is why teams usually add a short README note in the prefix.
  • Does git subtree add require the two repositories to share history?
    No. The upstream is unrelated to your project, and subtree handles the graft itself, so you do not need `--allow-unrelated-histories`. That is also why the merge is unusual: the two sides have no merge base, and the content is placed at the prefix rather than at the root.

saying these in an interview costs you the question

  • Claims subtree stores a pointer, so clones need an extra init step
  • Thinks upstream commits get their paths rewritten under the prefix
  • Believes subtree pulls new upstream code automatically on clone
  • Says it can be run from any subdirectory of the repo
  • Confuses the prefix directory with a Git remote name

context