What are the trade-offs of git subtree versus git submodule for someone who only clones your repo?
answer
- Who pays: maintainer or consumer?
- One stores content, one stores a pointer
- Plain clone versus an extra init step
- Repository size and mixed log traded away
- Version pin in tree, or only in commit messages
basics
~20 sSubtree stores the dependency's actual files in your repository, so a plain clone builds immediately; submodules store only a pinned pointer, so consumers need an extra recursive step. Subtree costs repository size and a mixed history.
solid answer
~40 sWith `git subtree` the vendored code is real content in your tree. `git clone`, `git archive`, a release tarball and a CI checkout all get it with no extra command, and nobody lands in a detached HEAD inside a nested repository. The cost is on your side: your repository carries the dependency's bytes (and, unless you imported with `--squash`, its commits), your `git log` mixes two projects, and the dependency's exact version is only discoverable from commit messages. A submodule keeps your repository small and pins an explicit upstream SHA that is visible in the tree, but every consumer must remember `git clone --recurse-submodules` or `git submodule update --init`, and forgetting it is the classic broken-onboarding bug. Choose subtree when consumers matter more than repository hygiene.
go deeper
Recall the one-sentence contrast: subtree puts the real files in your repo so a plain clone works, while a submodule stores a pointer that consumers must initialise.
Explain both costs concretely — repository growth and interleaved history on the subtree side, the empty-directory-after-clone failure and detached HEAD on the submodule side.
Frame it as who absorbs the pain. Show you would count consumers and tooling that only does plain checkouts, and mention that with subtree the version pin lives in commit metadata rather than in the tree.
Own the standard: pick a default for the organisation, say what evidence would change it, and account for the one-off migration cost of switching a path from one mechanism to the other across every existing checkout.
## Two ways to depend on another repository Git offers two built-in mechanisms for keeping another project's code inside yours, and they sit at opposite ends of one trade-off. **Submodule**: your tree contains a special entry — a *gitlink* — that records the path plus one commit SHA of the other repository, with the URL stored in a tracked `.gitmodules` file. The other project's objects live in their own repository, fetched separately. **Subtree**: the other project's files are copied into a directory of your repository as ordinary blobs and trees, imported by `git subtree add --prefix=<dir>` and refreshed by `git subtree pull --prefix=<dir>`. ## What the consumer experiences This is the axis interviewers care about, because it is where teams get burned. With a subtree, someone who runs `git clone <your-repo>` has a complete, buildable checkout. So does `git archive`, so does a downloaded snapshot, so does a CI job that does a plain checkout, and so does a colleague who has never heard of subtree. Nothing can be forgotten because there is nothing to remember. With a submodule, a plain clone produces an empty directory where the dependency should be. The consumer needs `git clone --recurse-submodules`, or `git submodule update --init` after the fact. Every tool in the chain — CI checkout steps, release tarball generation, IDE clone helpers — must also be taught to recurse. Once inside a submodule the working tree sits at a detached HEAD by default, which reliably confuses people who edit there. ## What each costs the maintainer Subtree costs you repository weight and history clarity: - The dependency's bytes are in your object database forever, and every clone of your repository pays for them. - Imported without `--squash`, the upstream project's commits become ancestors of yours, so `git log`, `git shortlog` and even `git bisect` traverse someone else's work. - The pinned version is not visible in the tree. It exists only in the `git-subtree-split: <sha>` line of import commits, so "which version of this library are we on?" is a history question, not a file question. - Sending local fixes upstream needs a deliberate `git subtree split` / `git subtree push`, not a normal push. Submodules cost you consumer friction, but buy clarity: - The pinned SHA is an object in your tree, so a diff literally shows the version bump as a one-line change. - Your repository stays small; the dependency's history is never yours. - Working *in* the dependency is natural, because it really is a separate repository with its own branches. ## Choosing Prefer subtree when the population of consumers is large or uncontrolled: open-source projects that want `git clone && make` to work, teams whose CI or release process only ever does a plain checkout, code that ships as a source tarball, or a small vendored dependency you rarely update. Prefer submodules when the dependency is big, actively co-developed by the same people, or needs an obvious auditable version pin — and when you can guarantee everyone's tooling recurses. A useful decision heuristic: subtree moves the pain to the maintainer, submodule moves it to the consumer. Count consumers. ## Switching between them Going from submodule to subtree is straightforward: remove the submodule and its `.gitmodules` entry, commit, then `git subtree add --prefix=<same-dir>` the same URL and ref. Going the other way means deleting the vendored directory, committing, then `git submodule add`. Neither direction is free — both rewrite what the tree looks like at that path, and both mean everyone with an existing checkout does a cleanup step, so the switch is a decision you want to make once. ## What weak answers miss The answer "subtree is easier" is not enough on its own. The complete answer names *who* it is easier for and what that ease is paid for with: plain clones work, at the price of a heavier repository, a mixed history, and a version pin visible only from commit metadata. And it names the failure mode on the other side by its symptom — an empty directory after a plain clone, followed by a build that cannot find the dependency.
- With a subtree, how does a reviewer find out which upstream version is currently vendored?From history, not from files. The most recent subtree commit for that prefix carries a `git-subtree-split: <sha>` line naming the upstream commit that was imported, so you search the log for that prefix. This is strictly worse than a submodule, where the pinned SHA is an object in the tree and shows up directly in a diff.
- Which choice is friendlier to a shallow clone or an export of the repository?Subtree. `git archive` and a plain export contain the vendored files because they are ordinary tree entries, and a shallow clone still gets the current content. A submodule contributes only a gitlink, so an archive of the superproject has an empty directory at that path and the consumer must fetch the other repository separately.
- Does subtree remove the need to update the dependency?No. Updates are still an explicit act — `git subtree pull --prefix=<dir> <repo> <ref>` — and it can conflict if you have edited files inside the prefix. The difference from a submodule is only that consumers get whatever version you last imported without doing anything, not that the import happens by itself.
saying these in an interview costs you the question
- Says subtree keeps the repository small like a submodule does
- Claims a plain clone of a submodule superproject includes the code
- Thinks subtree pulls upstream updates automatically
- Believes the subtree's pinned version is visible as a tree entry
- Treats the choice as pure preference with no consumer impact