skip to content

questions

13

After cloning a repo with Git submodules, why are the submodule directories empty?

level: juniorimportance: must knowfreq 65%

answer

  1. A clone brings only what the repo stores
  2. The parent stores references, not content
  3. One flag at clone time does both steps
  4. --recurse-submodules, or update --init

basics

~20 s

A plain clone fetches only the superproject, which records submodules as commit references rather than files, so it creates the directories but leaves them empty. Clone with --recurse-submodules, or run git submodule update --init --recursive afterwards.

solid answer

~40 s

The superproject stores each submodule as a gitlink — a single commit id — not as content, so cloning it transfers no submodule files. Git creates the empty directory at that path and stops. Two fixes: clone with `git clone --recurse-submodules <url>`, which clones the superproject and then initialises and checks out every submodule in one step; or, in an existing clone, run `git submodule update --init --recursive`, where `--init` copies the URLs from `.gitmodules` into local config and `--recursive` handles submodules nested inside submodules. You can check the state with `git submodule status`: a leading `-` means not initialised. Setting `submodule.recurse` to `true` locally makes commands such as `git pull` and `git checkout` update submodules automatically, which prevents the recurring version of this problem.

code

console · 9 lines
console
$ git clone https://example.invalid/app.git
$ ls vendor/liblog
$ git submodule status
-9fceb02d0ae598e95dc970b74767f19372d61af8 vendor/liblog

$ git submodule update --init --recursive
Submodule 'vendor/liblog' registered for path 'vendor/liblog'
$ git submodule status
 9fceb02d0ae598e95dc970b74767f19372d61af8 vendor/liblog (v1.4.0)

go deeper

for a junior

Recall the two commands: clone with --recurse-submodules, or fix an existing clone with git submodule update --init --recursive. Know that the parent repository stores a reference, so nothing arrives without asking.

for a middle

Explain what each part does — init copies URLs from the tracked mapping into local config, update checks out the recorded commit, recursive handles nesting — and read the markers in git submodule status.

for a senior

Anticipate the recurrence: pulls and branch switches leave submodules behind unless recursion is enabled, so make the bootstrap idempotent and make the build fail loudly on an empty submodule path rather than with a confusing toolchain error.

for a principal

Treat this as an onboarding-cost decision. Submodules impose a predictable, permanent tax on every new clone and every branch switch; decide whether the exact pinning is worth it and, if it is, make the tooling absorb the tax.

## Why the directory is empty A submodule is recorded in the superproject as a gitlink: one tree entry naming a commit in a *different* repository. None of that repository's objects are stored in the superproject. So `git clone <superproject>` downloads exactly the superproject's objects, checks out its tree, and creates the submodule path as an empty directory — there is simply nothing to put in it yet. This is the single most common submodule complaint, and it typically surfaces as a build failure: headers missing, an import that resolves to nothing, a directory the toolchain expected to contain a project. ## The one-step fix at clone time `git clone --recurse-submodules <url>` clones the superproject and then, for every submodule, copies the URL from `.gitmodules` into local config, clones it, and checks out the exact commit recorded in the gitlink. It recurses into nested submodules. The older spelling `--recursive` still works as an alias. You can also add `--shallow-submodules` to fetch the submodules with limited history when you only need the pinned commit's content, and `--jobs <n>` to fetch several submodules in parallel. ## The fix after the fact If you have already cloned, run from the superproject: `git submodule update --init --recursive` Three parts, each doing something distinct: - **`--init`** performs the initialisation step: reads `.gitmodules` and copies each submodule's URL into your local `.git/config` as `submodule.<name>.url`. Without it, `update` skips submodules it considers uninitialised. It is the same work `git submodule init` does on its own. - **`update`** fetches the submodule if needed and checks out the commit recorded in the superproject's gitlink. - **`--recursive`** applies the whole thing to submodules nested inside submodules. Running it is idempotent, so it is safe to put in a project's bootstrap script. ## Diagnosing the state `git submodule status` prints one line per submodule with a leading marker: - `-<sha> path` — not initialised, i.e. no working-tree content. This is what you see right after a plain clone. - ` <sha> path` (leading space) — checked out at exactly the recorded commit; everything is as it should be. - `+<sha> path` — the submodule's checked-out commit differs from the one recorded in the superproject. Either you moved it deliberately, or you have a pending bump to commit. - `U<sha> path` — the submodule has unresolved merge conflicts. ## Making it not happen again The same emptiness reappears every time the superproject moves to a commit whose gitlink differs from what is on disk — after a `git pull`, after switching branches. By default Git leaves the submodule where it was and reports the difference; it does not silently move it. Setting `submodule.recurse` to `true` in local config makes the commands that touch the working tree — including `git pull` and `git checkout` — recurse into submodules and update them to the recorded commits. It does not cover initialising brand-new submodules that appeared, so the bootstrap command is still worth keeping. Related useful settings: `status.submoduleSummary` for `git status` to summarise submodule movement, and `fetch.recurseSubmodules` to control whether fetching the superproject also fetches submodule objects. ## What good onboarding looks like Because this problem is entirely predictable, teams that use submodules successfully do three things: they document the recursive clone command instead of a plain one; they provide a bootstrap script that runs `git submodule update --init --recursive` and is safe to re-run; and they make the build fail with an explicit message when a submodule path is empty, rather than with a confusing compiler error. None of that is Git functionality — it is the operational discipline the data model requires. ## Related cleanups `git submodule deinit <path>` reverses initialisation: it removes the working-tree content and the local config entry, leaving the gitlink untouched. It is the right way to temporarily stop carrying a submodule you do not need, and `--force` handles the case where the submodule's working tree has local modifications.

  • What does --init add in git submodule update --init?
    It performs the initialisation step: reading `.gitmodules` and copying each submodule's URL into your local `.git/config` under `submodule.<name>.url`. Without it, `update` skips any submodule it considers uninitialised, so a fresh clone would still end up with an empty directory. It is the same work `git submodule init` does, folded into one command.
  • After a git pull, a submodule still points at the old commit. Why doesn't Git move it?
    By default Git updates the superproject's working tree, including the recorded gitlink in the index, but leaves the submodule's own checkout alone — moving it could discard work you have in progress there. `git submodule status` shows it as `+` to flag the difference. Run `git submodule update`, or set `submodule.recurse` to true so pulls and checkouts recurse automatically.
  • How do you tell from the command line whether a submodule is initialised?
    `git submodule status` prints a leading marker per submodule: `-` means not initialised, a leading space means it is checked out at exactly the recorded commit, `+` means the checked-out commit differs from the recorded one, and `U` means unresolved conflicts. It is the fastest first diagnostic when a submodule-related build failure appears.

saying these in an interview costs you the question

  • Thinks the clone failed or the network dropped files
  • Says the submodule content is missing from the remote
  • Reaches for deleting and re-cloning instead of initialising
  • Believes git pull always updates submodules by default
  • Confuses initialising a submodule with adding a new one

context

open as a page

What does git worktree add create, and how does it differ from a second git clone?

level: juniorimportance: must knowfreq 45%

basics

~20 s

It creates an additional working directory with its own checked-out branch, backed by the same repository. Unlike a second clone it shares one object database and one set of branches, so it costs no extra fetch and no duplicated history.

open as a page

What are the trade-offs of git subtree versus git submodule for someone who only clones your repo?

level: middleimportance: must knowfreq 48%

basics

~20 s

Subtree stores the dependency's actual files in your repository, so a plain clone builds immediately; submodules store only a pinned pointer, so consumers need an extra recursive step. Subtree costs repository size and a mixed history.

open as a page

Which repository state do Git worktrees share, and which is private to each worktree?

level: middleimportance: must knowfreq 40%

basics

~10 s

Objects, refs, remotes and repository config are shared by all worktrees. Each worktree keeps its own working directory, index, HEAD, HEAD reflog and in-progress operation state, stored in .git/worktrees/<name> inside the original repository.

open as a page

What does git subtree add --prefix do to your repository's files and history?

level: juniorimportance: should knowfreq 30%

basics

~20 s

It fetches another repository and copies its content into the given directory as ordinary tracked files, then records the import as a commit in your history, so a plain clone of your repo already contains that code.

open as a page

Why is HEAD detached inside a Git submodule, and what does that risk for work committed there?

level: middleimportance: should knowfreq 50%

basics

~20 s

The superproject records an exact commit, not a branch, so submodule update checks out that commit directly and leaves HEAD detached. Commits made there sit on no branch and are easy to lose on the next update — check out a branch before editing.

open as a page

What does git submodule update --remote do that plain git submodule update does not?

level: middleimportance: should knowfreq 35%

basics

~20 s

Plain update checks out the exact commit the superproject has pinned. With --remote, Git instead fetches the submodule and checks out the tip of its configured branch, leaving a changed pin in your working tree that you still have to commit.

open as a page

What does the --squash flag change about a git subtree add or pull import?

level: middleimportance: should knowfreq 35%

basics

~10 s

With --squash, Git collapses the upstream commits being imported into one synthetic commit before merging, so your history gains a single snapshot entry instead of the other project's entire commit history.

open as a page

Why does Git refuse to check out the same branch in two worktrees, and what do you do instead?

level: middleimportance: should knowfreq 38%

basics

~20 s

All worktrees share one set of refs, so a branch is a single pointer. Two checkouts advancing it independently would silently clobber each other, so Git allows a branch in only one worktree. Use a new branch or a detached checkout instead.

open as a page

A teammate's clone reports a Git submodule commit it cannot fetch. What went wrong?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The superproject was pushed with a gitlink pointing at a submodule commit that exists only locally. Fetching the submodule finds no such object, so the checkout fails. Fix it by pushing the submodule commit; prevent it with push.recurseSubmodules.

open as a page

How do you send changes made inside a git subtree prefix back to the upstream repository?

level: seniorimportance: should knowfreq 28%

basics

~20 s

Use git subtree split to synthesise a history containing only the commits that touched that prefix, with the prefix stripped from the paths, then push that synthetic branch upstream — git subtree push does both steps in one command.

open as a page

A Git worktree's directory was deleted with rm -rf — what stale state remains and how do you clean it up?

level: seniorimportance: should knowfreq 26%

basics

~20 s

The repository still holds that worktree's administrative directory under .git/worktrees, so it appears in git worktree list as prunable and still occupies its branch. Run git worktree prune to drop the bookkeeping, or remove worktrees with git worktree remove.

open as a page