After cloning a repo with Git submodules, why are the submodule directories empty?
answer
- A clone brings only what the repo stores
- The parent stores references, not content
- One flag at clone time does both steps
- --recurse-submodules, or update --init
basics
~20 sA plain clone fetches only the superproject, which records submodules as commit references rather than files, so it creates the directories but leaves them empty. Clone with --recurse-submodules, or run git submodule update --init --recursive afterwards.
solid answer
~40 sThe superproject stores each submodule as a gitlink — a single commit id — not as content, so cloning it transfers no submodule files. Git creates the empty directory at that path and stops. Two fixes: clone with `git clone --recurse-submodules <url>`, which clones the superproject and then initialises and checks out every submodule in one step; or, in an existing clone, run `git submodule update --init --recursive`, where `--init` copies the URLs from `.gitmodules` into local config and `--recursive` handles submodules nested inside submodules. You can check the state with `git submodule status`: a leading `-` means not initialised. Setting `submodule.recurse` to `true` locally makes commands such as `git pull` and `git checkout` update submodules automatically, which prevents the recurring version of this problem.
code
console · 9 lines$ git clone https://example.invalid/app.git
$ ls vendor/liblog
$ git submodule status
-9fceb02d0ae598e95dc970b74767f19372d61af8 vendor/liblog
$ git submodule update --init --recursive
Submodule 'vendor/liblog' registered for path 'vendor/liblog'
$ git submodule status
9fceb02d0ae598e95dc970b74767f19372d61af8 vendor/liblog (v1.4.0)go deeper
Recall the two commands: clone with --recurse-submodules, or fix an existing clone with git submodule update --init --recursive. Know that the parent repository stores a reference, so nothing arrives without asking.
Explain what each part does — init copies URLs from the tracked mapping into local config, update checks out the recorded commit, recursive handles nesting — and read the markers in git submodule status.
Anticipate the recurrence: pulls and branch switches leave submodules behind unless recursion is enabled, so make the bootstrap idempotent and make the build fail loudly on an empty submodule path rather than with a confusing toolchain error.
Treat this as an onboarding-cost decision. Submodules impose a predictable, permanent tax on every new clone and every branch switch; decide whether the exact pinning is worth it and, if it is, make the tooling absorb the tax.
## Why the directory is empty A submodule is recorded in the superproject as a gitlink: one tree entry naming a commit in a *different* repository. None of that repository's objects are stored in the superproject. So `git clone <superproject>` downloads exactly the superproject's objects, checks out its tree, and creates the submodule path as an empty directory — there is simply nothing to put in it yet. This is the single most common submodule complaint, and it typically surfaces as a build failure: headers missing, an import that resolves to nothing, a directory the toolchain expected to contain a project. ## The one-step fix at clone time `git clone --recurse-submodules <url>` clones the superproject and then, for every submodule, copies the URL from `.gitmodules` into local config, clones it, and checks out the exact commit recorded in the gitlink. It recurses into nested submodules. The older spelling `--recursive` still works as an alias. You can also add `--shallow-submodules` to fetch the submodules with limited history when you only need the pinned commit's content, and `--jobs <n>` to fetch several submodules in parallel. ## The fix after the fact If you have already cloned, run from the superproject: `git submodule update --init --recursive` Three parts, each doing something distinct: - **`--init`** performs the initialisation step: reads `.gitmodules` and copies each submodule's URL into your local `.git/config` as `submodule.<name>.url`. Without it, `update` skips submodules it considers uninitialised. It is the same work `git submodule init` does on its own. - **`update`** fetches the submodule if needed and checks out the commit recorded in the superproject's gitlink. - **`--recursive`** applies the whole thing to submodules nested inside submodules. Running it is idempotent, so it is safe to put in a project's bootstrap script. ## Diagnosing the state `git submodule status` prints one line per submodule with a leading marker: - `-<sha> path` — not initialised, i.e. no working-tree content. This is what you see right after a plain clone. - ` <sha> path` (leading space) — checked out at exactly the recorded commit; everything is as it should be. - `+<sha> path` — the submodule's checked-out commit differs from the one recorded in the superproject. Either you moved it deliberately, or you have a pending bump to commit. - `U<sha> path` — the submodule has unresolved merge conflicts. ## Making it not happen again The same emptiness reappears every time the superproject moves to a commit whose gitlink differs from what is on disk — after a `git pull`, after switching branches. By default Git leaves the submodule where it was and reports the difference; it does not silently move it. Setting `submodule.recurse` to `true` in local config makes the commands that touch the working tree — including `git pull` and `git checkout` — recurse into submodules and update them to the recorded commits. It does not cover initialising brand-new submodules that appeared, so the bootstrap command is still worth keeping. Related useful settings: `status.submoduleSummary` for `git status` to summarise submodule movement, and `fetch.recurseSubmodules` to control whether fetching the superproject also fetches submodule objects. ## What good onboarding looks like Because this problem is entirely predictable, teams that use submodules successfully do three things: they document the recursive clone command instead of a plain one; they provide a bootstrap script that runs `git submodule update --init --recursive` and is safe to re-run; and they make the build fail with an explicit message when a submodule path is empty, rather than with a confusing compiler error. None of that is Git functionality — it is the operational discipline the data model requires. ## Related cleanups `git submodule deinit <path>` reverses initialisation: it removes the working-tree content and the local config entry, leaving the gitlink untouched. It is the right way to temporarily stop carrying a submodule you do not need, and `--force` handles the case where the submodule's working tree has local modifications.
- What does --init add in git submodule update --init?It performs the initialisation step: reading `.gitmodules` and copying each submodule's URL into your local `.git/config` under `submodule.<name>.url`. Without it, `update` skips any submodule it considers uninitialised, so a fresh clone would still end up with an empty directory. It is the same work `git submodule init` does, folded into one command.
- After a git pull, a submodule still points at the old commit. Why doesn't Git move it?By default Git updates the superproject's working tree, including the recorded gitlink in the index, but leaves the submodule's own checkout alone — moving it could discard work you have in progress there. `git submodule status` shows it as `+` to flag the difference. Run `git submodule update`, or set `submodule.recurse` to true so pulls and checkouts recurse automatically.
- How do you tell from the command line whether a submodule is initialised?`git submodule status` prints a leading marker per submodule: `-` means not initialised, a leading space means it is checked out at exactly the recorded commit, `+` means the checked-out commit differs from the recorded one, and `U` means unresolved conflicts. It is the fastest first diagnostic when a submodule-related build failure appears.
saying these in an interview costs you the question
- Thinks the clone failed or the network dropped files
- Says the submodule content is missing from the remote
- Reaches for deleting and re-cloning instead of initialising
- Believes git pull always updates submodules by default
- Confuses initialising a submodule with adding a new one