How does a Git partial clone with --filter=blob:none differ from a shallow clone?
answer
- Two different things get removed
- One cuts commits, the other cuts contents
- Missing objects arrive later, from somewhere
- The remote makes a promise
- blob:none versus --depth
basics
~20 sA shallow clone truncates history to a few commits. A partial clone with --filter=blob:none keeps the entire commit and tree graph but omits file contents, downloading each blob lazily from the remote the first time something actually needs it.
solid answer
~40 sThey cut along different axes. `--depth` removes **commits**: you get a truncated graph, so log, blame and merge-base stop at the boundary. `--filter=blob:none` removes **file contents**: the full commit and tree history arrives, so history commands work normally, but blobs are fetched on demand from the *promisor remote* the first time a checkout, diff or blame needs them. Git marks this by setting `extensions.partialClone`, `remote.origin.promisor=true` and `remote.origin.partialclonefilter` in the config. `--filter=tree:0` is more aggressive and omits trees as well, which is ideal when a job only walks commit metadata. The server must permit it (`uploadpack.allowFilter`). The tradeoff is that lazy fetches are network round-trips: a job that touches many old files can end up slower than a full clone, and a partial clone is unusable offline.
code
bash · 8 linesgit clone --filter=blob:none https://example.com/app.git
cd app
git config --get remote.origin.promisor
# true
git config --get remote.origin.partialclonefilter
# blob:none
git log --oneline | wc -l # full history is present
git rev-list --objects --missing=print HEAD | head -3go deeper
Know that both flags make clones smaller but in different ways: --depth drops old commits, --filter=blob:none defers file contents until something needs them.
Explain promisor remotes and lazy fetching, name the config keys a partial clone writes, and contrast which commands each option degrades.
Judge the failure mode you are adopting — deterministic local breakage versus a network-dependent long tail — and measure whether the filter actually took effect against your server.
Decide where in the estate each strategy belongs, including whether making every developer's clone depend on the host for object retrieval is an acceptable availability coupling.
## Two different economies A clone is expensive for two reasons: the number of commits, and the number of file versions those commits reference. Git offers a separate lever for each. **Shallow clone** (`--depth <n>`) removes history. You receive a handful of commits and the objects needed to check them out. Ancestry stops at the boundary recorded in `.git/shallow`. **Partial clone** (`--filter=<spec>`) removes object *content* while keeping the graph. With `--filter=blob:none` you get every commit and every tree in the history but no blobs except the ones needed for the initial checkout. History commands behave exactly like a full clone because they only need commits and trees. ## Promisor remotes and lazy fetch A partial clone is allowed to be incomplete, which normally would be corruption. Git records the exception in the config: `extensions.partialClone` names the remote that promises to supply missing objects, `remote.origin.promisor=true` marks that remote, and `remote.origin.partialclonefilter` stores the filter used. When a command needs an absent object, Git transparently fetches it from the promisor remote and continues. `git checkout` of an old commit, `git blame` on a long file, or `git diff` against an ancient revision each trigger such fetches. You can see what is missing with `git rev-list --objects --missing=print`. ## The filter vocabulary - `--filter=blob:none` — no blobs at all up front. The common choice. - `--filter=blob:limit=1m` — blobs under the size limit come along; large ones are lazy. Good for repositories where a few binaries dominate. - `--filter=tree:0` — no trees and therefore no blobs. The cheapest clone that still has the whole commit graph, useful for jobs that only read commit metadata, but almost every path-aware command triggers fetching. The server must advertise support; hosts enable it with `uploadpack.allowFilter`. If it is not enabled, the filter is silently ignored and you get a full clone. ## Which failure mode you are buying Shallow clone fails *loudly and locally*: `git describe` errors, merge-base is missing, blame is wrong. The failures are deterministic and offline. Partial clone fails *slowly and remotely*: everything works, but some operations pause for a network round-trip, and each pause depends on the promisor remote being reachable. A `git blame` across a decade of a file can issue a long sequence of fetches. On a laptop that goes offline, a partial clone can stop being usable for operations a full clone would have served locally. ## Combining them The two are orthogonal and can be combined — `git clone --depth 1 --filter=blob:none` — though the marginal saving over either alone is often small, and you inherit both failure modes. Partial clone also pairs naturally with sparse checkout: filter the blobs at transfer time, and materialize only the directories you need in the working tree. ## Rule of thumb If the job needs the *graph* but not most *content* — changed-file detection, version derivation from tags, history queries — prefer a partial clone. If the job needs only the *current tree* — compile, lint, package — a shallow clone is simpler and has no lazy-fetch tail. If the job needs both deep history and lots of content, clone fully and cache the clone between runs.
- What does --filter=tree:0 buy you over --filter=blob:none, and what does it cost?`tree:0` omits directory objects as well as file contents, so the clone carries only the commit graph. It is the cheapest option for jobs that just read commit metadata, such as counting commits or reading messages. The cost is that any path-aware command — checkout, diff, blame — must fetch trees before it can even discover file names.
- What happens if the server does not support filtering?Filtering is a negotiated capability; a server enables it with `uploadpack.allowFilter`. If the server does not advertise it, the filter has no effect and you receive an ordinary full clone. A pipeline that assumed a cheap clone therefore silently gets an expensive one, which is worth measuring rather than assuming.
- Is a partial clone safe to work in offline?Only for the objects you already have. Any operation that reaches an omitted blob needs the promisor remote, so offline blame, diff against old revisions, or checkout of an old commit can fail. If offline work matters, either clone fully or pre-warm the objects you expect to need while still connected.
Shallow clone is buying only this month's issue; partial clone is buying the full index of every article but downloading each article's text only when you open it.
saying these in an interview costs you the question
- Says partial clone truncates history like --depth
- Thinks missing blobs are gone forever
- Assumes the filter always applies regardless of server support
- Ignores that lazy fetch needs network on every miss
- Believes a partial clone is always faster than a full clone