What is Git's commit-graph file, and which operations does it make faster?
answer
- parsing commit objects is the bottleneck
- a sidecar index of commit metadata
- numbers that allow early cut-off
- helps merge-base and --contains most
- derived data, safe to delete
basics
~20 sThe commit-graph is a derived binary cache under .git/objects/info that stores each commit's parents, root tree and date plus generation numbers, so Git can answer history and reachability queries without decompressing commit objects from packs.
solid answer
~50 sWalking history normally means reading and inflating every commit object out of a packfile just to learn its parents and date. The **commit-graph** precomputes that: a binary file (at `.git/objects/info/commit-graph`, or a chain under `commit-graphs/`) listing each commit's object ID, root tree ID, parents, commit date and a generation number. Generation numbers let Git cut off traversals early when answering reachability questions, which is where the big wins are — `git log --graph` and `--topo-order`, `git merge-base`, `git branch --contains`, and ahead/behind counts. It is written by `git commit-graph write --reachable`, by `git gc` when `gc.writeCommitGraph` is on, and by the background `git maintenance` tasks; `core.commitGraph` controls whether Git reads it. It is pure cache: it holds no unique data, deleting it only costs speed, and commits created since the last write simply fall back to being parsed normally.
code
bash · 3 linesgit commit-graph write --reachable
ls .git/objects/info/commit-graph
git -c core.commitGraph=false log --oneline --graph | headgo deeper
Just recall that Git can keep a cache file describing commit parents so history commands run faster, and that it is generated locally rather than downloaded.
Explain what is stored — parents, root tree, date, generation numbers — and why that avoids inflating and parsing commit objects during a walk.
Demonstrate operational use: know that gc, fetch and maintenance can write it, that core.commitGraph toggles reading, and that deleting it is a safe first diagnostic because it is pure derived data.
Frame it against the other derived structures — multi-pack-index and bitmaps — and decide which acceleration a given repository's dominant workload actually needs.
## The problem it solves Every question about history shape — what are the parents of this commit, which commits are common ancestors, is A an ancestor of B, how far ahead is this branch — is a walk over the commit DAG. To take one step of that walk from a plain object store, Git must locate the commit object (a `.idx` lookup), read it out of the packfile (possibly reconstructing a delta chain), zlib-inflate it, and parse the text to find the `parent` lines and the committer date. That is a lot of work for a few dozen bytes of structure, repeated once per commit, in a repository that may have hundreds of thousands of them. ## What the file contains The commit-graph is a purpose-built binary file recording, for each commit it covers: the commit's object ID, its root tree's object ID, the list of parent object IDs (encoded compactly, with an overflow representation for octopus merges), the commit date, and a **generation number**. Entries are laid out for direct indexed access, so a parent lookup is a pointer chase in a memory-mapped file rather than an object read. The generation number is the interesting part. In its basic form it is one plus the maximum generation of the commit's parents — the length of the longest path back to a root commit. Because a commit can never reach an ancestor with a generation greater than or equal to its own, Git can prune huge parts of a traversal without visiting them: if you are asking whether A reaches B and B's generation exceeds A's, the answer is no, immediately. Recent Git also records corrected commit dates so that traversals ordered by date get the same early-exit benefit even when clock skew has produced commits whose recorded dates lie. ## Which commands get faster Anything reachability-shaped: `git log --graph` and `git log --topo-order`, which need to know the DAG shape before printing anything; `git merge-base` and therefore every merge and rebase that has to find one; `git branch --contains` and `git tag --contains`; ahead/behind counts, such as the "your branch is ahead by N commits" line; and `git log --oneline` over deep history. Commands dominated by *content* — diffs, blame on a large file, checkout — see much less, because their cost is in trees and blobs, which the commit-graph does not describe. ## Writing and reading it `git commit-graph write --reachable` writes a file covering every commit reachable from refs. `--split` writes an incremental layer instead of a full rewrite, producing a chain of files under `.git/objects/info/commit-graphs/` with a `commit-graph-chain` listing them; the layers are merged occasionally so the chain does not grow without bound. Modern Git writes the graph as part of `git gc` when `gc.writeCommitGraph` is enabled, and `fetch.writeCommitGraph` updates it after fetching. The `git maintenance` machinery includes a commit-graph task for scheduled background upkeep. On the read side, `core.commitGraph` decides whether Git consults the file at all; turning it off is the first diagnostic step if you suspect the cache. ## It is a cache, and that has consequences Nothing in the commit-graph is unique data — every field is derivable from the commit objects themselves. Deleting the file is always safe and costs only speed. Because it is written at a point in time, it covers the commits that existed then: commits created afterwards are parsed the old way, so a repository is never "stale" in a correctness sense, merely partially accelerated until the next write. For the same reason the graph only ever describes commits, not trees, tags or blobs, and it is not part of what a clone transfers — each repository builds its own. One related but distinct file deserves separating in an interview: the **multi-pack-index** (`git multi-pack-index write`) accelerates *object lookup* across many packfiles, and pack **bitmaps** (`git repack --write-bitmap-index`) accelerate computing which objects to send. All three are derived acceleration structures living beside the object store, but only the commit-graph is about the shape of history. ## How to talk about it The interview-grade summary is: it is a derived index over commit metadata with generation numbers that enable early termination in reachability queries; it makes DAG-shaped commands fast on large repositories; it is written by gc, fetch or maintenance and read under `core.commitGraph`; and because it is pure cache, the correct response to any suspicion about it is to disable or delete it and compare.
- What is a generation number and why does it speed up reachability queries?It is a per-commit value, one more than the maximum of its parents' values, so it bounds how far back a commit can reach. If you are asking whether A reaches B and B's generation is not smaller than A's, the answer is no without visiting a single parent — which lets Git terminate large traversals early instead of walking to the root.
- Is the commit-graph part of what a clone downloads?No. It is derived local data built from the commit objects the repository already has, so each repository writes its own — typically during gc, after a fetch when `fetch.writeCommitGraph` is set, or via a scheduled `git maintenance` task. Deleting it costs only performance, never correctness.
- How does the commit-graph differ from a multi-pack-index or a pack bitmap?All three are derived acceleration files, but they answer different questions. The commit-graph describes the shape of history — parents, dates, generation numbers. A multi-pack-index speeds up finding which pack holds a given object across many packs. Bitmaps speed up computing the set of objects reachable from a ref, which mainly helps object-set computations.
saying these in an interview costs you the question
- Thinks the commit-graph stores commit contents or diffs
- Believes it must be fetched from the remote to be valid
- Says deleting it corrupts or loses history
- Assumes it speeds up diff and blame equally
- Confuses it with the packfile .idx index