In Git, what does a commit record about its predecessors, and why is history a DAG?
answer
- Look at what a commit object contains
- Count the parent lines
- Order of parents is not arbitrary
- Hash includes the parent IDs
- Edges point backwards only
basics
~20 sEach Git commit stores an ordered list of parent commit IDs: zero for a root commit, one for an ordinary commit, two or more for a merge. Following those pointers backwards yields a directed acyclic graph, not a straight line.
solid answer
~40 sA Git commit object holds a `tree` (the snapshot), author and committer metadata, a message, and zero or more `parent` lines. The first commit in a repository has no parent, an ordinary commit has exactly one, a merge commit has two, and an octopus merge has more. A commit's ID is a hash over its own content **including** its parent IDs, so a commit can never point forward and can never gain a parent afterwards — changing a parent produces a different commit. That is why history is a **directed acyclic graph**: every edge points from child to parent, and no cycle can form. Branches and tags are just names pointing into that graph, so `git log`, ranges and reachability are all traversals over parent edges.
code
console · 8 lines$ git cat-file -p HEAD
tree 8f4b1c2a9d3e5f6071829304a5b6c7d8e9f01234
parent 3a1b2c3d4e5f60718293a4b5c6d7e8f901234567
parent 9f8e7d6c5b4a39281706f5e4d3c2b1a098765432
author Ada <[email protected]> 1710000000 +0000
committer Ada <[email protected]> 1710000000 +0000
Merge branch 'feature'go deeper
Be ready to say a commit points at a snapshot plus its parent commits, and that a merge commit has two parents. Knowing the arrows point backwards in time is most of the answer.
Explain that the commit ID hashes the parent IDs, so the graph cannot contain a cycle and any rewrite produces new IDs for the whole descendant chain.
Show how reachability over parent edges underpins release questions, range selection, and garbage collection, and how first-parent ordering keeps merge history readable in production repos.
Be able to argue what graph shape a team should aim for — how much branching and merging the DAG should carry before history stops answering the questions people actually ask of it.
## What a commit actually stores A commit in Git is a small object in the object database. Its content is: one `tree` line naming the root tree object (the whole project snapshot at that point), zero or more `parent` lines naming preceding commits, an `author` line, a `committer` line, optionally a signature, then a blank line and the message. `git cat-file -p HEAD` prints it verbatim. Notice what is absent: no diff, and no branch name. A diff is computed on demand by comparing one commit's tree against another's. A branch name lives in a separate ref file and merely points at a commit; the commit does not know which branches contain it. ## Zero, one, or many parents - **Root commit** — no `parent` line at all. Usually there is exactly one root, but a repository can have several (for example after grafting unrelated histories together, or when a branch was started with `git checkout --orphan`). - **Ordinary commit** — exactly one parent: the commit that was `HEAD` when you ran `git commit`. - **Merge commit** — two or more parents. The order matters and is not arbitrary: the **first parent** is the commit you were on when the merge ran, the second is the branch you merged in. A merge of more than two branches at once (an octopus merge) records three or more parents. ## Why "directed" and why "acyclic" Edges are *directed* because a commit records its parents and nothing records its children. Finding the children of a commit is a search: Git must walk the graph from the refs and see who lists it as a parent. That is exactly why `git log` walks backwards in time by default. The graph is *acyclic* by construction, not by a rule someone enforces. The commit ID is a cryptographic hash of the commit's bytes, and those bytes contain the parent IDs. To make commit A a parent of commit B, A's hash must already exist, so A must already be complete. A cycle would require a commit to contain its own hash, which cannot be constructed. An important corollary: because parents are hashed into the ID, rewriting any commit (amend, rebase, filter) necessarily produces a new ID for it **and** for every descendant. History is immutable; "rewriting" means building a new subgraph and moving refs onto it. ## What the graph buys you Everything higher-level is defined in terms of reachability over these edges. "Is this fix in the release?" is "is commit X reachable from tag v2.0?". Range syntax like `main..feature` is a set difference over reachable sets. Garbage collection keeps objects reachable from refs and reflogs and discards the rest. `git log --graph --oneline` draws the DAG in ASCII, which is the fastest way to see the shape of a merge. ## Frequent misconceptions Candidates often say a commit "stores the changes". It stores a full snapshot tree; the storage layer deduplicates unchanged content and packs deltas, but that is a storage detail, not the data model. Another is that a merge commit contains the merged content "twice" — it has one tree like every other commit, and two parents. A third is that a commit knows its branch; it does not, which is why deleting a branch can strand commits that nothing else points at.
- How many parents can a single commit have, and what produces more than two?Zero for a root commit, one for a normal commit, two for a standard merge. Three or more comes from an octopus merge — merging several branches in one `git merge` invocation. Octopus merges are rare because Git refuses them when any of the branches conflict.
- Why does amending an old commit change the IDs of every commit after it?A commit's ID is a hash of its content, which includes its parent IDs. Change anything in a commit and its ID changes; its child now records a different parent, so the child's hash changes too, and so on down the chain. Rewriting always produces a new subgraph.
- Given a commit, how does Git find its children?It does not store them. Git walks backwards from the refs (branches, tags, remote-tracking refs, and the reflog) and collects commits that list the target as a parent. That is why `git log <commit>` shows ancestors instantly but "what came after this" needs a full traversal.
Each commit is a photo of the whole project with a note on the back listing the photo (or photos) it came after — you can always walk backwards, never forwards.
saying these in an interview costs you the question
- Says commits store diffs rather than snapshot trees
- Thinks a commit knows which branch it is on
- Believes history is a linked list, never branching
- Claims parents of a merge are unordered
- Thinks amending edits a commit in place